RESEARCHMonitorWATCHLIST
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers introduced RepBench, a standardized benchmark dataset mapping internal LLM representations to 182 distinct capability taxonomies.
Open source