What Does the BenchMIRT Benchmark Actually Measure in LLMs?
AI News Brief: BenchMIRT is a new method for breaking down what LLM benchmarks actually measure, auditing them at the individual question level. Traditional benchmarks usually claim to measure a single capability (like safety, general reasoning, or instruction following), but individual questions often involve multiple capabilities at once. For example, BBQ, a benchmark that tests whether models rely on social stereotypes, has a question about a grandparent and grandchild booking an Uber that looks like it’s testing age bias, but actually also requires the model to correctly track character roles and reason from evidence rather than assumptions.