Pada si Iroyin
AtunseAI Understanding finifini

BenchMIRT rii awọn ami-ami LLM le dapọ ailewu ati awọn ifihan agbara ero

Ai2 ṣafihan BenchMIRT, ọna ti o ṣe itupalẹ awọn ibeere ala-kọọkan lati ṣe idanimọ iru awọn agbara ti wọn wọn. Idanwo awọn LLM 100 kọja awọn ipilẹ 16, awọn oniwadi rii pe diẹ ninu awọn igbelewọn darapọ ailewu ati awọn ifihan agbara-gbogboogbo ati pe awọn eto ibeere ti o kere julọ le nigbagbogbo ṣetọju pupọ ti…

5 min readRead the primary source
Source-provided image accompanying BenchMIRT finds LLM benchmarks can mix safety and reasoning signals
Iwe aṣẹ orisun akọkọOrisun ti o gbasilẹ
Olutẹwe
huggingface.co
Orisun ọna asopọ
huggingface.cohttps://huggingface.co/blog/allenai/benchmirt
Orisun iru
Iwe akọkọ - ikede osise, iwe, iforukọsilẹ, tabi oju-iwe ẹgbẹ akọkọ ti a ka taara.
AtokọLoye eyi ni iṣẹju 60

Bẹrẹ nibi

Awọn ofin bọtini

Awoṣe Ede nla (LLM)
Awoṣe ede ti a ṣe ikẹkọ lori titobi ọrọ corpora lati ṣe ipilẹṣẹ ati itupalẹ ọrọ.
Aṣepari
Idanwo idiwon tabi data ti a lo lati ṣe iwọn ati ṣe afiwe iṣẹ awoṣe.
Iwọn
Iye nọmba ti o kọ ẹkọ ti o ṣe iwọn awọn ifihan agbara ti n kọja nipasẹ nẹtiwọọki nkankikan.
Ṣe idanwo fun ara rẹAwọn awoṣe AI ti ṣalaye adanwo

Kini o ṣẹlẹ

Ai2 introduced BenchMIRT, a multidimensional method for auditing large language model benchmarks at the level of individual questions. The analysis used results from 100 open- LLMs across 16 benchmarks and more than 34,000 questions.

Ai2 says BenchMIRT extends Item Response Theory, a technique for inferring underlying abilities from patterns of test responses. Its multidimensional version estimates each model’s strength across latent capabilities, while also estimating each question’s difficulty and its ability to distinguish models that are stronger or weaker on those capabilities. The method is applied to both the model and question levels, allowing researchers to inspect what individual items appear to measure rather than relying only on an aggregate score.

The researchers trained BenchMIRT using results from 100 LLMs across 16 benchmarks containing more than 34,000 questions. Six benchmarks focused on general reasoning, including MMLU-Pro, GPQA, MATH, and BBH. Ten came from Ai2’s Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP, and XSTest. According to the source, the researchers did not tell the method which benchmarks represented which capabilities. It independently recovered two dominant dimensions: safety and general reasoning. Repeating the analysis from scratch produced the same two dimensions, which Ai2 presents as evidence that the result was stable within this set.

BenchMIRT largely matched the intended focus of many evaluations, but it also identified exceptions. BBQ, a intended to assess reliance on social stereotypes, aligned more strongly with general reasoning than with safety in the analysis. Ai2 says that a low BBQ score could therefore partly reflect difficulty tracking relationships or reasoning from the information in a question, rather than safety behavior alone. WMDP, which tests dangerous dual-use knowledge in biology, chemistry, and cybersecurity, was also more closely associated with general reasoning than safety. Because WMDP treats refusal or failure to provide dangerous information as the desired response, stronger general reasoning was associated with lower WMDP scores.

The analysis also found that a single can contain clusters of questions with different patterns. HarmBench’s standard and contextual harmful-request questions aligned more closely with safety, while its copyright questions were more closely associated with general reasoning. Ai2 does not conclude that these benchmarks are invalid. Its claim is narrower: aggregate scores can combine several kinds of information, and BenchMIRT can help identify which questions are driving those results.

Awọn alaye orisun: huggingface.co ↗

Kini idi ti o ṣe pataki

scores are often treated as measurements of a single capability, but BenchMIRT’s results suggest some scores combine multiple signals. That could affect how researchers compare models, design safety tests, and interpret claims about model capabilities.

The central implication is interpretive. A labeled as measuring safety may partly measure reading comprehension, reasoning, domain knowledge, or the ability to track context. If those components are not separated, a model’s score can be overread as evidence of a single capability. BenchMIRT offers researchers a way to inspect that mixture and describe results more precisely.

The method could also make evaluations more efficient. Ai2 ranked questions by how well they distinguished stronger from weaker models on the underlying dimensions, while retaining a mixture of easier and harder items. Across the 16 benchmarks, the source says that keeping 10% of the questions generally preserved nearly the same ordering of models on the inferred safety or reasoning abilities as the full set. Keeping 50% often matched the full ’s measure more closely. This could reduce evaluation cost or make repeated testing more practical, although the report does not establish how the result would transfer to other benchmark collections or newer models.

BenchMIRT also attempts to predict how a model would perform on a question that was not included in its observed responses. In Ai2’s experiments, it correctly predicted whether a model would answer a held-out question correctly 79% of the time, compared with 70% for a simpler approach that assumed the model would perform on each question about as well as on the overall. That result suggests question-level information can improve estimates of model behavior, but it is an experiment reported by the method’s creators, not an independent validation.

For the public, clearer interpretation matters because model scores influence how capabilities and safety are communicated. The source does not show that BenchMIRT changes any model’s real-world behavior, proves that a model is safe, or establishes a universal definition of reasoning or safety. It is an auditing and analysis method whose findings depend on the questions and models supplied to it.

Interactive Mechanism

Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ

Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Ibanisọrọ Erongba Ṣayẹwo+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Kini lati wo tókàn

The method needs testing on newer model generations and on different collections of benchmarks. Researchers will also need to weigh its efficiency and transparency benefits against the risk that question-level analysis could help developers remove difficult safety items and weaken evaluations.

The most important limitation is recency. Ai2 says every model used to train and evaluate BenchMIRT had been released by March 2025, so the source does not establish how well the method works on later generations of LLMs. Newer models could produce different response patterns, and the dimensions recovered from the selected set might not remain dominant.

choice is another open question. The dimensions BenchMIRT discovers depend on the evaluations it receives. Safety and general reasoning emerged from this project’s 16-benchmark mix, but a different set could surface other capabilities or produce different relationships among questions. The source does not report a broader comparison showing that the same structure appears across independent benchmark suites.

There is also a trade-off between efficiency and ranking accuracy. Ai2 says that when the objective is ranking models by predicted performance on randomly held-out items, the ’s average score performs slightly better than BenchMIRT. The method’s advantage is the finer-grained information it provides about individual questions and underlying capabilities, not necessarily a better overall ranking in every setting.

Finally, the transparency can be misused. The same question-level estimates that identify informative safety items could help someone remove or avoid those items and create a weaker evaluation that an unsafe model can pass. Ai2 acknowledges that risk and says existing tools already enable similar trimming. Future work should clarify how designers can publish useful diagnostics while protecting evaluation integrity.

Awọn itọsọna ti o jọmọ & awọn ibeere

Awọn awoṣe AI ti ṣalayeAI IkẹkọÌlànà Ìwà AIṢe idanwo ohun ti o mọ — gbiyanju idanwo AI ọfẹ kanWa ọrọ AI kan ninu iwe-itumọ waTẹle olutọpa idasilẹ awoṣe AI
Ṣe eyi wulo?