返回新聞
創新AI Understanding 簡報

BenchMIRT 發現 LLM 基準可以混合安全和推理訊號

Ai2 推出了 BenchMIRT,這是一種分析單一基準測試問題以確定其衡量哪些能力的方法。研究人員在 16 個基準測試中測試了 100 個法學碩士,發現一些評估結合了安全性和一般推理訊號,並且較小的問題集通常可以保留大部分…

5 min readRead the primary source
Source-provided image accompanying BenchMIRT finds LLM benchmarks can mix safety and reasoning signals
主要來源文件來源記錄
出版商
huggingface.co
來源連結
huggingface.cohttps://huggingface.co/blog/allenai/benchmirt
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
重量
一個學習的數值,用來縮放通過神經網路的訊號。
測試一下自己AI 模型解釋測驗

發生了什麼事

Ai2 introduced BenchMIRT, a multidimensional method for auditing large language model benchmarks at the level of individual questions. The analysis used results from 100 open- LLMs across 16 benchmarks and more than 34,000 questions.

Ai2 says BenchMIRT extends Item Response Theory, a technique for inferring underlying abilities from patterns of test responses. Its multidimensional version estimates each model’s strength across latent capabilities, while also estimating each question’s difficulty and its ability to distinguish models that are stronger or weaker on those capabilities. The method is applied to both the model and question levels, allowing researchers to inspect what individual items appear to measure rather than relying only on an aggregate score.

The researchers trained BenchMIRT using results from 100 LLMs across 16 benchmarks containing more than 34,000 questions. Six benchmarks focused on general reasoning, including MMLU-Pro, GPQA, MATH, and BBH. Ten came from Ai2’s Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP, and XSTest. According to the source, the researchers did not tell the method which benchmarks represented which capabilities. It independently recovered two dominant dimensions: safety and general reasoning. Repeating the analysis from scratch produced the same two dimensions, which Ai2 presents as evidence that the result was stable within this set.

BenchMIRT largely matched the intended focus of many evaluations, but it also identified exceptions. BBQ, a intended to assess reliance on social stereotypes, aligned more strongly with general reasoning than with safety in the analysis. Ai2 says that a low BBQ score could therefore partly reflect difficulty tracking relationships or reasoning from the information in a question, rather than safety behavior alone. WMDP, which tests dangerous dual-use knowledge in biology, chemistry, and cybersecurity, was also more closely associated with general reasoning than safety. Because WMDP treats refusal or failure to provide dangerous information as the desired response, stronger general reasoning was associated with lower WMDP scores.

The analysis also found that a single can contain clusters of questions with different patterns. HarmBench’s standard and contextual harmful-request questions aligned more closely with safety, while its copyright questions were more closely associated with general reasoning. Ai2 does not conclude that these benchmarks are invalid. Its claim is narrower: aggregate scores can combine several kinds of information, and BenchMIRT can help identify which questions are driving those results.

來源詳情: huggingface.co ↗

為什麼這很重要

scores are often treated as measurements of a single capability, but BenchMIRT’s results suggest some scores combine multiple signals. That could affect how researchers compare models, design safety tests, and interpret claims about model capabilities.

The central implication is interpretive. A labeled as measuring safety may partly measure reading comprehension, reasoning, domain knowledge, or the ability to track context. If those components are not separated, a model’s score can be overread as evidence of a single capability. BenchMIRT offers researchers a way to inspect that mixture and describe results more precisely.

The method could also make evaluations more efficient. Ai2 ranked questions by how well they distinguished stronger from weaker models on the underlying dimensions, while retaining a mixture of easier and harder items. Across the 16 benchmarks, the source says that keeping 10% of the questions generally preserved nearly the same ordering of models on the inferred safety or reasoning abilities as the full set. Keeping 50% often matched the full ’s measure more closely. This could reduce evaluation cost or make repeated testing more practical, although the report does not establish how the result would transfer to other benchmark collections or newer models.

BenchMIRT also attempts to predict how a model would perform on a question that was not included in its observed responses. In Ai2’s experiments, it correctly predicted whether a model would answer a held-out question correctly 79% of the time, compared with 70% for a simpler approach that assumed the model would perform on each question about as well as on the overall. That result suggests question-level information can improve estimates of model behavior, but it is an experiment reported by the method’s creators, not an independent validation.

For the public, clearer interpretation matters because model scores influence how capabilities and safety are communicated. The source does not show that BenchMIRT changes any model’s real-world behavior, proves that a model is safe, or establishes a universal definition of reasoning or safety. It is an auditing and analysis method whose findings depend on the questions and models supplied to it.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The method needs testing on newer model generations and on different collections of benchmarks. Researchers will also need to weigh its efficiency and transparency benefits against the risk that question-level analysis could help developers remove difficult safety items and weaken evaluations.

The most important limitation is recency. Ai2 says every model used to train and evaluate BenchMIRT had been released by March 2025, so the source does not establish how well the method works on later generations of LLMs. Newer models could produce different response patterns, and the dimensions recovered from the selected set might not remain dominant.

choice is another open question. The dimensions BenchMIRT discovers depend on the evaluations it receives. Safety and general reasoning emerged from this project’s 16-benchmark mix, but a different set could surface other capabilities or produce different relationships among questions. The source does not report a broader comparison showing that the same structure appears across independent benchmark suites.

There is also a trade-off between efficiency and ranking accuracy. Ai2 says that when the objective is ranking models by predicted performance on randomly held-out items, the ’s average score performs slightly better than BenchMIRT. The method’s advantage is the finer-grained information it provides about individual questions and underlying capabilities, not necessarily a better overall ranking in every setting.

Finally, the transparency can be misused. The same question-level estimates that identify informative safety items could help someone remove or avoid those items and create a weaker evaluation that an unsafe model can pass. Ai2 acknowledges that risk and says existing tools already enable similar trimming. Future work should clarify how designers can publish useful diagnostics while protecting evaluation integrity.

相關指引和測驗

人工智慧模型解釋人工智慧培訓AI 倫理測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?