Ku laabo Warka
Hal-abuurnimoAI Understanding warbixin kooban

Warqada DiREAG waxay soo jeedinaysaa hab wanaagsan oo lagu qiyaaso kalsoonida AI ee jawaabaha xisaabta

Warqad cusub oo arXiv ah ayaa soo bandhigaysa DiREAG, hab isku dara warbixinnada kalsoonida ee dhowr jawaabood oo caddaynaya jawaabaha suurtagalka ah, oo ay ku jirto suurtagalnimada in midna uusan sax ahayn.

5 min readRead the primary source
Primary-source image accompanying DirEAG paper proposes a better way to calibrate AI confidence in math answers
Dukumeentiga isha aasaasiga ahIsha la duubay
Daabacaha
arxiv.org
Xidhiidhka isha
arxiv.orghttps://arxiv.org/abs/2608.20717
Nooca isha
Dukumeentiga aasaasiga ah - ogeysiis rasmi ah, warqad, xereyn, ama bogga xisbiga koowaad waxaan si toos ah u akhrinay.
Dulucda sheekadaKu fahan tan 60 ilbiriqsi gudahood

Halkan ka bilow

Qodobbada muhiimka ah

Guud ahaan
Sida ugu wanagsan ee moodalku u qabto xogta cusub, ee aan la arkin ee ka baxsan habka tababarka.
Calibration
Sida fiican buundooyinka kalsoonida moodeelku u dhigmaan ixtimaalka saxda ah ee dhabta ah.
Benchmark
Tijaabo la habeeyey ama kayd xogeed oo loo isticmaalo in lagu cabbiro laguna barbar dhigo waxqabadka moodeelka.
Is tijaabiChatGPT & LLMs Kedis

Maxaa dhacay

Researchers propose DirEAG, a Dirichlet Evidence Aggregation method for calibrating verbalized confidence from language models solving mathematical problems. The paper reports better than direct averaging and heuristic aggregation across three datasets and three model families, while maintaining competitive answer selection.

A paper posted to arXiv on Aug. 21, 2026, proposes DirEAG, short for Dirichlet Evidence Aggregation. Its subject is a specific weakness in large language models used for mathematical reasoning: asking a model how confident it is does not automatically produce a confidence score that corresponds reliably to correctness. The authors describe black-box verbalized confidence as difficult to calibrate because the numerical meaning of a model’s answer can shift with the prompt, the model, or the dataset.

The method uses multiple confidence-steering prompts on the same problem. Each prompt produces an answer-confidence observation. Rather than simply averaging those confidence values, DirEAG converts each observation into soft evidence distributed across the candidate answers generated during the process. It also adds a null state representing the possibility that none of the candidates is correct. This design matters because it treats confidence reports as evidence about competing answers instead of assuming that every reported percentage is already on a common, meaningful scale.

The paper reports experiments on GSM8K, SVAMP, and GSM-Hard, three mathematical reasoning datasets named in the source. It tests models from the Qwen, Mistral, and Gemma families. According to the authors, DirEAG often produces better than direct confidence averaging and heuristic confidence-steering aggregation, while preserving competitive performance in selecting an answer. The source does not state the exact calibration scores, answer-selection scores, model versions, or experimental settings.

The authors also report ablation results indicating that two parts of the approach address different problems. Evidence aggregation combines the information from the answer-confidence observations, while a final binary- step handles another part of the calibration task. The paper is 16 pages long, contains three figures, says code is available, and is listed as accepted by PRICAI 2026. The source identifies the work as a preprint and does not describe independent replication or peer-reviewed results beyond that acceptance statement.

Faahfaahinta isha: arxiv.org ↗

Maxay muhiim u tahay

Language-model confidence statements can be difficult to interpret, especially when prompting changes the scale of the model’s self-reported certainty. A method that better separates answer selection from uncertainty estimation could help developers identify when a mathematical answer deserves additional checking, although the source provides no evidence of deployment beyond the reported experiments.

The practical issue is not simply whether a language model can produce a mathematical answer. It is whether a user or a downstream system can interpret the model’s stated certainty when deciding what to trust. If confidence values change meaning across prompts, then a high number in one setting may not be comparable with a high number in another. The paper’s central contribution is an attempt to address that comparability problem in a structured way.

DirEAG’s null state is a useful feature of the proposal because it allows the method to represent uncertainty that is not resolved by the candidate answers under consideration. A system that must choose among listed answers can otherwise appear more certain simply because it has no explicit place to express that all available candidates may be wrong. The source presents this as part of the method; it does not show whether the null state improves safety or decision-making in deployed systems.

The separation between answer selection and also gives the research a potentially useful diagnostic angle. A model can select the correct answer relatively often while still expressing confidence poorly, or it can have calibrated confidence while selecting answers less effectively. The reported ablations suggest that the authors do not treat those as the same objective. That distinction could be relevant to developers evaluating systems that need both correct outputs and reliable signals for when to request review.

The evidence remains limited to the paper’s own experiments. The source reports results on mathematical benchmarks and several model families, but it does not establish that DirEAG improves confidence estimates for coding, scientific analysis, medical questions, or other tasks. It also does not report how the method compares with every other uncertainty-estimation approach, whether it adds meaningful computational or latency costs, or whether its gains are large enough to change operational decisions. Those limits keep the result in the category of promising research rather than demonstrated production capability.

Interactive Mechanism

Farsamaynta Is-dhexgalka: Sida Dhabta Ay U Shaqeyso

U baadh tignoolajiyada hoose ee ka dambeeya horumarkan si isdhexgal leh.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Hubinta Fikradda Is-dhexgalka+10 Points
ChatGPT & LLMs Quiz

What is a common training objective for an autoregressive language model?

Maxaa la daawan doona xiga

The key next questions are whether DirEAG holds up outside the tested mathematical benchmarks, how large its gains are, and whether it remains useful across model sizes, prompting strategies, and real applications. The source does not provide numerical results, implementation details, or evidence from operational systems.

The first verification point is quantitative. The abstract says DirEAG often achieves better and competitive answer selection, but it does not provide the size, consistency, or statistical significance of those gains. Readers should look to the full paper and released code for the specific calibration metrics, confidence ranges, baselines, datasets splits, model checkpoints, and repeated-run results needed to assess how robust the claim is.

A second question is . GSM8K, SVAMP, and GSM-Hard all concern mathematical reasoning, so the reported evidence does not show whether the method works when answers are open-ended, evidence is incomplete, or correctness is harder to define. Testing across additional domains would clarify whether DirEAG captures a general property of verbalized model confidence or mainly addresses the structure of these benchmarks.

The method may also depend on how candidate answers and confidence-steering prompts are generated. The source does not say how many prompts are used, how candidate answers are selected, how the null state is parameterized, or how the final binary is trained. Those details could affect cost, reproducibility, and performance. It will be important to see whether the approach remains effective when prompts, model versions, or datasets change.

Finally, the relevant practical test is whether calibrated confidence improves decisions rather than only metrics. The source does not report deployment, user studies, human-review outcomes, or integrations with mathematical tools. Further work should examine whether DirEAG helps people identify incorrect answers, allocate verification effort, or avoid over-trusting a model, while also measuring any added complexity and failure cases.

Tilmaamaha la xidhiidha & su'aalaha

ChatGPT iyo LLMsMoodooyinka AI ayaa la sharaxayTababarka AIAnshaxa AITijaabi waxaad taqaan - isku day kedis AI oo bilaash ahKa raadi erey AI qaamuuskeenaRaac qaabka AI raadraaca sii deynta
Tan faa'iido ma u heshay?