뉴스로 돌아가기
혁신AI Understanding 브리핑

New benchmark shows Vietnam’s exam rubric can change how language models rank

A paper introduces THPT-Ladder, a 632-item benchmark that applies Vietnam’s 2025 national exam grading scheme to language models and reports materially different scores from standard proportional-accuracy measures.

6 min readRead the primary source
Source-page capture accompanying New benchmark shows Vietnam’s exam rubric can change how language models rank
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.18336
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
정규화
최적화 안정성을 향상시키기 위해 값을 일관된 규모로 변환합니다.
견고성
소음, 교대 또는 적대적인 입력 하에서 성능을 유지하는 모델의 능력입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Researchers introduced THPT-Ladder, a built from 632 items in 21 official Vietnamese exams across 11 subjects. It applies the 2025 National High School Graduation Examination’s convex scoring rubric, under which four true/false statements earn 0, 0.10, 0.25, 0.50, or 1.00 points rather than proportional credit. Across eight models, the paper reports lower official-rubric scores than proportional scoring and shows that rankings against human candidates can change substantially.

The arXiv record describes a problem with translating exam performance into a single accuracy number. In Part II of Vietnam’s 2025 National High School Graduation Examination, each question contains four true/false statements. The official rubric assigns 0 points when none are correct, then 0.10, 0.25, 0.50, or 1.00 points as the number of correct statements rises from one to four. That schedule is convex rather than additive: three correct statements receive 0.50 points, even though three out of four statements would correspond to 0.75 under proportional credit. The paper treats this difference as an evaluation issue, not merely a grading detail, because Part II accounts for 4.00 of the exam’s 10.00 points. The source establishes these rules as the paper’s description of the 2025 reform; it does not independently verify the examination policy outside the cited paper record.

The authors say they created THPT-Ladder with 632 items drawn from 21 official exams covering 11 subjects. The is reported to grade answers exactly as the ministry grades its students. The paper also says the ministry publishes marks for more than one million candidates, allowing the researchers to place model results within a human candidate cohort. Across eight models, the official rubric reportedly produces scores between 0.020 and 0.159 points lower per Part II question than proportional credit. The abstract does not list all eight models, explain how each model was prompted, identify the number of runs, or describe whether answers were generated under identical settings. Those details are important for interpreting the numerical comparisons and are not supplied by the authoritative text provided here.

The abstract gives two examples of the reported ranking effect. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall is said to move the model from the 90th to the 77th percentile among 481,293 candidates. The paper also reports that a model at Claude Sonnet 5’s accuracy level could receive between 0.869 and 0.932 points per question, depending on how its errors are distributed. In the authors’ account, accuracy alone does not predict the penalty because the official score depends on which statements are answered correctly together within each question. The source does not provide the underlying response files, confidence intervals, item-level breakdowns, or an independent audit of those calculations, so these findings should be treated as reported preprint results rather than settled measurements.

소스 세부정보: arxiv.org

왜 중요한가요?

The paper argues that ordinary accuracy metrics can describe a model as more capable than an institution’s actual grading rules would certify. That matters for AI evaluations built from human exams, especially when results are used to compare models or make claims about educational competence. The findings are limited to the reported benchmark and models; the supplied source does not establish how broadly the result generalizes beyond Vietnam’s exam format.

The central implication is methodological. A can preserve the appearance of rigor while measuring a different construct from the one used by the institution whose exam it borrows. Under proportional scoring, each additional correct statement is treated as contributing the same amount. Under the Vietnamese scheme described by the paper, partial knowledge is rewarded unevenly, and some combinations of correct and incorrect statements receive less credit than their raw accuracy would suggest. A model’s final score therefore depends not only on how many statements it gets right, but also on the pattern in which those statements occur. For readers comparing models, that means a leaderboard based only on aggregate accuracy may not answer the practical question of how a model would perform under the official certification rule.

The reported human-cohort comparison makes the issue more concrete. The paper says Qwen3.5-27B’s 0.042-point difference on the 2025 History exam changes its position from the 90th to the 77th percentile among 481,293 candidates. If reproduced, that is not a cosmetic change in presentation: it alters the population comparison a reader would draw from the same model responses. The result does not show that the model became less capable, nor does it establish that percentile placement predicts classroom or examination outcomes. It shows that the choice of scoring rule changes the measurement.

The distinction is especially relevant when exam benchmarks are used to communicate claims about reasoning, knowledge, or readiness for educational tasks. The practical public impact is mainly on how AI performance is reported and interpreted. Researchers, educators, policymakers, and journalists may rely on scores without seeing how a grading formula handles partial answers. The paper’s argument supports reporting the official score alongside accuracy when the benchmark is intended to represent a real examination. It also suggests that evaluation designers should disclose whether a benchmark’s scoring rule is additive, proportional, thresholded, or otherwise non-linear. At the same time, the supplied source does not show that THPT-Ladder is representative of all language-model testing, that Vietnam’s scheme is common internationally, or that a convex rubric is inherently superior. The contribution is a warning about construct validity and comparability, not evidence that one universal scoring method should replace another.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Models Explained Quiz

What is the best response when AI Models Explained makes a mistake in production?

다음에 무엇을 볼 것인가

The key follow-up is whether the full study documents its prompts, model versions, sampling conditions, scoring implementation, statistical uncertainty, and data availability. Researchers should test whether the same gap appears under other non-additive grading systems and on independently reproduced runs. The paper is an arXiv preprint, so the supplied source does not establish peer-review status or independent replication.

The first verification question is reproducibility. The full paper should clarify the exact 632-item composition, the 21 exams and 11 subjects included, answer , handling of abstentions or ambiguous outputs, and the implementation of the ministry’s rubric. It should also identify all eight models and their precise versions, because model names alone may not determine behavior. Prompt wording, system instructions, decoding settings, date of access, and the number of independent generations could all affect scores. None of those details appears in the supplied arXiv abstract, so they remain meaningful unknowns.

The second question is statistical . The headline examples involve a 0.042-point shortfall, a percentile shift from 90th to 77th, and a reported score range of 0.869 to 0.932 at a stated accuracy level. To assess their stability, readers would need item-level results, uncertainty estimates, sensitivity analyses, and information about the human comparison cohort. It would also be useful to know whether the ranking changes persist across subjects and exam years, or whether they are concentrated in particular question types. The abstract reports results across eight models and a broad , but it does not establish the variance of those results or the extent to which individual exams drive the aggregate gap.

The broader test is whether the partial-credit gap generalizes. Other educational systems may use penalties, thresholds, multi-part questions, or rubrics that reward different forms of partial knowledge. Applying the same analysis to those settings could reveal whether this is a Vietnam-specific feature or a wider weakness in exam-based language-model evaluation. Independent teams should also compare official-rubric scores with other measures of educational performance rather than assuming that either accuracy or the convex score is a complete measure of competence. Finally, the arXiv record identifies the work as submitted on 18 August 2026 and presents it as a preprint. The supplied source gives no peer-review outcome, external replication, deployment, or evidence that exam authorities have adopted the .

관련 가이드 및 퀴즈

AI 모델 설명AI 윤리AI 트레이닝알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.
이것이 유용하다고 생각하시나요?