What happened
Researchers introduced THPT-Ladder, a benchmark built from 632 items in 21 official Vietnamese exams across 11 subjects. It applies the 2025 National High School Graduation Examination’s convex scoring rubric, under which four true/false statements earn 0, 0.10, 0.25, 0.50, or 1.00 points rather than proportional credit. Across eight models, the paper reports lower official-rubric scores than proportional scoring and shows that rankings against human candidates can change substantially.
The arXiv record describes a problem with translating exam performance into a single accuracy number. In Part II of Vietnam’s 2025 National High School Graduation Examination, each question contains four true/false statements. The official rubric assigns 0 points when none are correct, then 0.10, 0.25, 0.50, or 1.00 points as the number of correct statements rises from one to four. That schedule is convex rather than additive: three correct statements receive 0.50 points, even though three out of four statements would correspond to 0.75 under proportional credit. The paper treats this difference as an evaluation issue, not merely a grading detail, because Part II accounts for 4.00 of the exam’s 10.00 points. The source establishes these rules as the paper’s description of the 2025 reform; it does not independently verify the examination policy outside the cited paper record.
The authors say they created THPT-Ladder with 632 items drawn from 21 official exams covering 11 subjects. The benchmark is reported to grade answers exactly as the ministry grades its students. The paper also says the ministry publishes marks for more than one million candidates, allowing the researchers to place model results within a human candidate cohort. Across eight models, the official rubric reportedly produces scores between 0.020 and 0.159 points lower per Part II question than proportional credit. The abstract does not list all eight models, explain how each model was prompted, identify the number of runs, or describe whether answers were generated under identical settings. Those details are important for interpreting the numerical comparisons and are not supplied by the authoritative text provided here.
The abstract gives two examples of the reported ranking effect. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall is said to move the model from the 90th to the 77th percentile among 481,293 candidates. The paper also reports that a model at Claude Sonnet 5’s accuracy level could receive between 0.869 and 0.932 points per question, depending on how its errors are distributed. In the authors’ account, accuracy alone does not predict the penalty because the official score depends on which statements are answered correctly together within each question. The source does not provide the underlying response files, confidence intervals, item-level breakdowns, or an independent audit of those calculations, so these findings should be treated as reported preprint results rather than settled measurements.
Read the primary source: arxiv.org ↗
Why it matters
The paper argues that ordinary accuracy metrics can describe a model as more capable than an institution’s actual grading rules would certify. That matters for AI evaluations built from human exams, especially when benchmark results are used to compare models or make claims about educational competence. The findings are limited to the reported benchmark and models; the supplied source does not establish how broadly the result generalizes beyond Vietnam’s exam format.
The central implication is methodological. A benchmark can preserve the appearance of rigor while measuring a different construct from the one used by the institution whose exam it borrows. Under proportional scoring, each additional correct statement is treated as contributing the same amount. Under the Vietnamese scheme described by the paper, partial knowledge is rewarded unevenly, and some combinations of correct and incorrect statements receive less credit than their raw accuracy would suggest. A model’s final score therefore depends not only on how many statements it gets right, but also on the pattern in which those statements occur. For readers comparing models, that means a leaderboard based only on aggregate accuracy may not answer the practical question of how a model would perform under the official certification rule.
The reported human-cohort comparison makes the issue more concrete. The paper says Qwen3.5-27B’s 0.042-point difference on the 2025 History exam changes its position from the 90th to the 77th percentile among 481,293 candidates. If reproduced, that is not a cosmetic change in presentation: it alters the population comparison a reader would draw from the same model responses. The result does not show that the model became less capable, nor does it establish that percentile placement predicts classroom or examination outcomes. It shows that the choice of scoring rule changes the measurement.
The distinction is especially relevant when exam benchmarks are used to communicate claims about reasoning, knowledge, or readiness for educational tasks. The practical public impact is mainly on how AI performance is reported and interpreted. Researchers, educators, policymakers, and journalists may rely on benchmark scores without seeing how a grading formula handles partial answers. The paper’s argument supports reporting the official score alongside accuracy when the benchmark is intended to represent a real examination. It also suggests that evaluation designers should disclose whether a benchmark’s scoring rule is additive, proportional, thresholded, or otherwise non-linear. At the same time, the supplied source does not show that THPT-Ladder is representative of all language-model testing, that Vietnam’s scheme is common internationally, or that a convex rubric is inherently superior. The contribution is a warning about construct validity and comparability, not evidence that one universal scoring method should replace another.
What to watch next
The key follow-up is whether the full study documents its prompts, model versions, sampling conditions, scoring implementation, statistical uncertainty, and data availability. Researchers should test whether the same gap appears under other non-additive grading systems and on independently reproduced runs. The paper is an arXiv preprint, so the supplied source does not establish peer-review status or independent replication.
The first verification question is reproducibility. The full paper should clarify the exact 632-item composition, the 21 exams and 11 subjects included, answer normalization, handling of abstentions or ambiguous outputs, and the implementation of the ministry’s rubric. It should also identify all eight models and their precise versions, because model names alone may not determine behavior. Prompt wording, system instructions, decoding settings, date of access, and the number of independent generations could all affect scores. None of those details appears in the supplied arXiv abstract, so they remain meaningful unknowns.
The second question is statistical robustness. The headline examples involve a 0.042-point shortfall, a percentile shift from 90th to 77th, and a reported score range of 0.869 to 0.932 at a stated accuracy level. To assess their stability, readers would need item-level results, uncertainty estimates, sensitivity analyses, and information about the human comparison cohort. It would also be useful to know whether the ranking changes persist across subjects and exam years, or whether they are concentrated in particular question types. The abstract reports results across eight models and a broad benchmark, but it does not establish the variance of those results or the extent to which individual exams drive the aggregate gap.
The broader test is whether the partial-credit gap generalizes. Other educational systems may use penalties, thresholds, multi-part questions, or rubrics that reward different forms of partial knowledge. Applying the same analysis to those settings could reveal whether this is a Vietnam-specific feature or a wider weakness in exam-based language-model evaluation. Independent teams should also compare official-rubric scores with other measures of educational performance rather than assuming that either accuracy or the convex score is a complete measure of competence. Finally, the arXiv record identifies the work as submitted on 18 August 2026 and presents it as a preprint. The supplied source gives no peer-review outcome, external replication, deployment, or evidence that exam authorities have adopted the benchmark.


