Torna alle notizie
InnovazioneAI Understanding briefing

Il documento PhysElite riporta che il principale LLM multimodale ha risolto il 33,7% dei problemi di fisica a livello delle Olimpiadi

Un nuovo benchmark di 11.586 problemi di fisica bilingue basati su diagrammi riporta che il più potente dei 18 modelli linguistici multimodali testati ha raggiunto solo il 33,7% di precisione di risposta.

5 min readRead the primary source
Primary-source image accompanying PhysElite paper reports that leading multimodal LLM solved 33.7% of Olympiad-level physics problems
Documento di origine primariaFonte registrata
Editore
arxiv.org
Collegamento alla fonte
arxiv.orghttps://arxiv.org/abs/2608.25097
Tipo di fonte
Documento principale: un annuncio ufficiale, un documento, un documento o una pagina proprietaria che leggiamo direttamente.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Modello linguistico di grandi dimensioni (LLM)
Un modello linguistico addestrato su enormi corpora di testo per generare e analizzare testo.
Punto di riferimento
Un test o un set di dati standardizzato utilizzato per misurare e confrontare le prestazioni del modello.
Set di dati
Una raccolta di esempi strutturati o non strutturati utilizzati per la formazione, la convalida o il test.
Mettiti alla provaQuiz sulla spiegazione dei modelli di intelligenza artificiale

Cosa è successo

Researchers introduced PhysElite, a bilingual multimodal designed to test large language models on Olympiad-level physics problems. The contains 11,586 problems, visual diagrams, Chinese-English solution derivations, and final answers. The paper reports results from evaluations of 18 open-source and closed-source multimodal language models, with the strongest model reaching 33.7% answer accuracy.

The paper presents PhysElite as a for Olympiad-level physics reasoning by multimodal large language models. It was submitted to arXiv on August 25, 2026. The authors argue that existing physics benchmarks are limited because they do not contain enough high-difficulty problems and do not comprehensively cover visual information, physics knowledge areas, and step-by-step solution processes.

According to the paper’s abstract, PhysElite contains 11,586 Olympiad-tier problems in both Chinese and English. Each problem is paired with a visual diagram, a bilingual derivation broken into steps, and a final answer. The source says the has been released, although the supplied text does not include the destination link or describe its license, format, or access requirements.

The authors evaluated 18 multimodal language models, including both open-source and closed-source systems. The strongest model reached 33.7% answer accuracy. The researchers also conducted step-level process evaluation to identify where models failed within a reasoning chain. The source does not provide the names of the models, the number of attempts per problem, the scoring rules, or the detailed results by problem type.

Dettagli della fonte: arxiv.org ↗

Perché è importante

The result suggests that strong performance on existing physics or general reasoning tests may not translate to reliable solutions for difficult, diagram-dependent, multi-step physics problems. PhysElite is intended to make those weaknesses more visible by evaluating both final answers and individual stages of a model’s reasoning process.

The addresses a practical weakness in how AI reasoning systems are often assessed. A model can produce fluent explanations or perform well on shorter questions while still failing when it must interpret a diagram, select relevant physical principles, carry out several linked deductions, and arrive at an exact answer. PhysElite is designed around that combined challenge rather than treating physics as text-only question answering.

The reported 33.7% result is consequential because it places a clear limit on what the tested systems demonstrated on this particular high-difficulty task. It does not show that multimodal language models are incapable of solving Olympiad physics, and it does not establish that every model performs at the same level. It does show, according to the paper, that even the strongest evaluated system did not solve most problems correctly.

The step-level evaluation could be more useful than a single final score for researchers and users. If failures occur mainly while interpreting diagrams, selecting equations, performing calculations, or connecting intermediate results, developers may be able to target those weaknesses more precisely. The supplied source does not state which stages produced the most errors, so the ’s diagnostic value cannot yet be assessed in detail.

For the public, the findings offer a caution against treating general-purpose AI systems as dependable substitutes for expert physics reasoning. A system that gives a plausible but incorrect multi-step solution may be difficult for a nonexpert to check. The ’s bilingual and visual design may also help expose limitations that are hidden by evaluations focused on English text or short final-answer formats.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verifica concettuale interattiva+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Cosa guardare dopo

The main questions are whether the and evaluation code are sufficiently accessible for independent replication, how performance varies across physics topics and visual formats, and whether future models improve on the without relying on memorized solutions. The source does not identify the tested models or provide the evaluation breakdown, so the reported 33.7% should not be used to rank particular systems.

Independent replication is the immediate test. Researchers will need to inspect the released problems, diagrams, solutions, and scoring procedures to determine whether the is free of ambiguities, duplicated items, or other factors that could affect results. The source confirms a release but does not provide enough information to evaluate its accessibility or reproducibility.

Future coverage should show how the 33.7% figure breaks down across physics subjects, difficulty levels, diagram types, languages, and answer formats. It is also important to know whether models receive the same prompts, whether tools such as calculators are allowed, and how partial credit is handled. None of those details appears in the supplied source.

The may become more informative if later evaluations compare model versions over time and report both final-answer accuracy and step-level reliability. Researchers should also examine whether models can recognize uncertainty, identify incorrect intermediate steps, or request clarification when a diagram is ambiguous. The paper’s abstract does not report those behaviors.

The most important unknown is how representative PhysElite is of real physics work. Olympiad problems are intentionally difficult and may not reflect classroom exercises, laboratory analysis, engineering design, or research practice. The source establishes the ’s scope and reported result, but it does not establish how performance on PhysElite predicts performance in those other settings.

Taken together, the available description supports a limited interpretation of the result. The 33.7% figure belongs to the reported answer-accuracy evaluation of the 18 tested multimodal language models on this . It should be read alongside the ’s bilingual problems, visual diagrams, derivations, and final answers, as well as the separate step-level process evaluation. Because the supplied source does not include model names, scoring rules, attempt counts, topic breakdowns, or the evaluation code, readers cannot use the abstract alone to reconstruct the comparison or determine which parts of the task drove the result. Those missing details are therefore central to understanding subsequent reports about PhysElite.

Guide e quiz correlati

Spiegazione dei modelli di intelligenza artificialeTrasformatoriFormazione sull'intelligenza artificialeEtica dell'IAMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker del rilascio del modello AI
Lo hai trovato utile?