What happened
The Korea Herald reported that South Korea’s Ministry of Science and ICT released detailed second-phase scores for its Sovereign AI Foundation Model Project. SK Telecom ranked first overall with 70.6 points, followed by Upstage with 69.9, LG AI Research Institute with 69 and Motif Technologies with 65.8.
The Korea Herald reported that South Korea’s Ministry of Science and ICT released detailed scores from the second phase of its Sovereign AI Foundation Model Project on Aug. 27. The report said the ministry released the results after Motif Technologies, which was eliminated in the second phase, requested them and the other shortlisted teams consented. The source described the project as an effort to develop domestic foundation-model capabilities and expand South Korea’s AI ecosystem. The detailed release set out the project’s overall results and the circumstances in which those scores became public after the request and consent.
According to The Korea Herald, SK Telecom received the highest overall score, 70.6 points out of 100. Upstage scored 69.9, LG AI Research Institute scored 69, and Motif Technologies scored 65.8. The evaluation combined benchmark assessments worth 40 points, an expert evaluation worth 35 points and user evaluations worth 25 points. The user component included professional users and members of the public. These figures describe the combined assessment used for the overall ranking, rather than a result from one evaluation category alone.
The report said the benchmark portion drew 25 points from the Artificial Analysis Intelligence Index and 15 points from a National Information Society Agency benchmark. Motif Technologies ranked first in the AAII portion with 11.9 points, ahead of Upstage at 9.4, SK Telecom at 8.8 and LG AI Research Institute at 7.8. In the NIA benchmark, which The Korea Herald said covered seven areas including Korean-language capabilities and reliability, SK Telecom scored 13.4 and Upstage 13.3, followed by LG AI Research Institute at 12.8 and Motif Technologies at 12.7. The component figures show how the benchmark portion was divided between the two reported sources.
Read the source: koreaherald.com ↗
Why it matters
The results provide a public view of how South Korea is comparing domestic AI efforts across global benchmarks, Korean-language capability, reliability, expert judgment and real-world usability. They also show that overall rankings varied substantially by evaluation method.
The results matter because they make a government-backed comparison of South Korean AI projects more specific than a simple announcement of selected companies. The Korea Herald’s account shows that the leading team overall was not first in every component. Motif led the global benchmark category, LG AI Research Institute led the expert evaluation and public evaluation, while SK Telecom led the professional-user evaluation. The overall result therefore reflected a combined assessment rather than a single benchmark result.
That spread illustrates how evaluation design can shape conclusions about AI capability. A model’s performance on an international benchmark, its Korean-language and reliability results, expert assessments and practical usability may measure different qualities. The Korea Herald reported that 10 external experts from industry, academia and research institutions participated in the expert evaluation, while 49 AI startup executives took part in the professional-user evaluation and 185 members of the public participated in the public evaluation. The different categories and participant groups help explain why the component results did not produce one identical ranking.
For South Korea’s AI strategy, the scores offer a basis for deciding which domestic systems receive continued attention under the sovereign-AI program. The ministry told The Korea Herald that the project was not intended simply to rank teams or eliminate participants, but to encourage qualitative growth and expansion of the country’s AI ecosystem. That stated goal suggests the scores may be used as feedback for model development and national capability-building, not only as a competition table. The report does not establish whether the evaluated systems are broadly available, how they compare with leading international models, or whether the scores predict performance in operational deployments. Those unanswered questions limit what can be concluded from the rankings alone.
What to watch next
The ministry said it will review Motif Technologies’ objection and make a final decision within 15 days, according to The Korea Herald. The next significant developments are the outcome of that review, any changes to the project’s participating teams and future six-month revisions to its evaluation targets.
The most immediate development is the ministry’s response to Motif Technologies’ objection. The Korea Herald reported that Motif had previously challenged the second-phase outcome and that the ministry would review the objection under established procedures, with a final decision expected within 15 days. The source does not say what remedy Motif is seeking or whether the objection could change the rankings or project participants. The review is therefore the next point at which the reported second-phase outcome could receive an official update.
The ministry also said the project’s development targets will be revised every six months as “moving targets,” reflecting the rapid pace of AI progress, according to The Korea Herald. Future revisions could therefore change the balance among benchmark performance, expert review and user experience. The report said the evaluation criteria and methodology were finalized after prior consultation with the shortlisted teams, but it does not provide the full test questions, model versions, training data, scoring rules or individual evaluator comments. Those omissions leave important details for later reporting or future revisions.
Readers should watch for evidence about whether these evaluations translate into practical public or commercial use. The Korea Herald’s report provides scores and participation counts, but it does not independently verify the ministry’s underlying data or assess the models itself. It also does not identify deployment commitments, costs, safety testing, compute resources, licensing terms or access conditions. Those details will be important for judging whether the project represents durable sovereign AI capacity or primarily a government evaluation exercise. The scores establish a reported comparison, while practical use would require evidence beyond the evaluation results.

