The second phase evaluation of the 'Reader AI Foundation Model' project, led by the Ministry of Science and ICT, has concluded, revealing the global benchmark results for domestic AI models. Notably, Motif Technologies' 'Motif3' achieved the highest score among domestic models, drawing significant attention to the evaluation results.
According to the Artificial Analysis (AA) platform's 'Intelligence Index (AAII)' released on August 13, Motif's 'Motif3' scored 47 points, placing it among the top 10 global large language models (LLMs) and marking the highest ranking for a model developed by a domestic company.
Following Motif3, Upstage's 'Solar Open2' scored 37 points, SK Telecom's (SKT) 'A.X K2' received 35 points, and LG AI Research's 'K-Exa One 2.0' scored 31 points. The gap between the top and bottom scores among the four models participating in the Reader AI project is 16 points.
AA is a platform that comprehensively analyzes AI model performance, quality, and cost. It provides metrics that consider not only model performance but also price, output speed, latency, and context window, which are crucial in the actual service implementation process.
Notably, the AAII combines benchmark results from various fields, such as mathematics, science, and coding, to evaluate the overall capabilities of AI models. This approach is distinctive as it aggregates multiple evaluation results rather than relying on individual benchmarks. However, it does not include a separate metric for assessing Korean language proficiency, which limits its ability to fully reflect model performance in the domestic usage environment.
Another noteworthy aspect is that the size of the models did not correlate with their AAII scores. Among the models participating in the second phase evaluation, the largest model is LG AI Research's K-Exa One 2.0, developed with a total of 750 billion parameters. SKT's A.X K2 follows with 688 billion parameters. In contrast, the highest-scoring Motif3 has a total of 314 billion parameters, while Upstage's Solar Open2 consists of 250 billion parameters.
However, the AAII scores do not directly translate to the results of the second phase evaluation of the Reader AI project. The second phase evaluation consists of 40 points from benchmarks, 35 points from experts, and 25 points from users. Notably, this evaluation included a public assessment panel of 200 citizens who directly used the four models and evaluated their performance.
The gap with the top global models remains significant. In the AAII, Anthropic's 'Claude Opus5' scored 63 points, 16 points higher than Motif3. OpenAI's 'GPT-5.6 Sol' and xAI's 'Grok 4.6' also scored 61 points each.
Chinese AI models have also shown remarkable progress. Moonshot AI's 'Kimi K3' scored 60 points, just one point behind GPT-5.6 Sol. Alibaba's 'Q1.8 Max' scored 58 points, Z.ai's 'GLM-5.2' scored 53 points, and DeepSeek's latest model 'V4 Flash' scored 52 points, all exceeding Motif3's score by 5 to 13 points.
The Ministry of Science and ICT is expected to announce the results of the second phase evaluation of the Reader AI project soon. Through this evaluation, three of the four elite teams currently participating will advance to the next stage.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.