Thursday, August 27, 2026

Motif, despite ranking first in AAII, was eliminated; "The evaluation results should be disclosed and reviewed again"

Input
2026-08-27 11:12:28
Updated
2026-08-27 11:12:28
Bae Kyung-hoon, Deputy Prime Minister and Minister of Science and ICT, delivers congratulatory remarks at the first presentation event for the independent AI foundation model project held on Dec. 30 last year at COEX in Gangnam District, Seoul. Newsis News Agency

[Financial News] Motif Technologies, which was eliminated in the second evaluation of the Independent AI Foundation Model Project (K-LLM), has changed its earlier position that it would not file an objection and has submitted a formal complaint to the Ministry of Science and ICT.
Motif is asking the ministry to publicly verify the detailed evaluation results, after MSIT explained that the company was eliminated because its expert and user scores were relatively low. Motif, however, drew a clear line, saying, "This objection is not intended to demand reselection for the K-LLM project."
Eliminated despite topping AAII: "The detailed evaluation results should be disclosed"

Motif said in a statement on the 27th that it had submitted an objection to MSIT over the second-round K-LLM results.
In its objection, Motif requested that the ministry disclose the standards used for expert and user evaluations, along with the item-by-item assessment details, and review the results again.
The company had initially said it would not challenge the results announced on the 18th, but reversed course 10 days later.
Motif said it "respected and accepted the evaluation results themselves," but explained that it filed the objection because MSIT later said Motif 3 had only scored highly in AAII and had issues in expert and user evaluations, while also presenting a new evaluation direction that would reduce the weight of benchmarks in the next round.
The company appears to have raised concerns that, by publicly saying Motif was strong only in benchmark scores but weak in real-world expert and user assessments, MSIT could create the impression that the model has usability problems in future market competition.
The second-round K-LLM evaluation consisted of 40 points for benchmark scores, including 25 points for AAII and 15 points for NTIA, 35 points for expert evaluations, and 25 points for user evaluations. Motif recorded the highest score among participating models in the Artificial Analysis Intelligence Index (AAII), a comprehensive performance metric for global AI models, with 47 points. Upstage scored 37 points, SK Telecom 35 points, and LG AI Research 31 points. According to Motif, Motif 3 ranked 10th globally among model developers and first outside the U.S. and China.
However, Motif finished last in the overall score combining expert and user evaluations, failing to advance to the third round.
Motif argued that, since AAII is a comprehensive indicator that covers both performance and usability, the specific reasons for the discrepancy between benchmark results and the overall assessment should be disclosed.
"Foundation models should be evaluated on scalable base performance"

The company also argued that the performance gap shown in the benchmark results was not sufficiently reflected in the final scoring. The raw score gap between Motif, which ranked first in AAII, and LG AI Research, which ranked fourth, was 16 points, but that difference shrinks to 4 points when converted to a 25-point scale, it said.
Motif emphasized that "the value of a foundation model comes not only from the usability of a chatbot, but from its base performance, which can be expanded into multiple fields," adding that "rather than lowering the benchmark weight, the actual performance gap should be properly reflected in the scoring."
It also raised the possibility that user evaluations may have been influenced by the developer's brand recognition. Motif said, "When ordinary users evaluate models from large companies and startups, knowing the developer in advance can create bias, so we requested a blind evaluation," adding, "Since the performance gap in the benchmark changed significantly after the user evaluation, the evaluation method and conditions must provide convincing grounds for how they affected the results."
Despite the objection, it will not take part in the third round

Motif made clear that this objection was not intended to overturn the elimination result and rejoin the third-round evaluation.
Motif said, "Regardless of the outcome of the objection, we will not participate in the third round of the K-LLM project," explaining that the purpose was "to clarify the kind of foundation model and values the project seeks through open discussion of the evaluation process."
MSIT plans to select the final two teams early next year from among the three teams that advanced to the third round: LG AI Research, SK Telecom, and Upstage. The ministry is expected to review issues raised in the second-round evaluation, including benchmark reliability and score adjustments, before setting the criteria for the third round.
Ryu Jemyoung, the second vice minister of MSIT, said when the second-round results were announced on the 18th, "We will use the objection process to see whether there are any additional areas that need to be improved in the first and second evaluations," adding, "We will consult with the three companies and announce the third-round evaluation plan as soon as possible."

[email protected] Choi Hye-rim Reporter