Five K-LLM teams release AI training data for free, opening 1.56 trillion tokens
- Input
- 2026-08-27 14:00:00
- Updated
- 2026-08-27 14:00:00

[Financial News] The Ministry of Science and ICT said it will make the AI training data of five elite teams that took part in the first evaluation of the Independent AI Foundation Model Project (K-LLM) available to the public for free. The 29 data sets, built by NAVER Cloud, Upstage, SK Telecom, NC AI and LG AI Research, amount to 1.56 trillion tokens, enough to train large AI models with around 70 billion to 80 billion parameters.
The Ministry of Science and ICT announced on the 27th that it will release the 29 data sets secured by the five elite K-LLM teams during the first-stage evaluation on the AI Hub website.
The newly opened data were planned by each team to match its model development strategy and were secured using the government's budget for data construction and processing. They include large-scale datasets needed for pretraining large AI models, multimodal data such as video and audio, and Red Teaming data used to test model safety.
NAVER Cloud built 2.34 million public videos and 520,000 broadcast videos, along with 15.5 million text records and 12 million voice question-and-answer pairs based on those videos. The text data consist of captions and other content that describe video scenes sentence by sentence, and can be used to improve AI's video understanding and train image generation models.
Upstage is releasing 1 trillion tokens of pretraining data and 500,000 post-training data records. The pretraining data can be used for the 'from scratch' approach, in which an AI model is trained anew from the ground up. The post-training data focus on developing AI agents with reasoning, judgment and execution capabilities beyond simple question-and-answer tasks.
SK Telecom built high-difficulty, multi-step reasoning data in specialized fields such as mathematics, science and law, as well as post-training data based on voice and images. It also included about 10,000 Korean-style Red Teaming data records that reflect domestic laws and the common values of Korean society.
NC AI prepared seven types of training data based on real-world industrial data, including technical documents from manufacturing and voice recordings from customer service consultations. The materials are designed to enhance large language model (LLM) capabilities such as contextual understanding, step-by-step reasoning and question answering.
LG AI Research built a multimodal Physical AI dataset tailored to Korean home environments. It filmed more than 50 household tasks in 50 homes across the country, securing over 170,000 video clips and layering object, segmentation, posture and contextual information.
The National Information Society Agency (NIA) and the Telecommunications Technology Association (TTA) received the data from each team after the first-stage evaluation and verified its quality, as well as whether it contained personal information or harmful content. Data subject to licensing restrictions were excluded from public release.
Under the conditions for participating in the K-LLM project, at least 50% of the data secured with government funds must be opened to the public. NAVER Cloud, Upstage, SK Telecom and NC AI decided to release all of the data that passed quality checks. LG AI Research selectively released data in a way that meets the mandatory disclosure requirement of at least 50%.
Anyone in South Korea, including domestic companies, researchers and students, can download and use the released data for free from AI Hub. The Ministry of Science and ICT plans to open additional government-funded data after quality checks during the second-stage evaluation.
The released data are valuable assets that contain the AI training strategies of the elite teams, and they will serve as a foundation for the growth and self-sufficiency of the AI ecosystem, said Kim Kyung-man, director general of the AI Policy Bureau at the Ministry of Science and ICT.
[email protected] Choi Hye-rim Reporter