Government's $1.6 Billion AI Data Project Faces Criticism for Waste and Errors

By Kim Bongcheol Posted : August 12, 2026, 15:44 Updated : August 12, 2026, 15:44

The South Korean government has invested 1.63 trillion won ($1.6 billion) in an artificial intelligence (AI) training data project, but numerous instances of duplicated and unusable data have been identified. Critics argue that inadequate pre-review, quality control, and post-management have led to inefficient spending of taxpayer money.


According to an audit report released by the Board of Audit and Inspection on August 12, the Ministry of Science and ICT has been developing AI training data from 2017 to 2024. As of November 2025, 908 types of data have been made available on the AI Hub.


In addition, 26 public institutions, including the Ministry of Food and Drug Safety, Seoul City, and the Korea Expressway Corporation, have released 313 types of data through their own portals. However, the lack of a system for sharing construction plans or assessing the usability of existing data has resulted in the duplication of similar datasets.


The Ministry of Science and ICT initially spent 7.38 billion won to develop four types of data related to wildlife, household waste, concrete cracks, and eggs. Yet, five other institutions, including Seoul City and the Korea Expressway Corporation, spent an additional 610 million won to create similar datasets.


In Daejeon, data on household waste showed that 5,200 out of 9,000 images (57.8%) were similar to existing data from the Ministry of Science and ICT. Similarly, about 25,000 out of 60,000 images in the wildlife dataset from the Korea Expressway Corporation were found to be similar to previously collected data.


There were also instances where similar data was created by different agencies in the same year without coordination. While the Ministry of Science and ICT invested 23.3 billion won in data related to pills, oral health, and autonomous driving from 2021 to 2023, three other agencies, including the Ministry of Food and Drug Safety, separately spent 840 million won on similar datasets.


In 2023, the Personal Information Protection Commission developed 1,000 images for oral health, all of which were similar to data created by the Ministry of Science and ICT that same year. Additionally, 1,411 images for pill identification created by both the Ministry of Science and ICT and the Ministry of Food and Drug Safety overlapped.


Quality control of the data was also found to be lacking. Among the top 20 public institutions in terms of construction performance, four had no separate quality control standards, and nine did not conduct third-party quality verification before delivery. The Ministry of Food and Drug Safety and Korea East-West Power did not require third-party verification or internal checks by the executing agencies.


Of the 114 types of data constructed by five institutions, 96 types (84.2%) lacked a syntax rule specification to assess the accuracy of sentence structure and format, making quality checks impossible.


When the Board of Audit and Inspection conducted a sample verification of six datasets constructed by Gyeonggi Province and the Korea Expressway Corporation, one dataset was found to be untestable. The remaining five datasets had format accuracy ranging from 77.3% to 99.2%, falling short of the standard of 99.5%.


The 'Road Driving CCTV Training Dataset' constructed by the Korea Expressway Corporation in 2025 lacked bounding box coordinate information for 2,680 images depicting objects like cars, rendering it unusable for training autonomous driving technology.


In the 'Large Waste Learning Data' created by Seoul City in 2020, 538 out of 2,303 images (23.4%) had mismatched file names and actual images. For instance, files labeled 'bag' contained images of boxes or tables instead.


Issues also arose when construction companies went out of business, leaving data errors uncorrected. The Korea Intelligent Information Society Agency constructed 721 types of data through 1,270 companies from 2021 to 2024, with 213 types (29.5%) coming from 65 companies that either closed or faced internal issues.


Despite receiving complaints about errors in 11 types of data, the agency closed the complaints without making corrections or engaging third parties, citing the closure of the construction companies. Data with confirmed errors continued to be publicly available on the AI Hub.


The Board of Audit and Inspection has instructed the Ministry of Science and ICT to collaborate with the Ministry of the Interior and Safety to review the potential for alternative use of existing data and to establish a system for sharing and coordinating construction plans among agencies. It also called for the establishment of common quality control standards and procedures for internal checks and third-party verification applicable to public institutions.


The Korea Intelligent Information Society Agency has been directed to directly correct errors in data constructed by closed companies or to establish a third-party maintenance system. The Ministry of Science and ICT stated that it would verify the quality of existing data through a comprehensive investigation across ministries and address deficiencies with the help of data and AI specialized companies.


The Board of Audit and Inspection noted, “The production of unusable data due to duplicated datasets and inadequate quality control systems has occurred, and post-management has been insufficient. The ongoing complaints from AI companies about the lack of usable data highlight the urgent need for improvements in both the quantity and quality of data to achieve the goal of becoming a leading AI nation.”





* This article has been translated by AI.

Copyright ⓒ Aju Press All rights reserved.