Who Owns the Korean-Language Model?
The sovereign AI debate is consumed as though building a single homegrown model settles the matter. What actually demands scrutiny is which of the three—data, compute, or evaluation—the state truly controls. Subsidies should flow to scorecards, not models.
AI Summary
Korea's sovereign AI debate has narrowed to a question of GPU counts and parameter sizes, but this framing mistakes a model—a fast-depreciating output—for a durable national asset. The article argues that state resources should instead flow to Korean-language evaluation benchmarks, open training data, and shared compute pools, because the market will not supply these public goods on its own. Until Korea establishes a dedicated institution to define and continuously audit how well AI performs in Korean across high-risk domains such as healthcare and public administration, linguistic sovereignty will remain an illusion.
The Illusion That Owning One Model Equals Sovereignty
As the term 'sovereign AI' has hardened into a political slogan, the picture has grown simple: if the state nurtures one or two large models that handle Korean well, linguistic sovereignty is secured. So budget meetings always stall at the same point—how many GPUs to buy, how many parameters to scale up to.
This picture is only half right. A model is an output, not an asset. A bigger model arrives in six months, and the weights produced with last year's subsidy go stale fast. If what the state bought with tax money becomes obsolete within a year, that is a consumable, not sovereignty. The moment we define linguistic sovereignty as model ownership, we step onto a treadmill where we pay the same cost all over again every year.
The question needs to change. Do we actually understand this technology as a society? Before AI is a productivity tool for individuals, it is public infrastructure that decides what a society knows—and does not know—in its own language. If that is the case, then the unit for allocating resources should not be the 'model' but the three pillars supporting that infrastructure: data, compute, and evaluation.
Of Data, Compute, and Evaluation: Which One Should the State Control?
Think of it as cooking. Training a model is the cooking. Compute (GPU) is the fire, data is the ingredients, and the evaluation benchmark is a scorecard recording the diners' tastes. Nearly all of Korea's current debate is fixated on the size of the flame—arguing over how many more burners to add.
But even with the same fire, inferior ingredients yield different food; without a scorecard, there is no way to ever determine who cooked better. Compute is the domain where the market works best—it can be rented from the cloud when there is money. Korean-language public data and Korean-language evaluation standards, by contrast, are not something the market will produce on its own. There is no money in it. The exact point where the market fails is where the state must step in.
The core point is this: a model can be borrowed, but a scorecard cannot. How often does an AI conducting medical consultations in Korean hallucinate? What does it omit when summarizing administrative documents? How does it handle regional dialects and the speech patterns of elderly users? A yardstick for measuring these things does not emerge from translating English-language benchmarks. We must define it ourselves—and it is an asset that will outlast any single homegrown model by far.
Why Subsidies Should Flow to Scorecards, Not Weights
A counterargument follows. Evaluation alone does not grow an industry; ultimately, you have to run large models yourself for capabilities to accumulate. This has merit. A country that has never trained a foundation model end to end does not understand the floor of that technology.
But capability and ownership are different things. Training capability stays in people and code; model weights stay on disk. If subsidies buy weights, that money disappears in a year. Spend the same funds on Korean-language evaluation benchmarks, open training data, and a shared compute pool, and companies, universities, and startups will each train dozens of models on top of that foundation. Instead of buying one, you are building the soil for a hundred to grow.
There is one more common trap here. When benchmark scores become the goal itself, everyone builds models that excel only on that test. The moment a metric becomes a target, the metric breaks down. That is why evaluation must not be a fixed exam created once and left alone, but a living institution that the public continuously updates and audits. This cannot be done by the private sector. The exam-setter cannot also be the exam-taker.
The Blanks Korea Has Yet to Fill
The problem is that the institutional actor responsible for creating and managing this scorecard simply does not exist in Korea. The Ministry of Science and ICT handles model-training subsidies; the Personal Information Protection Commission oversees the boundaries of data use. Yet there is effectively no agency that continuously measures whether Korean-language AI is fit for public use, makes those standards transparent, and keeps them updated. Compare this to the United Kingdom, which established an AI Safety Institute and elevated model evaluation itself to a state function—Korea still leaves evaluation to the voluntary self-promotion of companies.
The blanks are concrete. In high-risk domains such as healthcare, public administration, and education, who sets the minimum standards a Korean-language model must meet? Under what conditions will publicly held legal precedents, administrative documents, and health data be opened for training? Is there a channel for citizens to report malfunctions of AI in their own language, with those reports feeding into future evaluations? Who verifies whether these models are actually usable on the ground in non-capital cities like Busan? Right now, every one of these questions goes unanswered.
The minimum threshold of civic AI literacy is also decided here. In the AI era, what citizens need is not the ability to write better prompts, but the instinct to ask: 'What was this model evaluated against?' That question can only exist if the scorecard is publicly available.
Many countries will become proficient at using AI. Bigger models and faster compute ultimately converge on a question of money. But countries that define for themselves what to measure in their own language are rare. Linguistic sovereignty goes not to whoever has the largest model, but to whoever holds the yardstick to score it. Being a country that genuinely understands AI comes before being a country that merely uses it well.
This article was automatically translated from the Korean original by AI. For the authoritative version, read it in Korean.
한국어 원문 읽기 →