SOUTH+BRIDGE
Culture AI-translated

Korean Territory Within the Weights

The place Korean occupies within global models is not a matter of expression but of territory. The state has yet to even measure what exactly we lose when token share decreases.

Transcript · June 6, 2026 · 5 min read

AI Summary

Large language models allocate disproportionately small space to Korean compared to English—typically around 1% of training data—resulting in higher costs per query, shallower reasoning, and less contextual depth for Korean users. While Korean companies are developing domestic models, South Korea lacks national frameworks to measure, track, or strategically expand Korean's 'territorial' presence within global AI systems. The author argues this is not a translation issue but a matter of sovereignty that requires public monitoring of Korean token share, inference accuracy, and usage costs across major models.

Korean Territory Within the Weights

What You Can't See Through a Translation Lens

People tend to think of large language models as giant translation machines. When you ask in Korean and get a Korean answer, they imagine there's a separate Korean room and English room inside the model. Wrong. There are no rooms inside the model. Words and sentences are scattered and embedded across a single mass of coordinate space consisting of billions of numbers—weights. Which language occupies more space in that territory depends on how much and how diversely that language was included in the training data.

This is where the first misconception begins. Just because a model answers smoothly in Korean doesn't mean it understands Korean as deeply as English. Whether GPT series, Gemini, or Llama, analyses of publicly available training corpora show English is overwhelming while Korean typically remains a small fragment of around 1 percent. The model fills gaps in Korean knowledge with a worldview learned from English. So when you ask about Korean legal systems or local issues in Busan, the sentences are fluent but the content often turns out to be answers that dress up an American newsroom's common sense in Korean.

Tokens as a Unit, Territory as a Metaphor

If you view language only as a 'means of expression,' this gap seems trivial. Translation works, after all. But inside the model, language is not expression—it's real estate. Imagine the weight space as a map. English is a continent, and Korean is a narrow coastline attached to the edge of that continent. A narrow coastline doesn't mean a shortage of vocabulary. It means the pathways through which contexts, nuances, case law, proverbs, administrative terminology, and regional sentiment expressed in Korean can enter the model's reasoning are narrow.

Tokens are the surveying unit of that territory. Korean is an agglutinative language, so it uses more tokens than English to convey the same meaning. The same question costs more, and the same context window holds less. The territory is narrow, yet the admission fee is expensive. This isn't a matter of inconvenience but of structure. Korean users receive more expensive, shallower reasoning with less context. Has any Korean ministry ever examined this asymmetry even once with numbers?

So Individual Usage Methods Are Not the Point

Here comes a common rebuttal. "Aren't Korean companies building their own Korean language models? The market will solve it." Half true. Naver's HyperCLOVA, LG's EXAONE, and Kakao's attempts are genuine territorial expansion. But the market only expands territory; it doesn't measure and protect it. Companies have motivation to boast about their own models' Korean performance, but no motivation to publicly track whether Korean is shrinking or growing year by year within global models. That's the state's job.

The moment you narrow the problem to individual prompt skills, the essence disappears. Lectures teaching "how to use AI well" are abundant. The question we're actually not asking is this: Within the models our next generation will routinely delegate thinking to, how much space do knowledge and sentiment accumulated in Korean occupy? Who measures that space?

The Blank Korea Has Yet to Define

Here lies an institutional void. Korea possesses public language data. The National Institute of Korean Language's corpora, the Ministry of Government Legislation's legal data, vast administrative documents and case law. Yet there are no national standards on what conditions these enter global model training under, whether they should enter, or what we receive in return. Copyright law amendment discussions remain focused on creator protection, and nowhere in the administration has the perspective taken root of treating public data as a strategic asset for expanding 'Korean territory within model weights.'

The accountability void is the same. When global models answer Korean contemporary history incorrectly to Korean teenagers, or respond to Busan disaster response information in American style, whose responsibility is that error? We have no indicators to measure hallucinations arising from gaps filled with English data as 'cultural loss.' What is not measured is not managed, and what is not managed quietly shrinks. Korean's place can diminish year by year in this way, without anyone ever having decided it.

This is not an emotional appeal but a measurable proposal. We could start with a single public observatory that annually publishes Korean token share, Korean inference accuracy, and Korean usage costs across major global models. Just as exchange rates are announced daily, this would be the work of announcing the exchange rate our language occupies inside the machine's mind.

A Country That Understands Properly, Rather Than Uses Well

Many countries will become good at using AI. Everyone accesses the same models, after all. The real difference splits afterward. Between countries that can measure what proportion of our language and knowledge this technology contains and raise that proportion as national strategy, and those that cannot. The former protect Korean territory within models; the latter rest assured by fluent answers, unaware their territory is shrinking.

Do we understand this technology socially? Not yet. Korean shrinking within the weights is not a translation problem but a sovereignty problem. Before becoming a country that uses AI well, we must become a country that properly understands where our language stands inside the machine.

This article was automatically translated from the Korean original by AI. For the authoritative version, read it in Korean.

한국어 원문 읽기 →