The Model Is Free, Data Is the Rent
As open-weight models flood the market, weights have become a common commodity. Competitive advantage has shifted to proprietary data and feedback loops, now transformed into rental assets. We examine why Korean companies' closed-domain data is the real moat.
AI Summary
As open-weight AI models from Meta, Alibaba, DeepSeek, and Mistral have made model weights effectively free, the true competitive moat has shifted from model capability to proprietary data and self-reinforcing feedback loops. Big Tech releases its models for free not out of altruism but to commoditize competitors while retaining durable advantages through user-generated feedback and closed-domain datasets. For Korean companies, the critical choice is whether to treat their world-class domain data — in healthcare, semiconductors, logistics, and port operations — as a strategic asset to monetize, or to remain passive customers paying rent inside someone else's data loop.
What Went Free in a Week
Over the course of 2025, Meta's Llama, Alibaba's Qwen, DeepSeek, and Mistral released their weights one after another. Models that were once assets worth tens of millions of dollars became files you simply download from Hugging Face. The performance gap narrowed as well. A pattern kept repeating in which open-weight models caught up to closed models' benchmarks with a lag of just a few months.
Stop there, and you fall into a familiar conclusion: "AI has been democratized; now anyone can use a good model." That is only half true. The fact that weights have become free means weights are no longer a moat. The moat hasn't vanished — it has simply moved.
The Real Reason Big Tech Opens Its Models
Reading Meta's release of Llama as an act of charity is a misreading. Releasing a model for free collapses the pricing on competitors' closed models. The margins OpenAI was charging through its API come under pressure. Meta never intended to make money from models — it makes money from advertising and recommendation algorithms. So it deliberately drives the model layer down to cost, cutting off competitors' revenue streams. Commoditization is strategy, not accident.
So where does Big Tech maintain its edge? Three places. First is the inference infrastructure — Nvidia GPUs, cloud compute, and electricity. Second is the feedback data users generate through their interactions with the model. Third is proprietary data that originates exclusively from specific domains. Model weights can be replicated by anyone, but the preference signals produced by hundreds of millions of real users cannot.
That is also why OpenAI offers ChatGPT for free. It isn't making money from the model; it is using usage logs to train the next one. Users are, in effect, a data factory. Capital is moving in the same direction. From 2024 to 2025, the center of gravity in U.S. venture investment shifted from foundation model startups toward data infrastructure, evaluation tooling, and domain-specific applications. Money flowed to companies that feed data to models and refine that data, rather than to companies building new models. Scale AI's valuation is the signal.
From Weights to Rent
The definition of an asset has changed. The old moat was "our model is smarter." Today's moat is "we run a loop that continuously makes the model smarter using data only we possess." Models have gone free like assets you buy once and own outright, while data and feedback loops have become rental assets renewed every month. Those who pay rent are those without data; those who collect it are those who hold data.
One dangerous position exists in this structure: companies trying to differentiate through prompts by placing a thin application on top of someone else's model. Models are free and prompts are replicable, so there is no moat here. The moment a company holding the data loop absorbs the same functionality, these players disappear.
A strong counterargument surfaces here: "Public data already abounds, and synthetic data is pouring out. The advantage of proprietary data will soon be diluted." That is a fair point, and it holds in general text domains. But synthetic data only recombines what the model already knows — it cannot generate signals that can only emerge from real-world measurement. Actual surgical outcomes in the operating room, correlations between factory equipment vibration and defect rates, variables in early-morning logistics flows: these are products of measurement, not simulation. The more closed a domain, the less synthetic data can fill it.
Are Korean Companies Suppliers, Customers, or Architects?
Return to the opening question. Are Korean companies suppliers, customers, or standard architects in this competition? In the foundation model race, most are customers. That game has already been decided by capital and GPU scale. But shift to the data rental economy, and the story changes.
Look at the areas where Korea is strong. In healthcare, the single national health insurance system means standardized billing and prescription data accumulates in one place. In manufacturing, Samsung Electronics, SK Hynix, and Hyundai Motor's process data accumulates at a density found nowhere else in the world. In logistics, Coupang's dawn-delivery operational data is itself a domain asset that cannot be learned away. It is closed data produced only through measurement — data that neither Meta nor OpenAI can scrape. In Busan alone, port cargo volume and vessel arrival and departure data accumulates at the port authority. Refined into a form that can feed models, it becomes a moat in its own right.
| Korea's Strong Domains | Proprietary Data Generated Only Through Real-World Operation | |
|---|---|---|
| Healthcare | Single-payer national health insurance system | Structured claims and prescription data |
| Manufacturing | Samsung Electronics · SK Hynix · Hyundai Motor | World's highest-density process data |
| Logistics | Coupang dawn delivery | Operational data impossible to replicate synthetically |
There is one prerequisite: this only applies to companies that treat data as an asset. If a company leaves its data buried in an ERP system and relies on external APIs, Korean companies end up renting back — at a fee — models built on their own data. They sit as customers despite holding the assets needed to be a supplier.
The Cost of Waiting
Silicon Valley news is not someone else's story. Behind the headline that models have gone free lies a structural shift: competitive advantage has been reorganized around data rent. This shift will directly determine Korean companies' next cost structures. Companies that refine their domain data and run the loop will be on the side collecting rent. Companies that wait and watch will be on the side paying rent every month. The judgment that "models are free, so we can relax" is the most expensive judgment of all. If you don't build the data pipeline now, in a year you'll find yourself inside someone else's loop, with the value of your own data converted into rent.
This article was automatically translated from the Korean original by AI. For the authoritative version, read it in Korean.
한국어 원문 읽기 →