Who Gets Paid When AI Learns From You?
The overdue argument for treating your data as work, not exhaust.

Let me walk you through the economics of something you almost certainly did today. You used a search engine, or a social feed, or a maps app. In doing so, you produced behavioural data: what you searched, what you clicked, where you paused, what you bought. That data trained the models that make those products more valuable, which made those products more valuable to advertisers, which made the whole system more valuable to its owners. The value generated by that cycle runs into hundreds of billions of dollars a year. You got a free service. This is the deal the modern data economy runs on, and for most of its history nobody has seriously questioned it.
The data labour argument is simple enough: if human behavioural data is the input that makes AI systems valuable, then the people who generate that data are performing a kind of work, and they have some claim on the value it produces. It is not a radical claim. It is a standard economic principle applied to a system carefully structured to avoid that implication. The industry's response has been to frame data as a passive by-product and the free service as your compensation. Both framings are convenient. Neither is obvious.
What I want to think about here is why the data labour argument has not gained more traction, what actual compensation would look like, and whether the case is as strong as its advocates claim, or whether it glosses over some things that are genuinely complicated.

The strongest version of the argument rests on a specific claim: behavioural data is not just an input to AI, it is the primary ingredient without which those systems have very little commercial value. A language model trained on the writing and conversations of millions of people is enormously valuable. The people whose writing it learned from got nothing. A recommendation system trained on the viewing and buying habits of billions produces ad revenue that dwarfs the cost of running the service. The relationship between the contribution and the value is direct, measurable, and systematically untranslated into any payment for the contributor.
The comparison with other kinds of work is imperfect but useful. Farmers who grow crops get paid, even if they do not capture the full value of what they grow. Factory workers who turn raw materials into finished goods get wages. The idea that data generation is different in kind, that it doesn't count as productive contribution, is not economically obvious. It is a framing choice, one that has been embedded in law and regulation partly through the industry's own advocacy. The evidence that data really is valuable in the labour sense is strong, and the value is not evenly distributed either. Some people's data is worth much more than others.

The argument runs into genuine complications too. The value of any single person's data is tiny. AI models get their power from combining vast numbers of contributions, not from the specific value of any one. Any individual data point is worth a fraction of a cent, which means a strict per-contribution compensation scheme would produce trivial payments and administrative costs that would eat the payment itself. The collective is enormously valuable. The individual, alone, is not.
The free service argument, while self-serving, is not entirely without merit either. For a lot of people in a lot of contexts, the services being offered are genuinely useful, and the implied trade for data is a real one even if the currency is not cash. The honest question is not whether the exchange has zero value to the user. It is whether the terms are fair, given the enormous asymmetry in power, information, and choice on the two sides of it.
There is also a question about what people would actually prefer. Data dividend proposals often assume everyone wants cash. In practice, people's revealed preferences about data are mixed. Many will happily trade a lot of it for convenience they value. The bigger problem may be less about cash and more about the absence of meaningful choice, transparency, and control over the terms. Real data sovereignty might matter more to more people than a data cheque.

The proposals that have been developed most seriously fall into three shapes. Data dividend schemes would require platforms to share some portion of profits with the users whose data produced them. Data cooperatives would pool individual data assets and negotiate collectively for their use. Data ownership frameworks would give people property-like rights over their personal data, including the right to license it, withhold it, or receive payment when it produces commercial value.
Each of these has merits and each has real challenges. Dividend schemes need an attribution mechanism that is politically contested. Cooperatives need collective organisation among populations that are dispersed and have very different interests. Ownership frameworks need legal infrastructure that does not currently exist. None of these problems is impossible. Serious people are working on them. But none of them is close to being law in any major country either.
What is more immediately achievable, and I think more likely to change things for the better, is a combination of real transparency, genuine consent, and meaningful data minimisation. Not compensation in the strict economic sense, but at least an answer to the information and power asymmetry that is the most immediate problem. The data labour question deserves to be part of the conversation about what the data economy should look like. Whether it ends in cash dividends, cooperatives, or stronger rights, the current arrangement, where the value of billions of people's data flows entirely to a handful of platforms, is not a law of nature. It is a policy choice. And policy choices can change.
You might also like
View all
Are We Giving AI Too Much Control?
Control does not transfer in one moment. It seeps out one small decision at a time.

The Internet Is Being Flooded With AI Slop
Near-zero production costs, and what happens to the signal beneath the flood.