Back to blog
DATAECONOMYTECHNOLOGY

Winners in AI Are the Ones With the Best Data

The competitive moat in AI is not the model. It is what the model was trained on.

Sahir Maharaj smiling in glasses and a deep blue embroidered jacket11 min read
A heavy steel vault door slightly ajar in a dark room with a warm glow spilling out from behind it
The moat is not the model. It is what is stored behind that door.

There is a phrase in AI circles that has become so common it risks losing its meaning: data is the new oil. Like most analogies that achieve that level of ubiquity, it is both illuminating and somewhat misleading. Oil is a fixed resource: you find it, extract it, burn it, and it is gone. Data compounds in ways that oil does not. The more behavioural data a platform accumulates about its users, the better its models become, which attracts more users, whose behaviour generates more data, which improves the models further. Oil wells deplete. Data advantages grow. The analogy is better understood as: data is not the new oil, data is the new land, and the enclosure is already well underway.

The competitive dynamics of AI are being determined more by data advantages than by algorithmic sophistication, and this is a point that is often underappreciated in coverage that focuses on model architectures, parameter counts, and benchmark performance. The most important AI systems in commercial deployment are not competing on the theoretical frontier of what is technically possible. They are competing on practical performance in specific domains where the quality of performance is determined primarily by the quality of the domain-specific training data. In healthcare AI, the system with the most and best clinical data wins. In fraud detection, the system with the longest and most comprehensive transaction history wins. In recommendation systems, the system with the most detailed behavioural history of the relevant user population wins.

A stack of shiny golden coins beside a small glowing microchip on a black polished surface
The chip does the work. The pile beside it decides who gets to use it.

This data-determined competitive dynamic has implications for the structure of the AI economy that deserve more systematic attention. If data is the primary source of competitive advantage in AI, the entities that have the most and best data have structural advantages that are very difficult to overcome through technological competition alone. And the entities that have the most and best data are, by and large, the entities that have been most successful at accumulating it over the longest period of time. The AI economy is not starting fresh: it is inheriting the data advantages of the pre-AI digital economy, and those advantages compound.

The compounding nature of data advantages is what makes them so significant as competitive moats and so difficult to address through conventional competition policy. A company with more data trains better models; better models attract more users; more users generate more data; more data trains even better models. Each iteration of this cycle extends the lead of the data-rich incumbent over the data-poor challenger. The challenger not only needs to close the data gap that currently exists; it needs to close the gap that will exist at the time it is able to close the current gap, which has grown in the interval. This is a self-reinforcing dynamic that produces concentration without requiring any active exclusionary behaviour from the incumbent.

A snowball rolling down a snowy hill leaving a wide compacted trail behind it in soft winter light
Every turn of the cycle adds to the pile. Every turn makes catching up harder.

The implication for new entrants to AI markets is severe. A startup that builds a technically superior model faces the question of how to get the training data that would make that model commercially competitive. If the best training data is proprietary to incumbents who have accumulated it over years of user engagement, the startup cannot buy its way in, cannot generate it quickly enough to close the gap before the incumbent's continued accumulation extends it, and cannot train adequately on publicly available alternatives that lack the specificity and recency of the incumbent's proprietary dataset. The technical innovation the startup brings is real, but it is insufficient to overcome the structural data advantage. This is why we see repeated cycles in which promising AI capabilities are developed by smaller companies or academic researchers and then absorbed by the large platforms that have the data to commercialise them effectively.

The data moat is domain-specific, which means the competitive dynamics vary significantly across different AI application areas. In domains where data is abundant and widely available, the moat is shallower and technical competition matters more. In domains where data is scarce, proprietary, or requires long accumulation periods to be useful, the moat is deep and the incumbents with the best proprietary datasets have structural advantages that amount to near-insurmountable barriers to entry for well-funded new entrants, let alone resource-constrained ones. Healthcare, financial services, and autonomous vehicle navigation are all domains where proprietary data accumulated over extended periods creates that pattern.

A world globe on a dark wooden desk beside a stack of leather bound books in warm lamp light
The strategic question is no longer only about companies. It is about countries.

The understanding that data is a primary source of competitive advantage in AI has moved from business strategy into geopolitics. The major AI powers are engaged in explicit competition for data resources as part of their competition for AI leadership. The specific data that matters for AI is not generic: it is behavioural data from large populations using digital services, and the populations that generate the most of it are concentrated in a small number of large countries. The geopolitical implication is that large countries with large digital populations have structural AI advantages that smaller countries cannot easily overcome, regardless of the quality of their AI research.

The structural advantages that data creates in the AI economy are real and durable, but they are not immutable, and understanding what can and cannot be changed is important for anyone thinking about how the AI economy might be made more competitive and more equitable. Mandatory data sharing requirements, which force dominant platforms to share specific categories of non-sensitive behavioural data with competitors under standardised conditions, are one mechanism that could reduce data moats without eliminating the incentive to collect data. Data portability requirements, which give users the ability to take their data from one platform to another, reduce the lock-in that helps incumbents maintain their data advantages. Open dataset initiatives, which make high-quality training data available to all researchers and companies rather than concentrating it in proprietary systems, can reduce advantages in specific domains. None of these mechanisms is perfect or cost-free, but they represent the kinds of structural interventions that could meaningfully reduce the competitive advantages that current data accumulation creates.

The honest assessment is that the data advantage dynamic in AI is producing a level of market concentration and structural inequality that the current regulatory toolkit is not well equipped to address, and that the political will to apply more powerful tools is not currently present in most major jurisdictions. The winners in AI are the ones with the best data, and the best data is concentrated in the hands of a small number of incumbents whose data advantages are growing faster than any regulatory response. Whether that dynamic produces a stable and reasonably competitive AI economy or an increasingly concentrated one depends on choices being made now. The choices that would produce a better outcome are known. Whether they will be made is the question.

AI ECONOMYDATA MOATSCOMPETITION POLICYPLATFORM POWERGEOPOLITICS