Data Was the New Oil.
AI is shifting strategies - from extractive data consumption to collaborative stewardship, requiring new infrastructures designed for both human and machine interaction.
Long before large language models, there was a simple idea that shaped the modern internet:
Data is the new oil.
It captured something real. Data could be extracted, refined, turned into enormous economic value. Platforms were built to capture it. Empires were built on it. But the metaphor was always slightly wrong. Oil runs out. Data does not. In fact using it does the opposite.
The Platform Assumption
Through the 2000s and 2010s, platforms behaved as if data wasn't just oil — but some weird kind of 'renewable oil'. Every interaction produced more of it. Posts. Clicks. Likes. Connections. Behavioural traces. The system fed itself. The more it was used, the more data it generated.
This created a powerful, implicit belief: Data is abundant, self-replenishing, and safe to extract indefinitely. For over a decade, that assumption held.
The Age of Brute Force
Large language models inherited this worldview — and scaled it to breaking point. Between 2018 and 2023, a new paradigm emerged. Models were trained on vast swathes of the internet — websites, forums, books, code, documentation. Anything accessible became training material. The logic was simple and industrial: More data → better models. More scale → emergent capability. And it worked. But it depended on a fragile premise: that the supply of high-quality human-created data would remain open and inexhaustible.
The extraction phase didn't trigger alarm at first. Large-scale crawling looked like search indexing — familiar behaviour. Then LLMs changed the value dynamic. When models began to answer questions directly, summarise articles, generate substitutes for original content — they stopped indexing the web and started competing with it. The system reacted. Publishers blocked crawlers. Platforms restricted APIs. Legal challenges multiplied. The open web became a negotiated surface. And for a moment, it looked like the whole story of AI and data would be one of conflict, restriction, and diminishing returns.
Data is the new soil
The deeper issue wasn't access. It was the underlying model. Data is not oil. And it's not wind, either — not something clean and inexhaustible that you simply harvest. It's closer to soil. To understand why that matters, think about what happened to farming in the 1970s and 80s. The Green Revolution had transformed agriculture. New fertilisers, pesticides, monocultures. Yields soared. It looked like a permanent abundance. But intensive extraction came at a cost invisible in the short term: soil depletion. Strip the nutrients, kill the microbiome, repeat the same crop year after year — and eventually the land stops giving. The reckoning forced a new philosophy. Not extraction, but stewardship. Crop rotation. Fallow periods. Recognising that soil is not a substrate to be used — it's a living system to be maintained. Data works the same way. Healthy data ecosystems depend on human effort and creativity.
When treated purely as a resource to extract, these systems degrade — content optimised for algorithms rather than meaning, valuable information strategically withheld, declining trust in the platforms themselves. The fertility of the system declines. LLMs didn't run out of data. They began to degrade the conditions that produce good data. And that forced the same reckoning that farming eventually reached.
Here Is Where It Gets Strange. You might expect the story to end here. AI over-extracted. The data ecosystem reacted. Restrictions tightened. We learned the lesson of the soil. Stewardship. Careful, consensual access. Perhaps licensing regimes. Perhaps a kind of data conservation movement. That story makes sense. But something else is happening at exactly the same time — and it moves in the opposite direction. Just as platforms are restricting AI's access to human-readable data, a new kind of data environment is being deliberately built — not to keep AI out, but to invite it in. Not scraped. Designed. Not protected. Offered. This is not a contradiction. But it requires explanation.
Two Eras. Two Relationships.
The first era of AI was about reading and writing. Models consumed human-created text. They learned to speak, reason, summarise. Their outputs competed with human writers — which is why the relationship with data became adversarial. The value flowed one way. The second era - the agentic era - is about acting. We no longer just want AI to write things. We want it to do things. Book the appointment. Navigate the claim. Act on our behalf within complex systems. And for that, the data problem changes completely. An AI that acts cannot reliably navigate ambiguous text, fragmented interfaces, and implicit rules. It needs something different:
Explicit permissions
Structured representations of state
Defined capabilities
Clear boundaries of authority
Auditable actions
In other words: systems that are legible to machines by design. Not the web as it was built for humans — but environments deliberately constructed so that agents can act within them safely, accurately, and with permission.
The U-Turn
So here is the strange double standard at the heart of the current moment. We spent years — rightly — trying to protect human-readable data from AI extraction. We built legal frameworks around it. We closed APIs. We sued. And now, quietly, we are beginning to build agentic data environments — structures specifically designed to be read and acted on by machines. The reversal isn't hypocrisy. It reflects a fundamental shift in what we want AI to do. When AI was consuming our content, the value relationship was extractive. AI took. We lost. When AI acts on our behalf — navigating government services, managing our data, executing transactions — the value relationship inverts. We gain. And we are willing to provide what it needs to do that well. This is the shift from accidental legibility to intentional legibility. From a world where machines adapted to human environments — to one where we are beginning to design environments that machines can act within.
The Rise of Agentic Legibility
This is already taking shape in domains like government services. Instead of forcing agents to scrape citizen-facing pages, you can provide:
Machine-readable policy — eligibility, rights, constraints
Structured service definitions — what can be done, and how
Executable interfaces — capabilities, not pages
Full audit trails of decisions and actions
This is not scraping. It is serving the machine — because serving the machine now means serving the citizen it acts for.
We are entering a new data economy. Not one where value comes from extracting data — but from designing and maintaining high-quality data environments. In this model:
data is cultivated, not harvested
access is granted, not assumed
use is governed, not opaque
Systems are no longer just human interfaces. They are shared infrastructures — for humans and machines to act within together.
What Comes Next
AI did not emerge because the web was designed for it. It emerged because it was powerful enough to exploit a web designed for humans. But the next phase will not be built on exploitation. It will be built on stewardship — and on something stewardship alone doesn't capture: deliberate design. The first era was about learning from the world as it is. The next is about designing the world so that intelligent systems can act within it — safely, meaningfully, and with permission. The soil metaphor still holds. But now we are not just learning not to deplete it. We are learning to grow something new in it.
Anyway...