Over the last two and a half decades, I have watched the enterprise pass through the internet, through SaaS, and through cloud. And each of those transitions taught us something that we promptly forgot by the time the next one arrived.
We are now in the middle of yet another phase shift, and the lesson we most need to remember is not a technical one at all, but financial. Artificial intelligence will not fail in the enterprise because the models are not capable enough. It will disappoint because we will consume it the way we consumed cloud in the early years – enthusiastically, without discipline, and without anybody in the organization being genuinely accountable for what was spent.
What makes AI different from every prior wave is how widely it cuts across the company. When enterprises bought cloud services in the past, they made a commitment before they truly knew what they would use it for, and that commitment was largely confined to infrastructure and engineering.
AI, however, is broader than that.
You make one commitment, and it is then drawn upon by R&D, by security, by finance, by operations, by go-to-market – by every function that can express its work as a prompt. That breadth is precisely what makes it powerful, and it is also what makes it so easy to lose control of.
This is why I believe every enterprise now needs what I have started calling a token operating model. Not a pilot budget, nor an innovation line item that somebody hides inside the IT allocation – but a genuine multi-year token budget with named ownership, forecast consumption, and a governance structure around it.
What we need is the same rigor that was eventually applied to the cloud.
The cloud lesson we should not need to learn twice
The original thesis of cloud was straightforward and, on paper, compelling. On-premises infrastructure ran at thirty to forty percent utilization. Move those workloads to the cloud and you would achieve eighty or ninety percent, and the savings would follow.
What happened across most organizations was that utilization settled at forty or fifty percent – the economics barely improved at all. While the technology was perfectly capable of delivering the promised outcome, the organizations were not managing it.
Because a commitment had been made, teams felt obliged to spend against it. Anyone could provision a service on demand, procurement had almost no visibility, and there was no control layer sitting between intent and consumption. An entire discipline – FinOps – had to be invented in response, and a meaningful number of workloads eventually came back on-premises altogether. That repatriation was not a verdict on cloud. It was a verdict on how we managed it.
I see the same pattern forming around tokens already. The same task gets run by multiple agents. Multiple functions query the same source, hitting the same data repeatedly, because nobody has built a caching layer or a retrieval layer that would allow the second person to reuse what the first person had already paid for.
This is duplication, and it is expensive, and it is entirely avoidable if somebody owns the problem before the invoice arrives rather than after.
Tokens do not respect geography
There is a second consideration that boards have not yet fully absorbed, and it changes the arithmetic considerably. When we placed work abroad, whether in India or in the Philippines, the entire economic case rested on labor arbitrage – the same task cost ten dollars an hour in one location and twenty-five or thirty in another.
Tokens carry no such arbitrage. A token costs the same whether the person invoking it sits in Bangalore, in Austin, or in Warsaw. It is traded in a global currency, at a single price, much like oil. And so, when a leader asks me whether AI is cheaper, my honest answer is that it depends entirely on where you were doing the work before and on how carefully you manage the consumption now.
Build for optionality, not for lock-in
The final element of a token operating model is flexibility, and here I hold a firm view. Make your commitment over three years or five – but preserve the liberty to buy tokens from whoever is delivering the best value at that moment, because value will keep shifting.
Just as we learned to architect for multi-cloud, we should now be architecting for multi-model and multi-provider. Nobody credibly knows what the landscape looks like in three years, and signing a decade-long agreement in a market moving this quickly removes the one advantage you have, which is the ability to negotiate.
Give a function a token budget, hold it accountable for the outcome, and measure it on transaction volume and quality rather than just on the activity. Do that, and AI becomes an operating discipline rather than an experiment – delivering results and driving success.
Fail to do it, and in four quarters we will be writing the same essay again, only this time about tokens, and how we need to rethink, again.
LLMs May (Yet) Fall Short of Saving the World">
LLMs">