Token consumption set to surge 24-fold by 2030, Goldman Sachs forecasts
Companies are confronting a mounting challenge in pricing AI services as the unpredictable consumption of tokens – the building blocks of large language models – makes traditional cost modelling all but impossible

The free versions of ChatGPT and its artificial intelligence rivals have become a familiar fixture of modern life, offering users an extraordinary bargain for services that have cost technology giants such as Microsoft, Google and Anthropic hundreds of billions of dollars to develop. Yet behind this apparent generosity lies a growing commercial dilemma, as the companies behind large language models seek to recoup their investments through paid-for versions with enhanced features for tasks such as coding and billing, while third-party firms build and sell agentic AI services trained for specific functions. The challenge, however, is that setting a price for these offerings has proved surprisingly difficult, with the economics of tokens – the mathematical building blocks that underpin LLMs and agentic systems – proving remarkably volatile.
When a user submits a prompt to an LLM, whether ChatGPT, Claude or Gemini, that query is broken down into tokens, which are processed by the model to generate a response, also delivered in token form before being converted back into readable text or executable commands. The process, however, is inherently unpredictable: subtle variations in phrasing can yield different answers, the same prompt does not always produce the same output, and different models generate divergent responses. In agentic systems, where multiple AI agents collaborate to make decisions and take actions, token consumption and unpredictability increase still further. As Simon Gooch of Saviynt, an identity management company incorporating agentic AI into its services, observed: "Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know."
While the cost of individual tokens has fallen sharply in recent years, according to analysis by Goldman Sachs, the volume of tokens consumed by businesses and consumers has surged. The bank forecasts that token consumption will increase 24-fold between 2026 and 2030, reaching 120 quadrillion tokens a month, as companies shift towards AI agents. Yet many organisations have only a tenuous grasp of how many tokens they are using – until they either exhaust their allowance or receive an unexpectedly large bill. Even Microsoft has reportedly reined back its engineers' use of some third-party coding tools, while Uber is said to have exhausted its AI coding token budget for an entire year in just a matter of months earlier this year. Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, noted that companies are frequently caught out as staff experiment with or implement AI internally, burning through tokens with little oversight. "People are finding it really hard to manage that cost… it's a non-deterministic output, so it's a non-deterministic value," he explained.
Companies are, however, finding ways to mitigate the uncertainty. Oliver King-Smith, founder of engineering software firm smartR AI, said smaller organisations can "fly under the radar and use [flat fee] personal accounts which I am sure the big vendors don't like", though he predicted that "this has to end at some point in time, because the big guys are taking a bath on those accounts." Once the major AI platforms face shareholder pressure to demonstrate profitability, he expects they "will start clamping down". King-Smith also argued that companies should be more discerning about which AI models they deploy, while Rob Steele, chief financial officer at UK accounting software firm iplicit, suggested that businesses needed to be far more precise with their prompts. "You wouldn't send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?" he observed.
The situation becomes particularly challenging when companies integrate AI into products that could be rolled out to thousands of users, Venters cautioned, as costs can escalate rapidly. Managers may find they require tokens not only for core software development but also for testing, security and implementing guard rails. "It's particularly hard when you're looking at agentic processes," he said, noting that deploying more AI agents can be achieved with a single click, whereas expanding the human workforce would entail careful deliberation over headcount and hiring. While token costs may be unpredictable, Venters suggested that companies might ultimately derive greater value from their token usage – "the more you give it, the more expensive it is, but the better the result may be" – though those costs must ultimately be passed on to customers.
"Nobody's really figured it out," said Bill Peterson, senior director of product marketing at Sumo Logic, whose firm is previewing new security services based on agentic AI and is in discussions with corporate customers about pricing models. Options under consideration include blanket price increases, payment by results, or charging for "bundles" of incidents, he revealed, adding that the company was "still having some fun conversations about this internally". Yet whatever structure is chosen, he warned, it could be upended if and when the large language model providers adjust their own pricing strategies. "You get into variable pricing, and it's changing every couple of months," he said. "Customers don't like that. That's not how anybody builds a budget."