Using the free version of ChatGPT or its rivals can feel like a steal, but for the providers the $2‑Billion investment into LLMs means they need to recoup costs. • They do this by offering paid tiers that unlock extra features such as code editing or billing tools.

The big shift now comes from "tokens" – the building blocks that both input prompts and model responses are broken into. • Because the number of tokens used per request fluctuates, even identical prompts can produce outputs that cost different amounts of tokens, leading to a wildly unpredictable bill.

Goldman Sachs predicts token consumption will increase 24‑times by 2030, to 120 quadrillion tokens a month, as companies move from single‑model usage to multi‑agent setups. • That growth is not matched by users’ ability to see how many tokens they are burning until the monthly bill arrives.

Microsoft has reportedly capped its employees’ use of third‑party coding tools, while Uber is reported to have spent its yearly AI coding token budget in under six months. • Both cases illustrate the scale of cost surprises that can appear.

Some start‑ups are responding by keeping usage below the monitoring radar: Oliver King‑Smith of smartR AI notes that smaller outfits use personal accounts with flat fees, a strategy the big vendors are unlikely to like. • He says “hitting a floor is inevitable – the big players will clamp down once shareholder pressure mounts.”

A CFO at an accounting software firm draws a parallel to grocery shopping: “You wouldn’t send someone home without a clear shopping list, right?” – meaning prompt clarity is critical to limit token waste.

When firms embed AI into products that reach thousands of users, token costs can snowball when the system is used for coding, testing, security or guard‑rails. • Venters warns that although the cost is variable, the act of giving the AI more information typically produces a better result.

The consequences spill to customer billing: Bill Peterson of Sumo Logic says the company is still exploring how to price new agent‑powered security services. • Idea options include flat‑rate price hikes, paying per result, or bundling incidents, but each faces the risk of price shifts from the underlying LLM providers.

In short, the industry faces a dual challenge: quantifying token consumption in real time, and building pricing models that aren’t upset by constant changes in token price and usage patterns.