As AI Spending Climbs, Enterprises Get Serious About Token Cost

As AI Spending Climbs, Enterprises Get Serious About Token Cost

With most products, the price is known before the purchase. With AI, enterprises often learn what the work cost only after the model has finished doing it. Users are familiar with consumption-based pricing in cloud usage, but with AI, some users don't know until they get the bill. Users say that's a problem and a warning for frontier AI model vendors.Adoption of AI is soaring, and massive IPOs by generative AI labs OpenAI and Anthropic are expected, but among customers, there's discontent. Opaque pricing and sometimes uncertain value are drawing concerns and complaints.Gartner recently forecast worldwide AI spending to reach $2.6 trillion this year, a 47% year-over-year gain. With costs rising, questions are being raised about the token pricing model, and alternative open-weight large language models (LLMs), including models built in China and Nvidia’s Nemotron, are getting more attention. OpenAI and Anthropic did not respond to requests for comment.Related:OpenAI Unveils GPT-Red to Test AI Model SafetyStarburst Data, a data platform vendor with more than $100 million in annual recurring revenue, runs most of its AI development production work on Anthropic, said CEO Justin Borgman. Claude and other tools are writing about a third of Starburst’s code, and the time from idea to production has dropped by up to 60%.Backward-Looking Cost ManagementToken costs are not out of control, but "they are rising faster than we thought" because they are using more of them, Borgman said. The problem is modeling the spending, he said. "We don't have good predictive modeling of our spend today."Instead, the company reviews its bill after the month closes and makes adjustments to optimize consumption. That might include examining whether a certain test was redundant, but these are after-the-fact adjustments."We're rear-looking," Borgman said.Borgman said he needs better internal tooling to understand what is likely to generate a lot of token consumption and what is not. But to Anthropic, "I would say they have to find ways to lower their costs faster if they want to capture the entire market," he said.Otherwise, he said they have to specialize around specific use cases where they can add value "beyond just the model itself."Enterprises are charged for what they send to the model in input tokens and for what the model returns in output tokens. But there is also another category, reasoning tokens, said Ken Stillwell, COO and CFO at Pega, an automation platform vendor. He described the cost of reasoning tokens as a "black box."Unpredictable CostsStillwell said his company has seen some variability in AI costs, but it hasn't been to the point where it's "shocking to the P&L."Related:Startup Raises $50 Million to Develop Sovereign AI InfrastructureStill, "the fact that you cannot predict what it's going to cost to use AI to solve a business problem, that is a problem," Stillwell added. That uncertainty alone, he said, will push people away from using AI.Stillwell's message to frontier model vendors is to make AI budgetable, not just powerful. He predicted that as models commoditize, vendors will have to cut prices and compete more on reliability, availability, and utility-like economics."I do have a lot of heartache with the amount of money that's being spent across the industry on AI," Stillwell said. "We really haven't locked in on where the real value use cases are yet -- it's a lot of speculation right now."Anthropic's own developer documentation, for instance, says customers are billed for a model's full internal reasoning and that usage varies by task.Self-Writing AI BillsShrihari Sridhar, senior associate dean at Texas A&M University Mays Business School, sees a problem in how AI costs are calculated."The billing is writing itself, which is a very unique thing in the history of humanity. Bills shouldn’t write themselves," he said. AI users are discovering that a large and often hidden share of bills is in reasoning tokens, Sridhar said. In his own use, the reasoning became so costly that the bill became “[random].”Related:Cost to Build Meta’s 5GW Louisiana AI Supercluster Hits $50 BillionSridhar said he expects AI pricing to shift to a task-based model, in which users are billed for completed tasks, such as resolved IT tickets. In such a model, buyers can more easily tie the outcome to ROI, he said.Changing customer behaviorToken prices have remained relatively stable across the major LLMs this year, according to ongoing data analysis by Silicon Data, which tracks pricing across hundreds of models."Customers who feel their AI bills are climbing are generally experiencing one of two things: higher usage volumes or a transition from promotional or subsidized packages to standard production pricing as AI deployments mature," said Carmen Li, CEO of the data analytics firm. In most cases, it's not because per-token prices have increased, she said.Spending has shifted in recent months. Silicon Data's benchmarking index rose from March through June, "not just because list prices increased, but because the token-maxing era encouraged customers to route more workloads to premium frontier models," Li said.But since June, the index has declined due to changes in customer behavior. "Enterprises have come under greater CFO scrutiny, optimized workloads, and increasingly shifted appropriate tasks to smaller or open-weight models," Li said.Looking ahead, pricing will also depend on GPU availability and competition from open-weight models. GPU availability is tied to AI providers’ ability to build data centers, she said.Token Discounting Era SunsetsLi's point is what users and analysts are saying: Token discounting is curtailing, alternative models are getting more consideration, and CFOs are playing a bigger role.The free- or discounted-token era is winding down, according to John D'Emic, co-founder and CTO of Revenium, which makes AI cost-tracking software. Many large enterprises consumed AI using credits bundled into negotiated cloud agreements with AWS, Azure or Google, or directly from the model vendors themselves, he said.As those discounts expire, it's intensifying focus on ROI, he said. D’Emic said he sees anticipated IPOs playing a role, pressuring frontier model vendors to scale back on subsidies.CFOs Want ControlThe increasing AI spend is also giving CFOs a larger role in AI cost management, said Scott Bickley, advisory fellow at Info-Tech Research Group. "Whenever the CFO grabs the reins like this, that's usually out of panic," Bickley said. "They've either been surprised by a bill already, and they're going to make sure that they're not surprised the second time," he said.Similarly, Prem Ananthakrishnan, managing director and global software and applied AI leader at Accenture, said that over the last few months, CFOs have begun taking a much more direct ownership role in AI spending. There is also an understanding that much of the spending comes from everything else AI needs -- governance layers, ongoing evaluations, and unplanned human oversight. These incur costs that can drive far higher bills than planned. Some cost pressure is unrelated to direct LLM pricing.Model Alternatives Offer Compelling EconomicsTo control costs, users are also considering a broader range of AI model options.Max Christoff, chief technology officer at Everlaw, which builds legal technology focused on litigation evidence, says he can tie AI spending to clear returns on investment but is also looking at lower-cost models to reduce token expenses. Before AI development tools arrived, a new feature generating about $1 million in annual revenue could take two engineers a year to build. But using frontier models, it might take one engineer roughly two months and about $10,000 in tokens, he said.Christoff is considering moving work that doesn't require frontier-level capability to lower-cost models such as DeepSeek, hosted privately on Amazon Bedrock. The model would run in Everlaw's own cloud environment without sending data to DeepSeek itself, he said, protecting the security of the data.The cost savings may be "anywhere from 5x to 10x cheaper for a given coding task" compared to the current state of the art from Anthropic and OpenAI, Christoff said. The quality might be "a notch down," but that matters less for "routine" coding tasks, he said."We're actively exploring it. The economics are frankly compelling," Christoff said.Similarly, Borgman, Starburst’s CEO, said they have been spending a lot of time investigating Nvidia’s open‑weight Nemotron model, among others, because of the company’s partnership with the chipmaker. He said the safest and most cost‑effective way to use these open source models is to “go all the way” and host them on premises. He says the token cost savings could be substantial.Users generally see AI models as increasingly interchangeable. If one model is ahead in capability, another might surpass it at another time.Token Price is Just One Driver of CostArnab Sen, senior vice president and head of data and AI engineering at AI and data science company Tredence, said that enterprises' ability to switch models will force vendors to cut prices. "Price is a function of consumption," he said. If a model is too expensive, the market won't adopt it, and vendors will have no choice but to respond when competitors offer comparable capability at "half the price, one-third the price," he said.Tredence's own token consumption is up sharply from a year ago -- "much, much more," Sen said, with AI-assisted teams delivering work 30% to 40% faster with fewer people. If frontier prices fall, Sen agreed that vendors are likely to make up the difference in volume as customers consume more.Another Approach: CapsAt Quadient, a communications automation vendor, CIO Nina Tatsiy is deliberately staying on Anthropic’s token‑capped Claude Team plan, even as she worries the enterprise tier effectively means “unlimited usage without guardrails.” By setting token limits, “people very quickly learn how efficient or inefficient they operate,” Tatsiy said. If a team runs out and asks for more, IT reviews how their models and prompts are set up. Under a chargeback model, departments billed for their AI consumption must show either cost savings or new revenue against that spend, she said.Kristian Luoma, co-founder of In Parallel, which makes software that gives AI tools business context to keep corporate teams aligned, said he compares AI to gasoline, because different mixtures do different things in different vehicles. But Luoma said his analogy breaks down at the pump."When you go to a gas station, you can compare the prices of gas," Luoma said. "You can't do that yet with tokens."

Original Source

Read the full article at Aibusiness →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.