The Token Bill Awaits: Inside the Industry’s Race to Manage AI’s Rising Costs
In today’s business environment, many organizations are reevaluating the costs tied to AI. By April, Uber had depleted its entire AI coding budget for 2026. Microsoft soon revoked its developers’ licenses for Claude Code after initial approval. A Priceline employee shared with TechCrunch that a routine renewal of their Cursor contract unexpectedly came with a price tag 4 to 5 times higher than anticipated.
Although per-token costs have fallen, the push for more AI integration and increasingly autonomous tools has led to a significant rise in token consumption. Companies that embraced unlimited subscription models in early 2025 are now scrambling to track their expenses, cut costs, and determine whether they can recoup any ROI amidst financial chaos.
At the same time, a new market is taking shape to tackle these issues. Startups, established companies, and a new standards organization are racing to equip businesses with the necessary tools and terminology to manage their spending.
“Just six months ago, discussions with clients revolved around ‘What can it do? Is it sufficient?’” stated Alexander Embiricos, head of enterprise at OpenAI, during an event in New York City this week. “Those conversations have evolved; they now center on ‘We’re investing heavily. What oversight and auditing capabilities do you offer? What token management controls do you have in place? How efficient are your models?’”
Amidst this shift, the Linux Foundation has announced the creation of the Tokenomics Foundation, a new standards entity focused on instilling financial discipline in AI token spending, akin to the FinOps strategy used for cloud expenses.
“In April and May, I began receiving feedback from companies stating: ‘We’re already three times over our entire 2026 token budget,’” recounted J.R. Storment, executive director of the FinOps Foundation under the Linux Foundation, in a discussion with TechCrunch. “We started hearing about crises that prompted a shift from token-maximizing and quick deployment to the necessity for guardrails and control measures.”
Tech industry leaders have expressed a strong desire to utilize the best models regardless of cost. Recent model launches, such as Anthropic’s Claude Opus 4.5, OpenAI’s GPT-5.1, and Google’s Gemini 3 Pro, have significantly enhanced the capabilities of autonomous tools, consequently increasing token usage. A notable instance involved a company that ran up a staggering $500 million bill for Claude after failing to enforce usage limits.
“It’s reminiscent of the crack-cocaine epidemic,” commented Chris Reed, senior director of IT finance at Priceline, noting that the company has started implementing token limits in certain departments. “They lure you in with trials to make you dependent, and before you know it, you’re hooked.”
Vitaly Gordon, CEO of Faros AI, recounted a conversation with a CTO who revealed, “One of my engineers spent $40,000 on tokens last month, and I genuinely don’t know whether to intervene or encourage others to do the same.”
A survey conducted by Faros in March showed that while output from 20,000 developers was rising, so too were bugs and rewrites. Research from Jellyfish indicated that developers who consumed the most tokens were roughly twice as productive as those utilizing AI less frequently, but they expended ten times the tokens to achieve that productivity.
Nicholas Arcolano, head of research at Jellyfish, informed TechCrunch via email that AI spending is skyrocketing primarily due to autonomous features, with per-developer token consumption increasing by about 18.6 times in just nine months. These figures complicate the equation regarding productivity in relation to financial investment.
“Whether the substantial spending yields benefits depends on the actual business value of the code delivered (e.g., revenue), which most companies struggle to measure,” stated Arcolano.
Some challenges in measurement stem from the extensive use of AI at present.
“Monitoring cloud costs involves hundreds of millions of rows of data monthly,” Storment explained. “In comparison, tracking token spending involves trillions of rows each month. You can’t just plug that into a simple spreadsheet. A comprehensive overhaul of your tools, specifications, and accounting systems is essential.”
At Priceline, Reed is already witnessing inconsistencies. He pointed out discrepancies between vendor reports and Priceline’s internal data.
“I started my career in telecom expense management, and I’m noticing many of the same patterns from telecom to cloud to AI,” he mentioned. “Every time a new component is introduced, it tends to be vulnerable to billing inaccuracies and opportunities for audits and optimization.”
A market is emerging to address these issues. Companies like Pay-i are focused on tracking, measuring, and optimizing the costs and performance of Generative AI investments. Pay-i also assists developers in monitoring expenditures, assessing usage, and charging users based on actual value rather than flat subscription fees.
Additionally, firms such as Jellyfish, Waydev, and Faros AI provide AI agent monitoring to demonstrate the ROI of developer tools. Storment notes that many of the 180 members of the FinOps Foundation are concentrating on this area.
Companies with established distribution channels are introducing new features to capitalize on these market opportunities. Recent developments include Ramp’s venture into AI expense management, while Datadog and New Relic have integrated services like cloud cost management, token-level observability, and GPU monitoring. At the forthcoming FinOps X conference, AWS is expected to unveil new financial management tools targeted at enterprise AI spending.
Tiffany Luck, a partner at NEA, suggests that token efficiency and observability will likely be incorporated at the “harness or app layer.” She mentioned Factory, a startup that develops AI agents for enterprises, which recently introduced a model router designed to intelligently select the appropriate model for each task.
Gordon anticipates that frontier labs and other model providers will adopt OpenRouter-style optimization to route queries to the most cost-effective models—a trend already visible in enterprise Claude expenditures.
“Your financial report for spending on Anthropic’s model might also reflect costs related to Sonnet or Haiku because they are intelligent enough to manage that,” Gordon remarked. “I believe this will increasingly become the standard.”
However, these tools are being developed without a unified language or agreed-upon definitions regarding token costs, outputs, and how to compare expenditures among service providers. This is where the Tokenomics Foundation aims to make a substantial impact.
The Foundation seeks to establish a comprehensive definition and framework for “tokenomics” alongside open standards, specifications, and metrics for AI token usage and billing; furthermore, it plans to introduce novel metrics for AI economics, such as cost-per-intelligence or tokens-per-watt. The organization intends to define metrics evaluating token factory efficiency and consumption. A formal launch is planned for July, with further member announcements scheduled for next week’s FinOps X conference.
“Token economics is inherently more abstract and elusive than anything we’ve handled at this scale previously,” remarked Nishant Gupta, chief availability officer at Salesforce. “It requires a different operational agility than the industry developed for cloud.”
Nonetheless, Goldman Sachs predicts that global token consumption will expand 24-fold by 2030. Businesses currently overspending need immediate solutions, while the foundation’s initial deliverables remain months away.
“Perhaps we’ve developed a steam engine, yet the assembly line is still a work in progress,” Gordon noted.
According to Arcolano, a sensible strategy is to promote widespread, moderate utilization.
“The best ROI comes from elevating the general middle group from low to moderate usage, rather than boosting heavy users,” he concluded.
Russell Brandom and Tim Fernholz contributed to this reporting.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


