Corporate America's AI Spending Spree Cools: Is 'Tokenmaxxing' Over?
A corporate fad known as “tokenmaxxing,” involving extensive use of AI, is facing a backlash as companies grapple with rising costs and limited productivity gains. The industry is shifting away from maximizing token usage towards more disciplined AI application, model routing, and exploring cost-effective open-source alternatives.A corporate trend dubbed “tokenmaxxing,” which involves maximizing the use of artificial intelligence (AI) technology, is encountering significant limitations as businesses implementing AI across all operations observe a rise in costs without a corresponding increase in productivity. This phenomenon, initially fueled by tech industry enthusiasm for extracting maximum AI-generated work from platforms like OpenAI’s ChatGPT and Anthropic’s Claude, has now shifted into a period of backlash.
Vincent Gusdorf, head of AI analytics at Moody’s Ratings and author of a new report advocating for a more disciplined approach to AI, highlighted the ease with which “you don’t need with AI.” “Tokenmaxxing” refers to the practice of maximizing the usage of “tokens,” the fundamental building blocks of generative AI. These tokens represent small pieces of text that an AI system processes or generates, with each token roughly equivalent to three-quarters of a word. AI products typically impose limits on token usage, with more expensive versions offering higher caps. Gusdorf noted that as AI bills accumulated, organizations began to recognize the substantial cost of these new tools and the necessity of employing them judiciously.
Just a few months prior, Silicon Valley executives had championed high token consumption as an indicator of high-performing employees. The archetypal “tokenmaxxer” was depicted as someone working tirelessly, leveraging numerous 24-hour AI agents to perform tasks. OpenAI CEO Sam Altman expressed enthusiasm in May for “tokenmaxxing startups,” anticipating their internal operations and product development. Nvidia CEO Jensen Huang controversially stated, “if your $500K engineer isn’t burning $250K in tokens, something is wrong.” Facebook parent Meta even organized an internal competition to reward token usage. This trend initially bolstered revenue for prominent AI large language model developers like Anthropic and OpenAI, but it waned as it became evident that it was not an optimal strategy for many other businesses.
Microsoft CEO Satya Nadella acknowledged the addictive nature of tokenmaxxing but cautioned in a recent blog post that customers are effectively paying twice for AI: first through token expenditure and second by feeding their proprietary data to these models. While promoting Microsoft’s own AI approach, Nadella’s remarks were notable for raising concerns about the data protection assurances provided by leading AI providers. Palantir CEO Alex Karp went further, stating to CNBC that something had gone “completely wrong,” relaying the sentiments of American businesses privately “livid” about excessive token costs yielding no value. Karp articulated the prevailing view among enterprises as: “‘I’m going to chillax and waste my time with tokens. I’m going to get no value and they’re going to get my IP.’”
Workplaces are now increasingly seeking improved “routing” for their AI tasks. Jue Wang, a management consultant at Bain & Company, observed that many large businesses her firm advises are scrutinizing the returns on their AI investments. Wang noted that token costs for these companies were “doubling, almost every other month.” She illustrated the financial impact: “Let’s say $200 per developer per month. Multiply that by 20,000 developers, which is often what we’re dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for.” This often implies avoiding the use of powerful, expensive AI models for simple tasks. Wang cited the example of Anthropic’s more capable Claude Opus 4.6, suited for complex tasks like software engineering or deep research, yet many companies default to using it for everything, including generating emails. This situation has spurred a demand for AI “model routing” tools, which automatically direct simpler queries to more affordable and efficient AI systems, reserving more powerful models for complex assignments.
Open-source AI models developed in China are emerging as more cost-effective alternatives. Software developer Hassan El Mghari, who leads developer experience at Together AI (a startup providing open-source AI models to developers), stated that the “ridiculous amount of money” spent on subscriptions to leading U.S. AI products has led many companies away from rewarding high usage. El Mghari suggests empowering employees to use AI when and as much as needed. Concurrently, enthusiasts of high token consumption are exploring new open-source models from Chinese startups, such as Moonshot’s Kimi or Zhipu’s GLM, which offer capabilities nearly matching top U.S. models at a significantly lower cost. Raffi Krikorian, Chief Technology Officer at Mozilla, acknowledged that this might “push tokenmaxxing a little bit further,” but believes the industry is largely recognizing that “tokenmaxxing is a dumb thing.” Krikorian drew a parallel to the outdated practice of measuring programmer productivity by lines of code, predicting that “tokenmaxxing is moving through the exact same pattern” and will be viewed as “an interesting blip that we’re all going to look back to laugh at in a year.”