Insights · No. 5

The Disconnect Between Tokenmaxxing and Outcome Focus

Token burn has become a proxy for AI adoption. Proxies get gamed — and then they get repriced to ROI.

Every new cycle goes through its goofy phases of adoption. In the enterprise AI market, we are going through that phase on the way to an AI-native solution that delivers high-margin productivity gains. The goofiness of the day centers around "inference" which is the cost of using AI and "the token" which is the unit of AI usage which have both been relegated to the safe harbor of "the cost of trying AI."

Inference can be overlooked or subsidized in a pilot or a test period, but can be a huge driver of costs. These costs need to be fully acknowledged and reconciled to develop useful business cases. If and while inference budgets exceed labor costs, AI adoption doesn't make sense. With the trend that is coming to be known as tokenmaxxing, the histrionics of the AI bubble are in full view. Token consumption has become the industry's proxy for adoption, rather than a real cost. It is the "kool-aid" of the AI bubble and its consumption is the signal for substantive AI adoption. The signal is being gamed well before it is being priced to a sustainable value. As adoption picks up, alarm bells are starting to go off as tokens move from experimental budgets to substantial operating line items.

The theater of "tokenmaxxing" is a pop culture phenomenon for geeks. Writing in The New York Times in March, Kevin Roose described engineers at Meta, OpenAI, and a widening circle of companies competing on internal leaderboards that rank how many tokens each person consumes [NYT]. Fortune and The Information detailed Meta's version, an intranet dashboard nicknamed Claudeonomics that ranked its 85,000-plus employees and handed out titles like Token Legend and Cache Wizard; in one 30-day window employees ran through more than 60 trillion tokens, the top user alone averaging 281 billion, over $1 million of compute in a month [Fortune]. Forbes found startups ranking staff up to the title of AI God [Forbes].

The reckoning is predictable. As adoption metrics give way to ROI discipline, we believe the advantage will not go to those who burned the most tokens on the newest frontier model, but to whomever treated inference as a cost of goods, matched the model to the use case, and kept the highest-value work under their own control. That gap, and why it favors vertical specialists over the frontier labs, is the subject here.

The adoption phase measures activity; the long-run ROI requirement prices the work (Lateral).

The adoption phase measures activity; the long-run ROI requirement prices the work (Lateral).

An economics principle called Goodhart's law holds that once a measure is turned into a target, it stops being a good measure. The mechanism is straightforward: a metric that reliably tracks some underlying goal only does so as long as no one is optimizing for the metric itself. The moment people are rewarded or judged on it, they optimize the number directly, and the correlation that made it useful breaks down.

The idea traces to British economist Charles Goodhart, who articulated it in 1975 in a monetary-policy context. His original formulation was more technical: any observed statistical regularity tends to collapse once pressure is placed on it for control purposes. He was pointing out that a relationship the Bank of England relied on to manage the money supply would degrade the instant policymakers tried to steer by it. The pithy version most people quote, "when a measure becomes a target, it ceases to be a good measure," was actually a later paraphrase by anthropologist Marilyn Strathern in 1997.

Token usage began as a reasonable proxy for AI adoption — a way for management to see whether an expensive capability was being picked up. The moment it entered dashboards, leaderboards, and performance reviews, it stopped measuring anything. Engineers have admitted, as reported in the Information, to padding context windows with documentation they didn't need, prototyping features they had no intention of building, and letting agents run in loops — not to produce output, but to avoid being seen as someone who "uses too little AI." When a company tells its workforce that AI-driven impact is a core performance expectation and then publishes a usage ranking, it should not be surprised that usage is what it gets.

The tokenmaxxing phenomenon is subsidized. The first relates to actual and perceived scarcity: we are in the massive infrastructure buildout phase of the cycle, where compute is the binding constraint. Consumption of a scarce resource reads as value — hence the proposition of every AI-savvy professional carrying an annual token budget. The second is a direct capital markets subsidy: the economics of many AI-based solutions currently run on a suspension of disbelief by VCs, underwritten by capital markets and by high-growth vendors who are themselves measured on adoption metrics and activity indicators rather than on returns. When both the buyer and the seller are graded on usage, nobody in the transaction is pricing the work - even if that what makes the most sense for the underlying customers.

"Cost of learning" theater

The leaderboards are the visible tip of a broader performative layer. Announcing a pilot, a partnership with Anthropic or OpenAI, a generous token budget — these have become signals of strength and leadership in a noisy market where fear of AI disruption moves from industry to industry in a manner that is sometimes arbitrary and frequently indiscriminate. Boards ask management teams for their AI story; usage metrics are the easiest story to tell, because they always go up. None of this is irrational at the level of the individual actor. It is what early innings look like: The capabilities are real. The measurement of the capability is not yet baked. Until business models and case studies materialize, activity substitutes for evidence. The formula works well for frontier models who are measured by investors on token consumption and for the investors behind those frontier models who want to see hockey stick growth, but not for the token consumers who have to foot the bill.

The cost this theater obscures is inference, waved away as a cost of learning rather than counted as a cost of goods. At scale the numbers are not rounding errors. Harvey, the legal AI company valued at $11 billion, has acknowledged that a single contract review across 100,000 contracts can generate a $20,000 token bill [Harvey co-founder comment, updated June 18, 2026]. Uber exhausted its entire 2026 agentic-AI budget within roughly four months [Forbes]. In many document-heavy tasks the token cost of an agentic pass now rivals or exceeds the cost of the paralegal or analyst whose work it was meant to replace, and legal AI providers are openly describing a token-price problem [Artificial Lawyer, updated June 2026]. Unless inference is priced into the business model from the start, the ROI that justified the adoption quietly inverts, and procurement is left holding an invoice that activity metrics were designed to defer.

The token reckoning is coming

Our view is that this window between activity and ROI closes quickly — plausibly within the next several quarters. Incremental model innovations are commoditizing for mainstream business use cases; for a growing share of workloads, the honest answer to "is the newest frontier model worth five times the cost for its incremental capability?" is already no. As token budgets formalize, willingness to pay collides with them, and CFOs and their procurement teams start asking the question that the activity metrics were designed to defer: what did we get for it? Adoption theater does not survive a collision with a CFO holding an invoice. The market's grading system will eventually reverts to ROI. It always does.

Sources

The 60 trillion tokens in 30 days stat in the graphic is from Meta's internal "Claudeonomics" leaderboard, and the original reporting was by The Information, which viewed a copy of the dashboard in early April 2026. Other citations are from The New York Times (March 2026), The Information and Fortune (April 2026), reporting on Meta's internal "Claudeonomics" Wall Street Journal (April 2026), reporting on AI-usage rankings in workplace performance culture; Third-party figures as reported; not independently verified. For more on Goodhart's law, Manheim and Garrabrant's "Categorizing Variants of Goodhart's Law" (2018) at https://arxiv.org/abs/1803.04585; Strathern, M. (1997), "'Improving ratings': audit in the British University system," European Review 5(3): 305–321. https://archive.org/details/ImprovingRatingsAuditInTheBritishUniversitySystem

Up next

Next month we pull the the five emerging rules of the road in AI-native businesses and introduce three key qualities that Lateral looks for in picking vertical-focused services and software companies that have an advantage in an AI-native market.

The Emerging Rules of the AI Market and the Three D's →

Disclosures

The views expressed in this white paper are the best-faith views of the principals of Lateral. This document is not primary research and should not be treated as such. Any such information regarding market forecasts and/or segmentation does not relate specifically to any investment strategy or offering of Lateral. This document is for informational purposes only and reflects the views of Lateral Investment Management as of the date of publication. It does not constitute investment, legal, or tax advice, nor an offer to sell or a solicitation of an offer to buy any security. Past performance is not indicative of future results.

Never miss an issue

Get new Insights by email.

Double opt-in. Unsubscribe any time.

Forward this article

Sends a formatted copy by email. Both addresses are recorded.