Current

The State of Token-Maxxing

What is a token?


A “token” in the context of generative Artificial Intelligence (e.g. large language models [LLM] like ChatGPT) is a numerical representation of a chunk of text which an LLM can process as input. When you ask ChatGPT a question, or give it a piece of code, this gets converted into tokens.




When you attach a data file like an Excel sheet to your query, this data is also converted into tokens.




3blue1brown has an incredible visualisation of how this happens.


What is token-maxxing?


Token-maxxing is the practice of maximising token-usage in the context of AI-usage, particularly in workplaces where is seen as a metric of productivity. A key risk of taking token-usage as a productivity metric is that it creates the perverse incentive of maximising token-usage, regardless of whether it leads to meaningful output.


Goodhart’s law states that “when a measure becomes a target it ceases to be a good measure”. In some companies, token-maxxing has definitely become a target. In the context of this article we refer to token-maxxing as the fruitless/manipulative version.


What has been said recently?


Token-maxxing has come under increased scrutiny in recent weeks. Nicholas Thompson, CEO of the Atlantic, recently pointed out that fruitless token-maxxing was a waste of resources, harmed the environment, produced false signals and was a security risk. The FT recently published on this phenomenon occurring at Amazon.


A HackerNews contributor wrote about a tool they’d created which simply “burned tokens” to increase token usage. It seems likely that such explicit and manipulative practices will be cracked down upon, and their perpetrators reprimanded or fired.


Whether this solves the problem will depend on whether most token-maxxing is done by this extreme group of users (1. Concentrated waste), or whether a lot of tokens are wasted by many/most users in a way that is far less obvious when looking at prompts (2. Distributed waste). This would also depend on whether the extreme wasters are also on average not producing as much value per token (shown in panel 3.).




One aspect of token-maxxing that does not seem to be getting that much coverage is this - it may well be a (seemingly) more acceptable way of monitoring employee activity, in sectors that have not traditionally done so. The output of software developers has long been measured through “KLOC” or “thousands of lines of code”, or in the case of IT support, through tickets closed.


But monitoring employees’ AI usage could serve as a proxy for their overall activity - a more comprehensive version of the Microsoft Teams “dot” (indicating whether a user is available or away, and how long they have been away for). And perhaps even one that could skirt labour regulations in some jurisdictions


The hazards of token-maxxing


So how exactly can token-maxxing be wasteful? Given that models charge for usage, the first issue is simply cost. Simon Couch estimated his own typical daily token usage at $15-20’s worth. Another is the electricity consumption. In Simon’s case he estimates that a typical day for him as a software developer his token-usage consumes around 1300 Wh of electricity - roughly equivalent to running an additional fridge the entire day.


Where to next?


It seems likely that more stories of this nature will come out. And, in time, an enterprise backlash may emerge and grow. This could lead to cutting back on AI spend and a hunt for alternative AI-productivity metrics. An interesting dilemma emerges for AI model developers: they want people to use as many tokens as possible for as long as possible. This means they are not incentivised to be token-efficient. This objective function is subject to certain constraints: if enterprises find token-spend to be wasted, they will reduce this spend at some point.


At the moment, LLMs are heavily subsidized by their developers - the average ChatGPT subscription probably costs Microsoft/OpenAI significantly more than its customer price. The only way this is sustainable is if 1) input costs for ChatGPT (energy usage and server infrastructure) drop significantly or 2) they raise prices or more widely implement a usage-based pricing policy.


Given what is happening to these input costs, 1) seems unlikely. Therefore when 2) occurs, it will be interesting to see how quickly prices rise and what enterprise sensitivity is to these price rises. In the interim, perhaps user and thereby market share growth, and with it funding and investment, will continue long enough that they can afford a few more quarters of subsidizing.


Mark Cuban attracted some controversy recently by suggesting that token usage be taxed at the provider level. This, he claimed, would incentivise LLM companies to be more efficient - to obtain a greater “useful output per token” ratio than perhaps is currently the case.

In the news

Which companies have been featured in the news for token-maxxing?



Company News source Date Summary
Disney Business Insider 27 Apr 2026 Internal AI dashboards showed heavy Claude and Cursor usage among Disney and ESPN product and tech employees.
Meta Fortune 9 Apr 2026 An employee-made leaderboard tracked token usage and encouraged competition among Meta staff.
Google Business Insider 26 Feb 2026 Some Google employees were told AI usage could factor into performance reviews.
Microsoft Business Insider 27 Jun 2025 Microsoft managers were told to consider internal AI-tool usage when evaluating some employees.

How to

How to token-maxx

Coming soon...

Report

Is your company token-maxxing?

We'd like to hear if your company is encouraging token-maxxing (e.g. token usage as a metric in performance reviews, AI-usage dashboards)

Tell us about it at hello@token-maxxing.io