What 100 Trillion Tokens Really Reveal About LLM Use

The December 2025 OpenRouter study of more than 100 trillion tokens revealed a platform where coding, creative roleplay and agentic workflows consumed an outsized share of model activity. Those findings remain useful, but they describe traffic routed through a multi-model inference service—not how everyone uses every large language model.
The evidence has since expanded without removing that boundary. A June 2026 paper based on licensed OpenRouter data covers 380 trillion tokens, more than 400 models and activity from January 2024 through April 2026; its authors estimate that the platform represented approximately 2% of global monthly LLM token consumption at the cited run rate. The larger dataset strengthens the case that long, tool-mediated and multi-model workloads are growing, while making clear that OpenRouter is a detailed window into one part of the market rather than a global census.
Token volume measures workload, not popularity
The most important distinction is between tokens and people. Token totals measure the amount of text processed or generated, so a long coding agent or persistent fictional conversation can outweigh many short factual queries. The study can show which workloads consume inference capacity, but it cannot establish how many individual people prefer each use case.
The same caution applies to sessions, commercial value and completed work. A large category may reflect extensive context, repeated tool calls or automated loops rather than a proportionally large audience. Nothing in token volume alone reveals whether a story was published, a software patch was accepted or a creator earned revenue from the resulting output.
OpenRouter’s audience also differs from that of a single consumer chatbot. Its common interface is designed for access to models from multiple providers, which makes the data particularly informative about API traffic, model selection and switching. Direct use of provider applications and APIs, private enterprise systems and locally hosted models falls outside this particular view.
Creative roleplay exposed a neglected form of demand
The study’s most striking creator-economy finding was the prominence of roleplay within its open-model segment. The category included character conversations, games, interactive narratives and other creative dialogue. It did not mean that roleplay dominated all artificial-intelligence activity, because the reported share applied to tokens in a defined model segment and observation window.
That narrower result is still significant. Conventional benchmarks favor contained problems with verifiable answers, while sustained fictional interaction depends on continuity, characterization, responsiveness and the ability to preserve a premise over a long exchange. High roleplay traffic therefore reveals demand for qualities that may be poorly represented by benchmark rankings.
Roleplay also crosses the boundary between entertainment and production. The same interaction structure can support dialogue exploration, character development, branching narrative design and game prototyping. The data identifies the behavior, however, not its eventual purpose: it cannot distinguish a private pastime from material that becomes part of a commercial creative project.
Coding changed the shape of measured consumption
Programming was the clearest broad shift in the original analysis. Coding workloads often carry repository context, documentation, error logs and prior revisions, making them substantially heavier than isolated question-and-answer exchanges. Their growing token share therefore reflects both greater use of coding systems and the unusual size of the requests those systems generate.
This is why a workload-level result should not be restated as a claim that most users are programmers. A software agent can inspect files, propose changes, invoke tools and revise its output within one continuing process. Each stage adds tokens, allowing a relatively concentrated group of intensive workflows to reshape the platform-wide distribution.
The pattern matters beyond software development because it shows models becoming components inside applications. Once an interface manages context, tools and repeated calls, the user may no longer select a model for every turn or even see every intermediate exchange. The unit of AI use begins to shift from a discrete chat message to an extended computational job.
Agentic use now appears beyond OpenRouter
Newer provider telemetry supports that directional change, although it cannot validate OpenRouter’s exact category shares. The June 2026 Anthropic Economic Index says Claude activity increasingly includes long-running agentic tasks through Claude Code and Cowork, and separates consumer conversations, coding sessions and first-party API traffic because those surfaces produce different patterns. This is independent evidence that chat transcripts alone no longer capture the full structure of model use.
The comparison also demonstrates why percentages from different providers should not be merged. Anthropic observes activity within its own products, while OpenRouter observes requests sent across a marketplace of models. Differences in interfaces, customers, pricing and available tools can change the task mix without either dataset being incorrect.
Agentic traffic presents an additional measurement problem: some model calls are initiated by software rather than directly by a person. A human may define the objective, but an application can generate many subsequent requests while reading context, using external tools and checking its own work. Token growth can consequently reflect deeper automation as well as rising human demand.
Multi-model use is the durable update
The expanded dataset adds an important dimension that the original public discussion could easily obscure: many OpenRouter accounts use models from multiple providers. Within this environment, model choice behaves less like allegiance to one universal assistant and more like infrastructure selection. Different systems can be used for code, narrative work, reasoning, speed or cost within the same broader workflow.
For creator tools, that pattern shifts value away from a model name alone. Project memory, permissions, reusable context and consistent output handling can matter when the underlying model changes between tasks. This is an inference from the observed multi-model behavior, not proof that any particular product architecture will retain users or produce better creative work.
The lasting conclusion is about the shape of AI activity, not the identity of a single winner. Creative conversations show demand for sustained fictional interaction; coding shows how large contexts and iteration drive consumption; agentic systems show software taking responsibility for sequences of model calls. Together, they explain why token traffic increasingly looks unlike a collection of simple chatbot questions.
The larger evidence base sharpens rather than overturns the original result. OpenRouter provides unusually granular visibility into routed model consumption, and its data captures changes that are difficult to see in surveys or benchmark tables. Its findings are most credible when treated as a detailed map of multi-model workloads—and not as a complete map of human engagement with artificial intelligence.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.