Enterprise AI’s 8.3× Usage Gap Is a Proxy—not Proof of Value

OpenAI’s headline 8.3× gap measures output-token intensity per active user, not productivity, profit or return on investment. It compares the top tenth of OpenAI’s enterprise customers on that usage measure with customers around the middle of the monthly distribution.
The gap is a useful proxy because longer, tool-using agent workflows can generate substantially more output than short assistant exchanges. But token volume alone cannot show whether that output was correct, accepted, economical or responsible for a business result.
What the 8.3× comparison measures

OpenAI’s Enterprise Signals definition ranks enterprise customers each month by output tokens per active user. It classifies the top 10% as “frontier firms” and customers between the 45th and 55th percentiles as “typical firms”; the ratio between the groups increased from 2.6× in January to 8.3× in June 2026.
This is a relative benchmark, not a published count of the tokens generated by either group. Because companies are ranked afresh each month, “frontier” is a position in that month’s distribution rather than a fixed cohort of named organizations. Dividing by active users partly adjusts for participation, but an organization-wide average still does not reveal how evenly usage is distributed among employees.
The population also matters. The analysis covers aggregated, de-identified activity from OpenAI’s global enterprise customer base, not every enterprise or all AI systems used by those customers. It therefore cannot measure work performed through competing models, internal tools or non-generative automation.
Why token intensity is informative
Output tokens reveal something that login counts and seat activation cannot: how much material the systems produce for each active user. An agent that searches, edits files or completes a multi-step coding task will often generate more output than an assistant answering a short question, so rising intensity can be consistent with delegation of more substantial work.
The surrounding adoption measures strengthen that interpretation. OpenAI’s enterprise research summary states that Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers in June. Among weekly active users, plugins were used by 21% at frontier firms and 9% at typical firms, while skills were used by 19% and 3%, respectively.
These figures describe different layers of adoption. Token volume measures intensity; the Codex share reflects the output mix between agentic and chat use; and plugin or skill adoption indicates greater use of reusable instructions, company context and connected tools. Together they support a claim about depth of use, but they still do not measure the quality or economic value of completed work.
Why more output is not more value

A long response can be irrelevant, repetitive or wrong, while a short answer can prevent an expensive error. Retries, agent loops, weak stopping conditions and unnecessarily verbose responses can all increase token totals without increasing accepted work. Changes in task mix, model behavior or configuration can also move consumption independently of business performance.
Productivity evidence needs a defined outcome and a credible comparison. The NBER study Generative AI at Work examined a staggered introduction of an AI assistant across 5,179 customer-support agents and found a 14% average increase in productivity, measured as issues resolved per hour. The result rests on observed work outcomes and deployment variation—not on the amount of text the assistant generated.
Usage rankings also cannot establish causation. Firms with stronger data infrastructure, technical capacity, governance or management processes may be more capable of adopting AI intensively and of improving performance. The 8.3× comparison does not isolate AI’s contribution from those pre-existing differences.
The four layers the headline metric cannot collapse

- Token volume: The amount of model output generated per active user. It is useful for tracking intensity, capacity and cost exposure, but not correctness or usefulness.
- Depth of use: Whether AI handles longer, repeatable or tool-connected workflows rather than isolated prompts. Tokens provide partial evidence; task type, feature adoption and usage distribution add essential context.
- Operational performance: Whether a workflow reaches an accepted result at the required quality, speed and reliability. Completion rate, correction rate, cycle time, human review and cost per accepted outcome belong here.
- Business outcomes: Whether operational changes create revenue, savings, additional capacity, service improvements or risk reduction after implementation and oversight costs. This is the layer required for an ROI claim.
The distinction is consistent with OpenAI’s guidance on agentic investments, which recommends measuring the full cost of reaching an accepted result, including model and tool usage, attempts, completion rate, latency and human review. It then pairs that cost with outcomes such as time saved, shorter cycle times, protected revenue, avoided risk or added capacity.
How to read the gap without turning it into a target
The 8.3× figure is a benchmark for usage intensity, not a level every enterprise should try to maximize. Lower intensity could reflect shallow adoption, but it could also result from shorter tasks, stricter controls or efficient workflows that need little generated output. High intensity could indicate sophisticated delegation—or costly retries and poor controls.
The appropriate unit for a value claim is therefore the workflow. Usage should be considered alongside accepted task completions, quality failures, review time, total operating cost and the business measure the deployment was intended to change. Comparisons also need a defensible counterfactual, such as performance before deployment, a phased rollout or a comparable group without access.
OpenAI’s widening gap shows that some customers are asking its systems to generate much more work per active user, while related feature data indicate deeper use of agentic capabilities. It does not show that those customers are 8.3× more productive, more profitable or further ahead economically. Those conclusions require outcome and cost evidence that the token benchmark does not contain.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.