Usage numbers are seductive. OpenAI calls the heaviest enterprise adopters "frontier firms," and its analysis shows they generate far more output tokens per active user than typical firms. OpenAI also says, in the same breath, that tokens are an imperfect measure of business value. Both things are true. Together, they expose a flaw running through a lot of executive dashboards right now: activity is being asked to prove value, and it can't.

Usage still matters. It shows whether people are actually trying the tools, where demand is forming, and which teams might be discovering repeatable work. But a token can belong to an accepted deliverable, an abandoned experiment, a retry loop, or a verbose answer nobody needed. The count cannot tell you which one happened.

OpenAI's Enterprise Signals analysis ranks firms by output tokens per active user and compares the most intensive users with a middle group. It also reports that intensive users adopt capabilities like plugins and reusable skills more often. Those are useful signals about depth of engagement. They are also vendor-reported observations from OpenAI customers – not evidence that the firms became more productive or profitable because they consumed more tokens. OpenAI explicitly warns that a short response can be valuable while a long one may add little. [1]

The dashboard is answering the wrong question

A usage dashboard answers: where is AI activity happening? A capital-allocation dashboard has to answer something harder: which workflows create acceptable results at a justified total cost and risk?

Confusing those two questions produces two specific mistakes. The first is rewarding volume. Teams learn that more messages, active seats, or tokens look like progress – even when the additional activity comes from correction, rework, or poorly constrained agents. The second is neglecting concise workflows. A short classification, retrieval, or routing decision can create substantial value without producing an impressive usage graph.

The companion non-peer-reviewed working paper, based on OpenAI administrative data, gives leaders another reason for restraint. It documents wide variation in adoption across organizations, roles, and tasks, and associates adoption with existing organizational and intangible investment. It does not establish that heavier ChatGPT Enterprise use caused higher productivity. The authors also distinguish the reach of a task across workers from the volume of messages that task generates. A common task need not be message-intensive, and a message-intensive task need not be economically important. [2]

None of that makes usage data useless. It changes its job. Treat usage as a discovery instrument – something that helps portfolio owners find workflows worth investigating. Don't treat it as the investment verdict.

Move from activity to accepted outcomes

Every priority workflow needs a measurement chain that connects model activity to an operating result.

Start with a named outcome. Define what the workflow is supposed to deliver and who accepts it. "Draft completed" is weak if a person must substantially rewrite the draft. "Case resolved" is weak if the customer returns with the same problem. The acceptance rule should reflect usable work, not merely model completion.

Count the full cost. Include model and tool usage, failed attempts, latency, engineering support, human review, correction, and operational overhead. OpenAI's own investment guidance recommends looking beyond token price toward useful work per dollar and cost per accepted outcome. That guidance is vendor-authored and points toward OpenAI services – but the accounting principle travels fine across providers. [3]

Pair speed with quality. Cycle time can improve while defects, escalations, or downstream cleanup rise. Track acceptance without material correction, the effort required to review, and the consequences of errors. For consequential workflows, include policy compliance and exposure to security, privacy, or customer harm.

Set a decision threshold. Before the pilot, define what evidence would justify expansion, redesign, or termination. Without a threshold, every ambiguous result becomes a reason to keep spending.

Use a measurement ladder

Executives don't need one universal AI metric. They need a sequence of metrics matched to the maturity of the workflow.

At discovery, use active users, repeat use, tokens, feature adoption, and qualitative demand to locate promising behavior. These measures can reveal that a team has found a recurring need – or that a tool is failing to earn attention.

At validation, measure accepted completion, correction effort, retries, latency, and performance on representative cases. Compare the AI-enabled process with the current process, not with a hypothetical fully autonomous future.

At the economic gate, combine accepted outcomes with full cost, cycle time, capacity created, revenue effect, or risk reduction. Choose the business measure that the workflow can plausibly influence, and be honest about where attribution remains uncertain.

At scale, watch whether quality, cost, and downside remain stable as volume and workflow variation increase. A pilot can look efficient because expert reviewers are quietly rescuing edge cases. Scaling removes that hidden subsidy.

This ladder also clarifies low usage. It may indicate poor adoption – but it may also describe a rare, concise, high-value task. Investigate the workflow before cutting it. The same discipline applies to high usage: inspect the outcomes before expanding it.

Make funding workflow-specific

An organization-wide adoption target encourages broad activity while obscuring where value is actually created. Portfolio reviews should instead require every material workflow owner to present the same compact record: intended outcome, acceptance rule, baseline, full cost, quality and risk measures, observed result, and next decision.

That record should lead to one of three actions. Expand a workflow when accepted outcomes improve enough to justify cost and risk. Redesign it when demand is real but retries, review, or failure modes erase the benefit. Stop it when activity persists without a defensible operating result.

Open the current AI dashboard and pick its highest-usage workflow. Before approving more capacity, seats, or integration work, require its owner to show accepted outcomes, total cost including review and retries, and a predeclared stop threshold. If the dashboard can't support that conversation, it's measuring adoption – not investment performance.