Your vendors keep announcing that AI is getting dramatically more efficient. Your finance team keeps approving larger AI budgets. Both statements can be true at the same time, and understanding why that happens is the decision you need to make before your next planning cycle.

The mechanism behind the contradiction

Vendors are reporting meaningful efficiency gains in AI. The work of running a trained model to produce an answer – inference – is becoming cheaper per unit. OpenAI says its model used 54% fewer output tokens than another leading model. This figure is vendor-reported, not independently verified [1]. OpenAI also published results for Jalapeño, its first custom inference chip, claiming higher AI work per watt and lower end-to-end latency than named comparison systems on a public benchmark [2]. These are vendor-run comparisons using vendor-selected baselines, not independent audits, but the claim signals where the industry is heading.

The efficiency claim is plausible and consequential. It does not tell you what happens to your total spend.

Efficiency does not mean lower budgets – it means more workflows

Economists call the pattern Jevons paradox: when a resource becomes cheaper to use, total consumption rises because users expand into applications that were previously uneconomical. The same dynamic can apply to inference.

When the cost of generating an AI output falls, workflows that were previously too expensive to automate start to clear the investment bar. A legal team that rejected AI-assisted contract review at last year's per-document cost may find the arithmetic works at today's. A finance team that shelved an automated variance analysis may revisit it. Each individual workflow becomes cheaper to run. But more workflows run – so aggregate AI spend can rise, not fall.

OpenAI frames this directly. In its infrastructure strategy, the company invoked Jevons paradox to explain why greater efficiency makes more uses worthwhile and expands total consumption [1]. OpenAI and Nvidia both benefit from that expansion. Nvidia forecasts about 70% revenue growth in fiscal 2028. [3] This is a company forecast reported by Reuters, not realized revenue, and memory-component shortages may constrain that expansion. Taken together, the disclosures point to a world where cheaper inference can mean much more inference, not the same amount at lower cost.

That is the right bet for them. It may not be the right assumption for your budget model.

The vendor pass-through gap

Falling costs on the supply side do not automatically become falling prices for buyers. OpenAI has stated its intent to carry efficiency gains through to users [1], but intent is not a pricing guarantee. Nvidia is forecasting higher revenue, not less. Supply constraints, including the memory-component shortages reported by Reuters, can limit how fast lower production costs reach buyers [3].

Buyers who assume vendor efficiency gains will translate directly into proportional price reductions are making an assumption that has not been verified.

The measurement problem in your current budget

Most AI budgets track the wrong thing: seats, platform access fees, or aggregate API spend. None of those numbers tell you whether a workflow is delivering value at an acceptable cost.

The right unit of measure is the cost of a successful completed business task. That means the full cost: model calls including retries, human review time, integration overhead, and failure costs when the output is wrong and requires rework. An accepted analysis. A closed support ticket. A reviewed and approved contract. Not API calls consumed, not tokens generated – the task, completed successfully, end to end.

That unit of measure does two things. First, it makes the Jevons effect visible. If your cost per successful task falls as vendor unit costs fall, you can see which previously rejected workflows now clear your investment threshold. Second, it exposes workflows where AI activity is high but successful task completion is low – where you are paying for volume and getting noise.

The decision you need to make

Stop budgeting AI as seats or fixed platform lines. Require workflow owners to report realized cost per successful completed task, including retries, review, integration, and failure costs. That reporting discipline forces a conversation the current budget structure prevents: whether the AI spend is actually producing completed work at an acceptable cost.

Re-score workflows your organization previously rejected on cost grounds. If vendor unit costs have fallen materially, the economics on those decisions may have changed. Run the numbers against the current cost structure before dismissing workflows that failed the bar twelve months ago.

Build explicit uncertainty into your assumptions. Vendor pass-through of efficiency gains is intent, not contract. Supply constraints on components like memory can hold prices up regardless of chip-level efficiency improvements. Budget for a range of outcomes on unit cost, not a single line based on current vendor pricing.

The operating consequence

The executive who treats falling AI unit costs as a signal to hold AI budgets flat risks two surprises: total AI spend may rise because more workflows become economical, while vendor prices may not fall as fast as the efficiency headlines suggest. The executive who re-prices the workflow portfolio against current unit costs, requires task-level cost reporting, and builds supply-side uncertainty into the model is working from the actual economics, not the expanding-spend narrative vendors have an interest in promoting.

Cheaper inference does not mean cheaper AI. It means more AI. Plan for that.