Insights

Agent Economics Are Not Token Prices

Agent Economics Are Not Token Prices

In two days, McKinsey published two documents that belong together. On August 24, QuantumBlack released a practical guide to the economics of agentic workflows. On August 25, the firm published its 2026 State of AI survey. Read side by side, they explain why this month’s price cuts will not show up in most companies’ results.

Start with the survey. Eight in ten respondents say AI has improved their own productivity. Yet the share reporting that AI contributes to their organization’s EBIT is 37 percent, essentially unchanged from a year ago. About one in five organizations say AI operating costs already constrain how they use AI, even as most expect to raise AI investment in the coming year. Agents are spreading, but unevenly: 40 percent of respondents from companies with more than $1 billion in revenue report scaling agents, up from 27 percent, while smaller organizations stayed flat at 22 percent.

The guide explains a large part of that gap. Its most useful heading is a modest one: tokens sometimes aren’t the highest cost. In McKinsey’s banking example, token charges are 20 to 25 percent of the variable cost of running a customer-service agent. Human oversight by risk and functional experts accounts for 70 to 75 percent. Customer-facing agents in some banks can cost $20,000 to $30,000 to run as a single-agent workflow and $100,000 to $200,000 as a multi-agent team. Volume changes the math: onboarding 2,500 customers a year with a conversational agent costs $10,000 to $15,000, and doubling the volume raises the cost only to $15,000 to $20,000.

That is a different conversation from this month’s rate cards. OpenAI, Anthropic, and Google spent late July and mid-August moving unit prices. McKinsey is telling buyers the unit is the wrong object. The metric it puts first is completed-work ROI: the fully loaded cost of finishing the job with people, agents, and deterministic systems, measured against the value produced.

An agent is a workflow with retries, tools, reviewers, and a failure mode that lands on a person. Pricing it like a cheaper employee who happens to bill in tokens misses most of the bill.

What actually changed

McKinsey’s July work had already warned, citing research on agentic coding tasks, that agents can consume on the order of 1,000 times more tokens than single-turn coding or chat, and that runs of the same task can vary by about 30 times. The August guide puts a total-cost frame around those numbers.

Four costs now sit on the same page.

First, tokens and thinking traces. A cheaper workhorse model still multiplies when the agent plans, calls tools, retries, and resends long-lived context.

Second, the fixed layer. McKinsey counts AI infrastructure and agent orchestration, including the data scientists who maintain the agent in production, as fixed costs per agent. Multi-agent teams make that layer heavier through handoffs, shared memory, and extra calls to reconcile conflicting steps. Volume only helps once this layer is paid for.

Third, human oversight. In McKinsey’s onboarding example, 10 to 20 percent of runs go to risk and functional experts. That reviewer belongs in the unit cost of the agent. Filing it under a separate “change management” line hides it.

Fourth, the work the agent is allowed to touch. Write actions, customer messages, filings, and production tickets carry a different expected-loss profile than a draft that never leaves Slack.

Finance teams that only watch dollars per million tokens will approve the wrong systems and kill the right ones. A high token bill on a workflow that removes a four-person queue can be cheap. A tiny token bill on a workflow that still needs a lawyer on every fifth output can be expensive.

Why owners should care now

Agent Economics Are Not Token Prices

The survey’s EBIT number and the guide’s oversight number describe the same problem from two ends. Agent pilots fail the way cheap-model migrations fail: the demo measures the model, and the production bill measures the system.

The failure mode is specific. A team shows that an agent “handles” onboarding or ticket triage. Token cost looks small next to a loaded salary. After go-live, exception rates stay high, context windows keep growing, and the same senior people who were supposed to be freed are now reviewing machine output. The token invoice is the only number that is easy to see, so it becomes the argument. The oversight cost stays in the salary budget and never joins the business case. That is how productivity rises at the desk while EBIT stays flat.

There is a second failure on the other side. Leaders freeze agent work because a multi-agent stack looks expensive on a slide, without asking whether volume will amortize the fixed layer. McKinsey’s arithmetic is the reminder: once the platform exists, incremental runs can be cheap. In its account-opening example, the fully loaded cost of onboarding one customer can fall from about $50 to $150 to about $10 to $30. That only holds if someone owns routing, caps, and review.

The survey carries a third signal worth watching. Thirty-two percent of respondents say their organizations decided against buying at least one software product or feature because they could build it in-house with agentic coding tools. That can be a real saving. It also moves a vendor’s cost into your own run budget, your own review queue, and your own liability.

Owner question

For the agent you are about to scale, what share of the run cost is tokens, what share is a human who still has to look, and who owns the run when both numbers move?

The false comfort of “the model got cheaper”

Agent Economics Are Not Token Prices

This month’s price cuts do not rescue a badly designed agent. They make it easier to run a badly designed agent more often.

A Luna-priced first hop is useful if the second hop is gated. It becomes expensive if every cheap call can open a tool that writes. A Flash-priced coding agent is useful if retries are capped and diffs are reviewed. It becomes expensive if it runs all night against a repo nobody is watching.

Vendor rate cards will keep falling and rising on their own calendar. Oversight ratios will fall only if the firm designs them down: better retrieval so the agent sees the right file once, clearer write permissions, a reviewer queue with an SLA, and a stop rule when variance between runs blows out.

McKinsey’s answer is an AgentOps capability, the agent equivalent of FinOps: a cross-functional team that manages spend continuously and brings AI unit economics into quarterly business reviews. That work belongs to operations. Procurement cannot do it alone.

DNLA Playbook for Agent Economics

  • Price the workflow. Tokens, orchestration, review time, exception handling, and expected loss from write actions belong on one sheet.
  • Separate fixed platform cost from variable run cost. Volume only helps after the platform is real.
  • Measure variance as well as averages. A 30-times swing between runs is a control problem.
  • Cap loops, retries, and context growth before go-live. Long-lived context is a recurring charge.
  • Assign a run owner. Someone has to stop an agent when oversight load or token variance leaves the band you budgeted.
  • Budget human review as a standing cost. If the design assumes 10 to 20 percent of runs need an expert, budget the expert.
  • Re-open build-versus-buy when agents start replacing software purchases. Building in-house with a coding agent still leaves you with run cost, review cost, and liability.

DNLA Take

DNLA Take

The mid-August price war answered a question most boards were already asking the wrong way. Token prices matter, but they do not decide whether an agent pays for itself. A survey in which eight in ten people feel more productive while fewer than four in ten report EBIT impact is what that mistake looks like at scale.

An agent pays for itself when the fully loaded run, meaning the model, the tools, the retries, and the human who still signs, is cheaper than the work it replaces, at a quality the firm can defend. That number does not live on a rate card. It lives in the workflow.

If you cannot say what an agent costs on a bad day, you do not have agent economics. You have a token invoice and a hope that oversight stays free.

Want the same rigor applied to your own AI system?

That's what a QAi Health Check is for.

Get in touch