Insights

A Consulting Firm Embedding a Frontier Model Is Still Selling Hours Unless It Sells Evidence

A Consulting Firm Embedding a Frontier Model Is Still Selling Hours Unless It Sells Evidence

Last week, Business Insider described the large consultancies racing to become “AI native”: new engineering tracks, rewritten job titles, and alliances that put a frontier model inside the delivery machine. PwC rebuilt its training around 30 core skills, half of them AI-centric, and opened the first engineering career track in its 170-year history. On August 23, Erik Brynjolfsson and Georgios Petropoulos of the Stanford Digital Economy Lab gave the commercial problem an economist’s frame. Consulting fees still rest on assumptions from a human-only era, and when firms use AI internally to make consultants faster, the information gap between firm and client initially grows, even as efficiency improves.

PwC has been saying the pricing part out loud since spring. In March, its US tax leader, Krishnan Chandrasekhar, told Bloomberg Tax: “Time’s becoming less and less of relevance.” The same month PwC launched PwC One, a platform where clients describe a problem, agents do the work, and PwC professionals review the output in the background, with subscription or consumption pricing on the table.

The industry is embedding models. Most of it is still billing the old unit.

Mid-market law and accounting firms face the same trap. A licensed professional who uses a model to draft faster is still on the hook for the advice. A consultancy that puts Claude, Gemini, or GPT-5.6 under every workstream has not changed the product if the client still buys a stack of hours and a slide deck.

Access to a frontier model is a cost of entry. The scarce thing is evidence that the work was reviewed, bounded, and owned.

What actually changed

The large firms now have what every mid-sized firm is being sold: a named model partner, an internal assistant, and a story about agents that replace junior production. They spent 2025 and the first half of 2026 putting Claude, GPT, or Gemini in the hands of tens of thousands of staff. Accenture alone signed deals with OpenAI and Anthropic eight days apart last December. That is real capacity, and capacity alone makes no new offer. The client cannot tell whether the memo was produced by a partner, a first-year, an agent with a reviewer, or an agent with no reviewer. The invoice still looks like staffing.

Three things moved at once.

First, production got cheaper inside the firm. Research briefs, first drafts, code, and market scans move faster. That shows up in utilization and in the temptation to keep the old rate card.

Second, clients can see the same models. A buyer who can get a first-pass market scan from Claude will not keep paying strategy-firm hours for the first pass. McKinsey’s State of AI survey, published on August 25, shows how far this has gone on the software side: 32 percent of respondents say their organizations decided against buying at least one software product or feature because they could build it in-house with agentic coding tools. Advice is next in line.

Third, liability did not move. Embedding Anthropic or OpenAI in the delivery platform does not transfer professional responsibility to the lab. PwC says as much about its own platform: accountability sits with the firm and its professionals, with a human reviewer on every PwC One engagement. If the agent writes a requirement, a control design, or a board paper, the consultancy owns the defect.

Why owners should care now

A Consulting Firm Embedding a Frontier Model Is Still Selling Hours Unless It Sells Evidence

For a mid-market firm, whether to buy the same logo the Big Four bought is the wrong question. The real question is whether your proposal still prices labor while your delivery now depends on a model you do not fully control.

The failure mode is already in the market. A firm advertises “AI-enabled delivery,” staffs fewer juniors, and keeps the hourly rate. The client pays for speed it cannot inspect. When the output is wrong, there is no log of which model ran, what context it saw, who approved the write into the client system, or what was rewritten after the model drafted it. Speed the client cannot inspect, with no record behind it, is undocumented leverage.

The opposite failure is also live. A firm panics, bans models in client work, and watches staff use them anyway. Shadow use without an evidence trail is worse than a named model with a review rule.

Brynjolfsson and Petropoulos explain why both failures hurt in the same way: each widens the gap between what the firm did and what the client can see. The product that survives closes that gap. It is a packet the client can audit: the task, the model path, the retrieval set, the human sign-off, and the residual risk the firm is willing to carry.

Owner question

If a client asked, tomorrow, which pages of last month’s deliverable were model-drafted and who accepted them, could you produce the file?

The false comfort of the alliance slide

A Consulting Firm Embedding a Frontier Model Is Still Selling Hours Unless It Sells Evidence

A partnership with a lab is a procurement event, and it proves nothing about quality.

The lab will mark output, change prices, and alter agent behavior on its own calendar. This month showed it, on watermarks and on rate cards. None of those changes tell a client whether your control design is safe to implement.

Selling evidence means you can show, for a given workstream, the default model, the tasks that are forbidden to the model, the review SLA, and the artifact that proves review happened. It also means pricing the scarce resource honestly. Chandrasekhar’s own suggestion points the way: a mix of models within one engagement, time-based pricing for the early brainstorming phase and value-based pricing once the options are clear. If the scarce resource is judgment plus a trace, price that.

Mid-market firms have an opening the platforms do not advertise. You cannot out-seat Deloitte. You can out-document it. A 40-person firm that can hand a client a provenance pack on every deliverable is selling something the 400,000-person firm often is not: inspectability.

DNLA Playbook for Evidence-Based Delivery

  • Separate internal speed from the client product. Faster drafts are a cost reduction. They become part of the offer only when the client can see the control around them.
  • Define model-allowed and model-forbidden work by artifact type: research scan, draft memo, control design, production config, client-facing advice.
  • Log the path: model, prompt class, retrieval set, reviewer, timestamp. If you cannot replay it, you cannot defend it.
  • Put residual risk in the statement of work, including who owns a defect when an agent wrote the first version.
  • Change the unit you sell where the work has actually changed. Hours remain honest for workshops, brainstorming, and board time. They are dishonest for a document an agent drafted and a partner skimmed.
  • Keep the alliance out of the quality story. Name the model in the appendix. Put the review rule on page one.
  • Reuse the same evidence pack internally. If staff cannot show how a deliverable was produced, it is not ready to send.

DNLA Take

DNLA Take

Consulting is repeating a move the professions already tried: put a frontier model under the existing factory and hope the factory becomes a product. It does not.

Clients will keep buying judgment. They will stop buying opaque hours for work a model can draft by Tuesday. The firms that last will sell a reviewed artifact with a trail. A cheaper bench with a better logo on the login screen will not hold.

If you cannot show the evidence, you are still selling hours. The model is just hiding how few of them you used.

Want the same rigor applied to your own AI system?

That's what a QAi Health Check is for.

Get in touch