Skip to content
All white papers
White paper Enterprise AI Value — Part 1 of 3

From Pilot to P&L: The Economics of Enterprise AI Value

Why about 95% of enterprise GenAI pilots never reach the P&L, the four gates that fix the operating model, and the unit economics that cut LLM run cost 88%.

6 July 2026 Toronto v1 11 min read
AI economics FinOps Value realization
On this page

Scope: enterprise AI programs at the point where pilots must justify themselves to a CFO. This part stands alone as the economics brief; Part 2 covers the failures that produce no error messages and the ten-year platform, and Part 3 specifies the implementation.

Two numbers describe enterprise AI in 2026. McKinsey estimates generative AI could add $200–340 billion of annual value in banking alone, 2.8–4.7% of industry revenues. MIT’s NANDA study found that about 95% of enterprise GenAI pilots produce no measurable P&L impact. Both are roughly right, and they are not in contradiction. The value is real; it dies in transit.

$200–340B
Annual GenAI value in banking
McKinsey's estimate — 2.8–4.7% of industry revenues
~95%
Pilots with no measurable P&L impact
MIT NANDA, 2025 — failure defined as no financial impact in the observation window
~30%
Projects abandoned after proof-of-concept
Gartner prediction for end-2025 — a forecast, not a measurement
−88%
LLM run cost on one workload
Routing, caching, and batching — quality held constant (worked example below)
Four numbers that frame this part — the prize, the failure rate, the forecast, and what disciplined unit economics recovers.

Having spent years building ML platforms inside large banks, I can report a consistent pattern: the constraint is almost never the model. The initiatives I have watched stall share causes that appear on no vendor architecture diagram: no named benefit owner, no pre-launch baseline, no integration into the production workflow, no known cost per decision, and no mechanism for terminating anything.

The failure-rate evidence

The MIT finding is widely misquoted as evidence that the technology does not work. The study — 150 executive interviews, 350+ employee surveys, ~300 deployments — defines failure as no measurable financial impact within the observation window, and its diagnosis is specific: failed deployments do not integrate into real workflows, do not retain feedback, and do not improve. Gartner predicted that ~30% of GenAI projects would be abandoned after proof-of-concept by end-2025, citing poor data quality, weak risk controls, escalating costs and, most often, unclear business value. MIT also documents a budgeting problem: spend concentrates in visible front-office pilots while measured returns concentrated in back-office automation. Allocation driven by enthusiasm performs measurably worse than allocation driven by portfolio review.

Pilots started 100
Reach a working demo ~75

Capability is rarely the constraint.

Touch a production workflow ~45

The integration gap — the pilot never enters the workflow.

Still running at T+12 months ~20

No feedback loop, quality static, users abandon.

Finance-verified P&L impact ~5
Illustrative funnel consistent with MIT NANDA (2025) and Gartner (2024) findings; exact attrition varies by portfolio. Attrition is back-loaded — pilots survive the demo stage and fail at workflow integration and measurement, stages that model improvements cannot address.

The value realization operating loop

The countermeasure is mechanical rather than cultural: six stages and four hard gates, all of them enforceable by a platform.

01
Intake
problem, workflow touchpoint, materiality tier
02
Value model
baseline capture, unit economics
03
Build
on the paved road, controls pre-wired
04
Launch
shadow mode, champion–challenger
05
Measure
benefit vs baseline, finance-verified
06
Scale or kill
quarterly portfolio review
Gate A after 2 · Value model Value hypothesis signed by the P&L owner — no signature, no build.
Gate B after 3 · Build Evaluations and validation pass, rollback tested — else no launch.
Gate C after 5 · Measure Benefit confirmed at T+90 against the pre-registered baseline — finance is the scorekeeper.
Gate D after 6 · Scale or kill Every quarter, every use case re-justified; zombies terminated by default.
Loop Measured learnings re-enter at stage 2 — the loop iterates rather than terminates.
The four gates convert proposals into evidence at defined points. Gate A determines most downstream outcomes.

A value hypothesis must be falsifiable: “Reduce average complaint-handling time 20% by Q2, measured against the 90-day pre-launch baseline, owned by the head of servicing operations.” It requires a number, a date, a baseline, and a named owner. Statements such as “improve customer experience” or “drive efficiencies” are unfalsifiable and should not pass intake. In my experience, the requirement to write this sentence eliminates roughly a third of proposals before any build spend — these are the same projects Gartner counted as abandoned after proof-of-concept, identified before they consumed eighteen months. The baseline is the other non-negotiable: measured for 90 days before launch, or the benefit is permanently unverifiable. Baselines are cheap before launch and impossible after.

Four attribution rules, agreed with finance before anything ships, keep claimed benefits credible. Pre-registered baselines only. State the counterfactual: headcount growth avoided and headcount removed are different claims with different evidence burdens. One benefit, one owner: the same saved hour cannot appear in three business cases, which is the mechanism by which claimed savings come to exceed a division’s cost base. Hours count only when the P&L owner signs what the freed capacity now does. The platform team claims enablement metrics (lead time, reuse, unit cost); the business claims outcomes. Programs lose executive sponsorship fastest when their claimed benefits fail finance review.

Worked example: benefit measurement for complaint triage

Complaint triage at a retail bank; figures are illustrative but proportionate to real deployments. Volume 400,000 cases/year; pre-registered baseline 12.0 minutes average handling time; measured at T+90 with AI triage and drafting: 9.6 minutes (−20%). Time released: 400,000 × 2.4 min = 16,000 hours/year; at $55/hour fully loaded, the gross claim is $880,000/year.

The figure that should reach the CFO is smaller. Run cost, built from the cost model in the next section, is ~$140,000/year — and tokens are the minor component; amortized platform share and maintenance dominate, a distribution that is consistently underestimated. The 16,000 hours are 8.2 FTE-equivalents spread across ~400 agents who handle complaints alongside other work: roughly ten minutes per agent per day, which is realized as queue absorption and lower overtime rather than headcount reduction. The finance-verifiable entry:

LineAmount
Gross claim — 16,000 hours × $55 fully loaded$880,000 / yr
Overtime reduction, measured+$210,000
Absorbed 9% volume growth without hiring, agreed value+$260,000
Run cost (cost model below)−$140,000
Remaining diffuse minutes across ~400 agentsclaimed at $0
Finance-verifiable net≈ $330,000 / yr

This is roughly a third of the gross claim, and the difference in behavior it produces is the practical distinction between MIT’s 5% and 95%: a program that publishes the $330k figure with finance co-signature retains credibility and funding for the next twelve use cases; a program that publishes $880k survives one audit cycle.

Cost per decision

The $140k run cost above was load-bearing, and most organizations cannot produce its equivalent. A large bank prices a wire transfer to four decimal places; few can price one AI-assisted decision. Spend visibility is now common — the FinOps Foundation’s State of FinOps 2026 reports 98% of organizations managing AI spend — but a spend total with no denominator reads as undifferentiated cost, and undifferentiated cost lines are the first cut in any austerity cycle.

The cost structure of AI systems makes this measurement urgent. Unlike conventional software, the dominant bill arrives after go-live: vendor material puts inference at up to 90% of ML infrastructure spend (AWS’s framing; independent 2026 estimates are nearer 80%). Inference cost is perpetual, volume-linked, and elastic without procurement events — a single team adding a retrieval step can triple token consumption overnight.

Worked example: LLM cost reduction at one million documents per month

Document triage — classify, extract, summarize, route — at one million documents per month, averaging 3,000 input and 400 output tokens per document. Illustrative mid-2026 list prices: frontier-class model at $2.50/$10.00 per million input/output tokens; small-tier model at $0.05/$0.40. Prices are the most perishable numbers on this page; the levers are durable. Verify current rates — our LLM Cost Calculator tracks list prices across providers and hosting tiers, and models routing, caching, and batching against your own workload.

StepChangeCost/docMonthlyReduction
BaselineEverything to the frontier model$0.0115$11,500
1 · Route by difficulty80% of documents are routine — small model at $0.00031 each; 20% escalate$0.00255$2,550−78%
2 · Cache the prefix2,000 of 3,000 input tokens are shared instructions/schema; cached input bills ~50%$0.00205$2,050−82%
3 · Batch the non-urgent70% of volume tolerates 24h turnaround at ~50% of interactive price$0.00133$1,330−88%
Baseline — everything to the frontier model $11,500/mo
Route by difficulty — 80% routine to the small tier $2,550/mo

−78% vs baseline

Cache the shared prompt prefix $2,050/mo

−82% vs baseline

Batch the non-urgent 70% at off-peak pricing $1,330/mo

−88% vs baseline

Same workload, same quality bar: $11,500 → $1,330 per month. All three reductions come from architecture — routing, caching, deferral — with quality held constant by the evaluation suite at every escalation threshold. No model change is involved.

Two implementation notes constitute the actual work. The 80/20 routing split is discovered rather than assumed: run the small model in shadow with the evaluation suite as referee, and set the escalation threshold where measured quality holds. Caching is prompt discipline: stable instructions first, volatile content last, enforced in code review rather than purchased. Implementation effort on a functioning platform is two to three engineer-weeks; the annualized saving on this single workload is ~$122,000, and the pattern repeats for every LLM workload because on a shared platform these are configuration defaults. The same discipline governs GPU spend: an eight-accelerator cluster at $35,000/month running 35% utilization has an effective compute price ~2.9× the invoice, so the durable levers are structural — scale-to-zero defaults, interruptible/spot capacity with checkpointing, right-sizing before committing — while discount percentages are pricing levers, perishable and revised quarterly. Reserved LLM capacity follows the same rule: it beats pay-as-you-go only above a sustained-utilization break-even (commonly 40–60%), computable only from 30–60 days of your own gateway telemetry, so commit late, for short terms, and only to the measured floor.

None of this works without attribution. Tags are enforced at provisioning: resources that cannot be attributed cannot be created. Every model call is metered at one gateway — if token spend is unattributable, the gateway is not actually in the request path. Budgets on metered products block rather than alert, with a signed exception path. One cost schema (FOCUS) normalizes cloud, SaaS and token invoices so a single dashboard can divide verified cost by business volume. Showback runs for two quarters before chargeback, so allocation errors surface while the stakes are informational.

Marginal cost across the use-case portfolio

Unit economics has a second denominator: the marginal cost of the next use case. Sculley et al.’s finding that the ML code is a small box inside vast surrounding infrastructure has a corollary: the surrounding infrastructure is nearly identical across use cases. Ingestion, pipelines, registries, deployment, monitoring, metering — 70–80% of any new use case is plumbing that previous use cases already built. In practice this is the difference between a first use case delivered in twelve weeks and a sixth delivered in three, and between a business case that must justify the platform and one that inherits it. The declining marginal cost of use cases two through n is the platform team’s quantitative justification. It also explains the survey finding that data teams spend roughly half their time on preparation and plumbing (Anaconda’s 2020 survey: ~45%): un-amortized plumbing is exactly the cost a platform converts from per-project expense into shared capital.

The gates compound the same way at portfolio level, by moving spend from late, expensive failure to early, cheap failure. Ten candidate use cases at $250k each to build and run to a twelve-month verdict: ungated, all ten get built, nothing is measured, so nothing formally fails — $2.5M, plus $300k/year of zombie run cost continuing indefinitely. Gated, four are rejected at the hypothesis stage ($5k each for a one-pager and a workshop), two fail at launch review ($80k each), and four winners run to verdict: ≈$1.18M to identify the same four winners. The gated portfolio also produces a review culture in which a Gate-A rejection costs a proposer one week, so proposal volume rises rather than falls. A portfolio in which nothing is ever terminated is not being managed; the kill discipline is what funds the winners’ scaling budget.

Due-diligence questions for executives and platform teams

For budget holders, six questions expose most unsubstantiated programs in one meeting:

  1. Who signed the value hypothesis?
  2. What was the baseline, and when was it measured?
  3. Which workflow step changes?
  4. What does a decision cost, and in which direction is it trending?
  5. What share of AI spend has a named owner?
  6. What was killed last quarter?

For platform teams, every mechanism above is a product requirement: intake that will not issue credentials without a registry entry and a signed hypothesis; baseline capture as a template step; shadow mode and champion–challenger as deployment defaults; gateway metering and blocking budgets on by default; and a quarterly portfolio report generated from the registry rather than assembled from slides.

A practical starting point requires no re-organization: take the two use cases nearest production, write the hypotheses they never had, obtain signatures or terminate them, and put one page in front of the executive committee containing hypotheses, baselines, verdicts, and kills. Most programs cannot currently produce that page.


Next — Part 2: Silent Failures and the Ten-Year Platform: the failure modes that produce no error messages, the incident discipline that detects them, and the design stance that keeps a platform viable for a decade. Part 3 specifies the implementation.

References

  1. MIT NANDA, The GenAI Divide: State of AI in Business 2025 — coverage: Fortune, Aug 2025.
  2. McKinsey & Company, Capturing the full value of generative AI in banking.
  3. Gartner, press release, July 2024 — a forecast, not a measurement.
  4. FinOps Foundation, State of FinOps 2026 and FOCUS specification.
  5. AWS, EC2 Inf1 — inference-share framing (vendor); independent 2026 estimates nearer ~80%.
  6. Microsoft Learn, Azure OpenAI Batch and prompt caching — representative batch/caching economics.
  7. D. Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015.
  8. Anaconda, 2020 State of Data Science — survey-based; illustrative.

All prices and study statistics are point-in-time as of mid-2026 and carry their original caveats. Worked examples are illustrative composites, not client data.

Bring this rigor to your own AI controls.

If this series maps to a problem on your desk, a short call is the fastest way to compare notes.