Sources and evidence
Agentic AI: Fund the Outcome, Not the Pilot · September 14–19, 2026
Return to the Weekly →AI writes faster. Does the team finish better?
Before expanding AI coding tools, require evidence that they deliver accepted work at a justified total cost and risk. Faster drafting is useful only if review, testing and correction do not consume the benefit.
METR’s three studies give leaders three practical lessons:
| What the research found | What leaders should do |
|---|---|
| Measured slowdown: In a 2025 experiment, 16 experienced developers took 19% longer across 246 tasks when allowed to use AI. They nevertheless believed AI made them faster. | Check completed work, not perceived speed. Measure the effort required to reach the same acceptance standard. |
| Uncertain current effect: A later experiment involving 57 developers and more than 800 tasks could not reliably establish the size of the current benefit because of participation, task-selection and measurement problems. | Test today’s tools on your work. Include representative tasks, failed attempts and corrections. |
| Reported value: A 2026 survey of 349 technical workers found median self-reported work-value gains of 1.4–2× across questions. These were perceptions, not verified financial returns. | Use employee feedback to find promising applications. Validate the benefit before including it in a funding case. |
Sources: METR’s July 2025 experiment, February 2026 experimental update and May 2026 survey.
The studies examine different populations, tasks and measures; their results cannot be combined into one productivity score. The actions and illustrative comparison on the preceding page are ZAPHAN’s interpretation.
What the evidence changes—and what it does not
Lower operating cost can matter without proving a profitable workload. Sangfor’s CNCF-hosted case reports monthly external invocation spending falling from RMB 400,000 to RMB 200,000. China Merchants Bank’s case reports accelerator utilization rising from 35% to more than 60%, and inference cost per million combined input/output tokens falling by more than 60% for comparable models and services. These are two organizations’ reported operating results, not independently verified, like-for-like buyer returns. Workload mix, quality, period and the complete cost boundary must be checked before using them in a funding case. Sangfor case; China Merchants Bank case. Case summaries adapted by ZAPHAN from CNCF-hosted material under CC BY 3.0; no endorsement is implied.
A native option changes what deserves comparison, not what deserves automatic adoption. GitHub’s September 14 release documents cost, quality and latency priorities for Copilot auto model selection. For eligible coding workflows, test that option before commissioning a custom routing layer. Availability, output quality, controls and full cost in your environment remain to be verified. This is a product development, not evidence of savings or a market-wide trend. GitHub release.
Supporting evidence: where FinOps can influence the choice
The State of FinOps 2026 associates senior-executive engagement with greater reported influence by FinOps practitioners over technology selection. It compares engagement at VP level and above with director-level-only engagement.
| Technology decision | Executive Summary | Detailed discussion |
|---|---|---|
| Which cloud services to use | 53% versus 12% | 53% versus 24% |
| Which cloud provider to choose | 47% versus 8% | 47% versus 16% |
| Whether to use cloud or a data center | 28% versus 6% | 28% versus 12% |
The report does not reconcile these differences. They are comparisons of reported influence—not savings, returns or proof of causation. ZAPHAN preserves both versions and derives no ROI multiplier from them.
Publisher clarification remains needed.
ZAPHAN’s practical application: bring FinOps into the production decision early enough to challenge costs and alternatives before service, provider and infrastructure commitments are made.
Source: State of FinOps 2026. Comparison compiled by ZAPHAN from FinOps Foundation material under CC BY 4.0.
Supporting context
AI value should be assessed against a defined business outcome. FinOps Foundation guidance connects technology costs with outcome measures, helping teams distinguish cheaper processing from better business performance.
The OECD’s review of experimental research finds that generative AI’s benefits depend on the task, user experience and implementation context. Those findings do not establish a transferable financial return.
For this decision, ZAPHAN proposes measuring total in-scope cost per accepted outcome against a baseline, while tracking quality, review and rework. Treat productivity improvements as evidence to evaluate—not as automatically realised financial savings.
Sources: FinOps Foundation, Unit Economics and FinOps for AI; Calvino, Reijerink and Samek (2025), The effects of generative AI on productivity, innovation and entrepreneurship, OECD Artificial Intelligence Papers No. 39. Summarised and adapted by ZAPHAN under CC BY 4.0.
This is an adaptation of an original work by the OECD. The opinions expressed and arguments employed in this adaptation should not be reported as representing the official views of the OECD or of its Member countries.
GAO’s selected federal-acquisition review supports scrutiny of cost, expertise and procurement lessons—not an all-market return estimate. GAO.
The evidence supports a careful test of each workload. It does not establish a universal productivity gain, a preferred vendor, or a standard payback period.
