A team reports that a new tool saves two hours a week per employee. A finance leader asks where those hours appear in the budget. Both statements can be reasonable because time saved, capacity released, and money saved are different outcomes. An investment case becomes more credible when it explains the connection between them.

The economics of AI in IT should begin with a defined unit of useful work. That could be an accepted software change, a resolved support request, or a correctly processed document. Tool activity is informative, but the organization ultimately needs to understand what it receives for the full cost of delivery.

Separate four kinds of benefit

Task efficiency means completing a specific activity with less effort. Team productivity concerns useful output from the resources available. Financial savings mean reducing an actual expense or avoiding a demonstrable future expense. Business value may include service quality, faster learning, or additional revenue.

These measures can move differently. Faster drafting might release time that a team uses to improve testing. That could be valuable without reducing payroll. A service desk might handle more demand with existing staff, creating cost avoidance only if additional staffing would otherwise have been needed.

State which benefit the proposal targets. If the case depends on reducing expenditure, identify the expense and the decision needed to remove it. If it depends on capacity, explain how that capacity will be reassigned and measured.

Treat productivity evidence with its limits

METR's early-2025 randomized study involved 16 experienced open-source developers working on 246 tasks in repositories they knew well. Allowing the tools studied increased completion time by 19%. [1] This was a result for a particular population, tool set, and task environment, not a universal estimate for software development.

In February 2026, METR said that selection effects and measurement problems made its later experiment an unreliable guide to the current productivity effect. It said developers are "likely" more sped up than in early 2025, but that its data is "only very weak evidence for the size of this increase." [2]

These findings support caution about generalizing a single result. They also show why self-reported impressions should not be treated as equivalent to measured outcomes. A local trial should record the tools, task types, users, and operating conditions under which its results were obtained.

From task savings to realized valueTask time saved is reduced by review and correction effort to produce net capacity. Management decisions can turn capacity into service gains or demonstrated spending reduction. Full costs must be included.CorpExcellence.comFOUNDATIONS / 07From task savings to realized valueReleased time has value only when its use is explained.Task time savedMeasure comparable workInclude the baselineNet capacitySubtract added reviewand correction effortService gainsMore useful outputor improved qualityExpense reductionOnly when actual spendingor justified future costfallsAssess full economicsImplementation + operation + maintenance + supervisionCompare cost per acceptable outcomeOutcomes are not automatic.Original conceptual diagramCopyright © 2026 CorpExcellence.com. All rights reserved.
Figure 7. Original conceptual model by CorpExcellence.com. Illustrative relationships, not measured results.

Count the full cost of the service

The license or model bill is only one cost category. Include implementation, integration, data preparation, evaluation, supervision, support, and ongoing changes. For hosted services, usage can create variable expenses. For self-operated infrastructure, capacity utilization and operational staffing become additional considerations.

The FinOps Foundation identifies AI-specific challenges around cost allocation, forecasting, and optimization. [3] A practical response is to assign costs to the service or business activity being evaluated, rather than leave them spread across subscriptions and cloud accounts. Our AI cost management research goes further on this topic.

Avoid double counting. If existing staff time is already included in the baseline and future-state comparison, do not add the same hours again as a separate expense. Likewise, do not count the same released capacity as both cash savings and additional output unless the allocation is explicit and feasible.

Use a hypothetical example to test the logic

Suppose a support team processes 1,000 eligible requests each month. An assistant reduces average handling time from 20 minutes to 15 minutes, including review and correction. The calculated difference is 5,000 minutes, or about 83.3 hours per month.

These figures are illustrative assumptions, not research findings. Their financial meaning depends on what happens next. If the same staff remain on payroll and no external spending changes, the 83.3 hours are released capacity. They become measurable service value if the team uses them to reduce backlog or improve resolution quality.

Now include the tool's operating cost, maintenance effort, and any increase in escalations. Check whether the eligible requests represent the real workload. A pilot restricted to simple cases cannot establish the benefit for complicated cases that consume most of the team's attention.

Measure cost per acceptable outcome

The FinOps unit-economics guidance connects technology spending with defined units of value. [4] For an AI service, a useful denominator might be successfully completed requests meeting a quality threshold. Cost per model call can help engineers diagnose expenses, but it is not the same measure.

Include failed attempts and retries in the numerator. Otherwise, a workflow that calls a cheap model repeatedly may appear economical despite a high cost per accepted result. Report quality and latency alongside cost so optimization does not silently degrade the service.

Choose a comparison period and hold the task definition reasonably stable. Record changes in demand, staffing, and process design that could explain the result. Where possible, compare similar work under different conditions rather than attributing every improvement after rollout to the tool.

Make the expansion decision explicit

A pilot should end with a decision supported by evidence: expand, revise, restrict, or stop. Identify the uncertainty that matters most, whether adoption, review burden, quality, or operating cost. Fund the next phase to resolve that uncertainty. Our business case and ROI research explores this decision further.

The diagram shows the conditions between faster tasks and realized value. Each connection requires a management decision or evidence, rather than an automatic conversion. A defensible investment case explains those connections well enough that another person can reproduce the calculation and challenge its assumptions.

References

  1. Becker et al., METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. July 2025.
  2. METR. We Are Changing Our Developer Productivity Experiment Design. February 24, 2026.
  3. FinOps Foundation. FinOps for AI. Living guidance.
  4. FinOps Foundation. Unit Economics. Living guidance.

Sources checked October 2026. CorpExcellence.com articles are best-effort research and analysis, not professional advice.