CFO guide to AI agents: fund the work you can verify

A promising pilot still needs an investment decision. Define the work you will accept, account for the full delivery cost, and separate released capacity from changed spending. Then give Finance a memo that makes the next commitment, and its limits, explicit.

A pilot can save time without earning a larger budget. Your funding decision depends on what happened to that time, the work, and the spending.

The demo may show completed conversations. Operations may report shorter handling times. Neither tells you whether the organization accepted the result or avoided an actual expense.

A useful CFO guide to AI agents starts with a narrower question: what evidence would justify the next commitment? You need a defined workflow, a credible comparison, full costs, and someone accountable for the result.

That evidence can support expansion. It can also support a smaller experiment, a narrower task, or a stop. The point is to distinguish those decisions before the annual forecast makes them look inevitable.

TLDR

  • Count accepted business outcomes, not executions or retries, and keep unresolved work visible.
  • Separate delivery cost, released capacity, and changed cash spending. They answer different investment questions.
  • Agree the next commitment, its evidence owner, and the condition that would make you reconsider it.

Start with the decision Finance must make

Your AI investment business case should name the commitment before presenting the forecast. Are you approving another test, more eligible work, or a production service with ongoing obligations?

Those decisions need different evidence. A short extension might resolve uncertainty about reviewer effort. A larger deployment requires confidence in operating costs and accountability for failures.

Consider a hypothetical service-response pilot. The agent retrieves records and prepares answers for human review. You’re deciding whether to expand its eligible workload.

A sourcing decision already exists: someone chose the platform or built the workflow. Reopening the entire build-versus-buy decision would distract from the question now facing Finance.

Instead, write a decision sentence with a condition: expand this workflow if the accepted work supports the cost and the operating owner accepts its controls.

Leave room for disagreement. The finance partner may accept the unit-cost calculation but reject the cash assumption. The service owner may accept the economics but question the workload exclusions.

A useful memo shows both objections.

A blended confidence score can hide them.

Define the work you will accept and the work you will compare

An accepted outcome is a business result that passes your stated requirements within a defined review period. Write those requirements before counting the pilot’s successes.

For the service-response workflow, acceptance might require the correct account, current supporting records, a supported answer, and completion of the required review. These are proposed criteria, not a universal standard.

A polished draft using another customer’s record fails. A correct draft awaiting review remains unfinished. A reviewed answer may count as accepted, but you should label it human-assisted.

One outcome, even when the agent retries

Separate incoming cases, eligible cases, executions, and accepted outcomes. One case might require several model calls, a retry, and a reviewer’s correction.

Those attempts generate costs. They don’t create additional business outcomes.

The FinOps Foundation’s token-economics guidance makes this distinction explicit: “Failed and discarded runs are cost, not completions.” Its outcome measure also requires a workload-specific quality floor.

Decide how you’ll treat reopened cases. If an accepted response later requires substantial correction, your measurement window needs a consistent adjustment rule.

Don’t silently drop those cases from the next report. Keep rejected, escalated, and unresolved work visible alongside accepted work.

Eligibility matters just as much. A pilot handling well-documented requests can’t establish the same acceptance rate for unfamiliar products or disputed account records.

A counterfactual that matches the workload

Your counterfactual describes what would happen without the proposed investment. It needs comparable workload, quality, service levels, and risk.

Compare reviewed agent responses with reviewed human responses, not an idealized manual process. Include the work people still perform after the agent finishes.

You can use a matched group or a controlled rollout where appropriate. If the comparison differs in complexity or seasonality, disclose that difference rather than hide it in an average.

Keep evidence ownership explicit. Operations defines acceptable work. Finance agrees how benefits will be recognized. Platform owners identify the resources consumed across successful and unsuccessful attempts.

That shared definition prevents an acceptance percentage from changing meaning between the pilot report and the investment memo.

What is cost per accepted outcome?

Cost per accepted outcome divides the cost of a defined workflow by its unique accepted results during the same period. Include unsuccessful attempts and required human work in the cost. Define acceptance before measurement. This unit-cost measure describes delivery economics; it does not establish financial return or prove that spending fell.

For recurring operations, the calculation is straightforward:

Recurring cost per accepted outcome = recurring in-scope cost ÷ accepted unique outcomes

The hard part is agreeing the terms. The FinOps unit-economics capability connects expenditure with meaningful units while balancing cost, quality, speed, and risk.

Your unit should reflect the work Finance is considering. A generated answer and an accepted response aren’t interchangeable denominators.

If another report uses “cost per successful outcome,” compare its acceptance rule before comparing the price. The label alone does not establish equivalent work.

Reconcile recurring cost before allocating setup

Include runtime, infrastructure, licenses, verification, incremental rework, and ongoing engineering or security work where they apply. Give every cost one home.

If runtime spending already includes failed calls, don’t add those charges again under a separate failure allowance. Human rework needs its own measured or disclosed basis.

Keep one-time setup visible. You can show a recurring unit cost and a full-period cost that includes setup, but label both clearly.

For detailed inputs, use the AI agent total cost of ownership breakdown. The broader AI agent cost management method covers unit-cost measurement and its operating cadence.

This memo uses those inputs to support a decision. It doesn’t need another cost taxonomy.

If no outcomes qualify, the unit cost is undefined. Report the spending and unfinished work rather than presenting zero as an attractive result.

Does the capacity gain change cash spending?

Released capacity changes cash spending only when it changes an actual outlay against a credible comparison. Faster work can be valuable without meeting that condition.

An employee might use the released time to handle a backlog or improve service. Those are potential operating benefits, not evidence that their salary disappeared.

HM Treasury’s Government Efficiency Framework, updated in November 2025, distinguishes reduced expenditure from productivity gains without reduced spending. Its scope is UK central government, not corporate accounting law.

The distinction is still useful for your review. Ask your finance team to establish the company’s own realization rule rather than borrow a convenient label.

A hypothetical memo that can survive its arithmetic

Every input below is assumed. This is a teaching example, not observed pilot data, a customer result, or a forecast for your workload.

Assume a monthly workflow contains 1,000 eligible cases and 800 accepted unique outcomes. The remaining 200 cases stay visible as rejected, unresolved, or escalated work.

Assume the cost categories include all work within that boundary. Retries consume resources but never increase the accepted-outcome count.

Hypothetical monthly inputAssumed amountAccounting treatment
Runtime, infrastructure, and licenses$2,400Includes successful, failed, and retried runs once
Routine verification$2,000Human review, excluding the separate rework category
Incremental failure handling and rework$1,200Additional work not counted elsewhere
Ongoing integration, security, and operations$800Recurring support, separate from initial setup
Total recurring cost$6,400Sum of the four non-overlapping categories
Accepted unique outcomes800Same workflow and monthly acceptance window
Recurring cost per accepted outcome$8.00$6,400 divided by 800
One-time setup cash$9,600Excluded from the recurring monthly total
Realizable monthly cash benefit$8,000Assumed avoided spending requiring independent evidence

In short: the example delivers an $8 recurring unit cost, but its cash result depends on the separate $8,000 benefit assumption.

For this illustration, assume the recurring costs are incremental cash outlays. If your numbers contain unchanged salary allocations, separate the economic and cash views first.

The assumed benefit could represent a cancellable external-service expense for the same accepted work. No real contract cancellation is claimed here.

Under constant monthly costs, benefits, and volume:

Hypothetical calculationResult
Monthly net cash: $8,000 − $6,400$1,600
Simple payback: $9,600 ÷ $1,6006 months
First-year net cash after setup: 12 × $1,600 − $9,600$9,600
First-year full cost per accepted outcome: ($9,600 + 12 × $6,400) ÷ (12 × 800)$9.00

In short: including setup changes the unit-cost view, while cash payback requires a positive recurring cash benefit after costs.

These calculations omit ramp-up, discounting, tax, working-capital effects, and terminal value. They are not net present value or an internal rate of return.

The broader AI ROI discussion explains why benefits and investment scope must match. Here, the decisive question remains whether the assumed cash change can actually occur.

Set the conditions that would change the decision

An investment memo should show what would overturn its recommendation. Otherwise, your review becomes a defense of the pilot rather than a decision about it.

Start with changes you can interpret separately. In the hypothetical, 640 accepted outcomes at the same $6,400 monthly cost produce a $10 unit cost.

That calculation says nothing about the resulting benefit. Fewer accepted outcomes may change avoided spending, service performance, or reviewer demand. Re-estimate those effects rather than holding them constant without explanation.

Now reduce the realizable cash benefit to $6,000 while leaving recurring cost unchanged. Monthly net cash becomes negative $400. The setup cost has no positive simple payback under those assumptions.

You don’t need a more elaborate spreadsheet to identify that problem. You need the person responsible for validating the benefit.

Scale, extend, narrow, or stop

Separate the evidence from the decision it supports. Use the following memo as a working document, not a certification checklist.

Memo fieldWhat you need to record
Decision requestedThe next commitment: expansion, bounded learning, narrower scope, or stop
Workflow boundaryEligible work, exclusions, required quality, and the accountable operating owner
Evidence statusObserved, assumed, or missing for every input, with a source and date
Accepted workUnique qualifying outcomes, review window, assisted work, and unresolved cases
Comparable baselineWhat would happen without the investment under similar conditions
Cost and benefitRecurring cost, setup, capacity value, and a separate incremental cash calculation
DownsideWorkload changes, lower acceptance, increased oversight, price exposure, and exit costs
ControlsData and action limits, incident ownership, review requirements, and recovery arrangements
Decision conditionThe evidence or threshold required for the next commitment
Review responsibilityFinance and operating owners, the review date, and the action if conditions fail

In short: make the assumption, its owner, and its consequence visible in the same document.

Scale when the evidence supports the expanded scope and the owners accept the operating conditions. Don’t assume the next workload resembles the carefully selected pilot.

Extend when a specific uncertainty can be resolved through a bounded experiment. Give that experiment a spending limit, a question, and an end date.

For example, you might test whether reviewer effort stays stable across a different request type. “Collect more data” is too vague to justify indefinite spending.

Narrow when value depends on a better-defined boundary. A workflow may work for current documentation but fail on disputed or incomplete records.

Stop when the evidence no longer supports the commitment, or when an unacceptable control failure requires intervention. A financial threshold cannot excuse an unauthorized action.

Record disagreement rather than forcing consensus into the model. If Finance challenges the avoided-hiring assumption, show the case without that benefit. If Operations challenges acceptance, show the disputed cases.

A supplier-exit concern belongs here too. Identify what you would need to move: approved source access, test cases, retained evidence, workflow ownership, and ongoing work.

Do not assign an invented dollar value to every risk. Missing information is a valid memo entry, provided it has an owner and changes the decision appropriately.

For the wider committee’s responsibilities, use the enterprise AI agent buying criteria. Keep this memo focused on the commitment Finance is being asked to make now.

These questions separate a useful pilot result from a defensible investment claim. Apply the same scope and evidence rules to each answer.

A CFO guide to AI agents is useful only if it changes the next decision. It should let an owner say yes, not yet, or no.

Choose one pilot. Ask Operations to sign off on accepted work and Finance to sign off on the cash assumptions. Put unresolved conditions beside the requested commitment.

Then take that memo to the review. A forecast earns attention; a commitment needs evidence.

Frequently Asked Questions

DEVREV

See Computer work for you

Your AI teammate that finds answers, takes action, and gets work done across every tool.