---
Title: "CFO guide to AI agents: fund the work you can verify"
Url: "https://devrev.ai/blog/cfo-guide-ai-agents"
Published: "2026-10-05"
Last Updated: "2026-10-05"
Author: "DevRev Editorial"
Category: "Growth & Revenue"
Excerpt: "A promising pilot still needs an investment decision. Define the work you will accept, account for the full delivery cost, and separate released capacity from changed spending. Then give Finance a memo that makes the next commitment, and its limits, explicit."
Reading Time: 11
---

# CFO guide to AI agents: fund the work you can verify

A pilot can save time without earning a larger budget. Your funding decision depends on what happened to that time, the work, and the spending.

The demo may show completed conversations. Operations may report shorter handling times. Neither tells you whether the organization accepted the result or avoided an actual expense.

A useful CFO guide to AI agents starts with a narrower question: what evidence would justify the next commitment? You need a defined workflow, a credible comparison, full costs, and someone accountable for the result.

That evidence can support expansion. It can also support a smaller experiment, a narrower task, or a stop. The point is to distinguish those decisions before the annual forecast makes them look inevitable.

## TLDR

- Count accepted business outcomes, not executions or retries, and keep unresolved work visible.
- Separate delivery cost, released capacity, and changed cash spending. They answer different investment questions.
- Agree the next commitment, its evidence owner, and the condition that would make you reconsider it.

## Start with the decision Finance must make

Your AI investment business case should name the commitment before presenting the forecast. Are you approving another test, more eligible work, or a production service with ongoing obligations?

Those decisions need different evidence. A short extension might resolve uncertainty about reviewer effort. A larger deployment requires confidence in operating costs and accountability for failures.

Consider a hypothetical service-response pilot. The agent retrieves records and prepares answers for human review. You’re deciding whether to expand its eligible workload.

A sourcing decision already exists: someone chose the platform or built the workflow. Reopening the entire [build-versus-buy decision](https://devrev.ai/blog/build-vs-buy-ai-agents) would distract from the question now facing Finance.

Instead, write a decision sentence with a condition: expand this workflow if the accepted work supports the cost and the operating owner accepts its controls.

Leave room for disagreement. The finance partner may accept the unit-cost calculation but reject the cash assumption. The service owner may accept the economics but question the workload exclusions.

A useful memo shows both objections.

A blended confidence score can hide them.

## Define the work you will accept and the work you will compare

An accepted outcome is a business result that passes your stated requirements within a defined review period. Write those requirements before counting the pilot’s successes.

For the service-response workflow, acceptance might require the correct account, current supporting records, a supported answer, and completion of the required review. These are proposed criteria, not a universal standard.

A polished draft using another customer’s record fails. A correct draft awaiting review remains unfinished. A reviewed answer may count as accepted, but you should label it human-assisted.

### One outcome, even when the agent retries

Separate incoming cases, eligible cases, executions, and accepted outcomes. One case might require several model calls, a retry, and a reviewer’s correction.

Those attempts generate costs. They don’t create additional business outcomes.

The [FinOps Foundation’s token-economics guidance](https://www.finops.org/wg/token-economics-saas/) makes this distinction explicit: “Failed and discarded runs are cost, not completions.” Its outcome measure also requires a workload-specific quality floor.

Decide how you’ll treat reopened cases. If an accepted response later requires substantial correction, your measurement window needs a consistent adjustment rule.

Don’t silently drop those cases from the next report. Keep rejected, escalated, and unresolved work visible alongside accepted work.

Eligibility matters just as much. A pilot handling well-documented requests can’t establish the same acceptance rate for unfamiliar products or disputed account records.

### A counterfactual that matches the workload

Your counterfactual describes what would happen without the proposed investment. It needs comparable workload, quality, service levels, and risk.

Compare reviewed agent responses with reviewed human responses, not an idealized manual process. Include the work people still perform after the agent finishes.

You can use a matched group or a controlled rollout where appropriate. If the comparison differs in complexity or seasonality, disclose that difference rather than hide it in an average.

Keep evidence ownership explicit. Operations defines acceptable work. Finance agrees how benefits will be recognized. Platform owners identify the resources consumed across successful and unsuccessful attempts.

That shared definition prevents an acceptance percentage from changing meaning between the pilot report and the investment memo.

## What is cost per accepted outcome?

Cost per accepted outcome divides the cost of a defined workflow by its unique accepted results during the same period. Include unsuccessful attempts and required human work in the cost. Define acceptance before measurement. This unit-cost measure describes delivery economics; it does not establish financial return or prove that spending fell.

For recurring operations, the calculation is straightforward:

`Recurring cost per accepted outcome = recurring in-scope cost ÷ accepted unique outcomes`

The hard part is agreeing the terms. The [FinOps unit-economics capability](https://www.finops.org/framework/capabilities/unit-economics/) connects expenditure with meaningful units while balancing cost, quality, speed, and risk.

Your unit should reflect the work Finance is considering. A generated answer and an accepted response aren’t interchangeable denominators.

If another report uses “cost per successful outcome,” compare its acceptance rule before comparing the price. The label alone does not establish equivalent work.

### Reconcile recurring cost before allocating setup

Include runtime, infrastructure, licenses, verification, incremental rework, and ongoing engineering or security work where they apply. Give every cost one home.

If runtime spending already includes failed calls, don’t add those charges again under a separate failure allowance. Human rework needs its own measured or disclosed basis.

Keep one-time setup visible. You can show a recurring unit cost and a full-period cost that includes setup, but label both clearly.

For detailed inputs, use the [AI agent total cost of ownership](https://devrev.ai/blog/ai-agent-total-cost-of-ownership) breakdown. The broader [AI agent cost management](https://devrev.ai/blog/ai-agent-cost-management) method covers unit-cost measurement and its operating cadence.

This memo uses those inputs to support a decision. It doesn’t need another cost taxonomy.

If no outcomes qualify, the unit cost is undefined. Report the spending and unfinished work rather than presenting zero as an attractive result.

## Does the capacity gain change cash spending?

Released capacity changes cash spending only when it changes an actual outlay against a credible comparison. Faster work can be valuable without meeting that condition.

An employee might use the released time to handle a backlog or improve service. Those are potential operating benefits, not evidence that their salary disappeared.

HM Treasury’s [Government Efficiency Framework](https://www.gov.uk/government/publications/the-government-efficiency-framework/the-government-efficiency-framework--2), updated in November 2025, distinguishes reduced expenditure from productivity gains without reduced spending. Its scope is UK central government, not corporate accounting law.

The distinction is still useful for your review. Ask your finance team to establish the company’s own realization rule rather than borrow a convenient label.

### A hypothetical memo that can survive its arithmetic

Every input below is assumed. This is a teaching example, not observed pilot data, a customer result, or a forecast for your workload.

Assume a monthly workflow contains 1,000 eligible cases and 800 accepted unique outcomes. The remaining 200 cases stay visible as rejected, unresolved, or escalated work.

Assume the cost categories include all work within that boundary. Retries consume resources but never increase the accepted-outcome count.

| Hypothetical monthly input | Assumed amount | Accounting treatment |
| --- | --- | --- |
| Runtime, infrastructure, and licenses | $2,400 | Includes successful, failed, and retried runs once |
| Routine verification | $2,000 | Human review, excluding the separate rework category |
| Incremental failure handling and rework | $1,200 | Additional work not counted elsewhere |
| Ongoing integration, security, and operations | $800 | Recurring support, separate from initial setup |
| Total recurring cost | $6,400 | Sum of the four non-overlapping categories |
| Accepted unique outcomes | 800 | Same workflow and monthly acceptance window |
| Recurring cost per accepted outcome | $8.00 | $6,400 divided by 800 |
| One-time setup cash | $9,600 | Excluded from the recurring monthly total |
| Realizable monthly cash benefit | $8,000 | Assumed avoided spending requiring independent evidence |

In short: the example delivers an $8 recurring unit cost, but its cash result depends on the separate $8,000 benefit assumption.

For this illustration, assume the recurring costs are incremental cash outlays. If your numbers contain unchanged salary allocations, separate the economic and cash views first.

The assumed benefit could represent a cancellable external-service expense for the same accepted work. No real contract cancellation is claimed here.

Under constant monthly costs, benefits, and volume:

| Hypothetical calculation | Result |
| --- | --- |
| Monthly net cash: $8,000 − $6,400 | $1,600 |
| Simple payback: $9,600 ÷ $1,600 | 6 months |
| First-year net cash after setup: 12 × $1,600 − $9,600 | $9,600 |
| First-year full cost per accepted outcome: ($9,600 + 12 × $6,400) ÷ (12 × 800) | $9.00 |

In short: including setup changes the unit-cost view, while cash payback requires a positive recurring cash benefit after costs.

These calculations omit ramp-up, discounting, tax, working-capital effects, and terminal value. They are not net present value or an internal rate of return.

The broader [AI ROI](https://devrev.ai/blog/ai-roi) discussion explains why benefits and investment scope must match. Here, the decisive question remains whether the assumed cash change can actually occur.

## Set the conditions that would change the decision

An investment memo should show what would overturn its recommendation. Otherwise, your review becomes a defense of the pilot rather than a decision about it.

Start with changes you can interpret separately. In the hypothetical, 640 accepted outcomes at the same $6,400 monthly cost produce a $10 unit cost.

That calculation says nothing about the resulting benefit. Fewer accepted outcomes may change avoided spending, service performance, or reviewer demand. Re-estimate those effects rather than holding them constant without explanation.

Now reduce the realizable cash benefit to $6,000 while leaving recurring cost unchanged. Monthly net cash becomes negative $400. The setup cost has no positive simple payback under those assumptions.

You don’t need a more elaborate spreadsheet to identify that problem. You need the person responsible for validating the benefit.

### Scale, extend, narrow, or stop

Separate the evidence from the decision it supports. Use the following memo as a working document, not a certification checklist.

| Memo field | What you need to record |
| --- | --- |
| Decision requested | The next commitment: expansion, bounded learning, narrower scope, or stop |
| Workflow boundary | Eligible work, exclusions, required quality, and the accountable operating owner |
| Evidence status | Observed, assumed, or missing for every input, with a source and date |
| Accepted work | Unique qualifying outcomes, review window, assisted work, and unresolved cases |
| Comparable baseline | What would happen without the investment under similar conditions |
| Cost and benefit | Recurring cost, setup, capacity value, and a separate incremental cash calculation |
| Downside | Workload changes, lower acceptance, increased oversight, price exposure, and exit costs |
| Controls | Data and action limits, incident ownership, review requirements, and recovery arrangements |
| Decision condition | The evidence or threshold required for the next commitment |
| Review responsibility | Finance and operating owners, the review date, and the action if conditions fail |

In short: make the assumption, its owner, and its consequence visible in the same document.

**Scale** when the evidence supports the expanded scope and the owners accept the operating conditions. Don’t assume the next workload resembles the carefully selected pilot.

**Extend** when a specific uncertainty can be resolved through a bounded experiment. Give that experiment a spending limit, a question, and an end date.

For example, you might test whether reviewer effort stays stable across a different request type. “Collect more data” is too vague to justify indefinite spending.

**Narrow** when value depends on a better-defined boundary. A workflow may work for current documentation but fail on disputed or incomplete records.

**Stop** when the evidence no longer supports the commitment, or when an unacceptable control failure requires intervention. A financial threshold cannot excuse an unauthorized action.

Record disagreement rather than forcing consensus into the model. If Finance challenges the avoided-hiring assumption, show the case without that benefit. If Operations challenges acceptance, show the disputed cases.

A supplier-exit concern belongs here too. Identify what you would need to move: approved source access, test cases, retained evidence, workflow ownership, and ongoing work.

Do not assign an invented dollar value to every risk. Missing information is a valid memo entry, provided it has an owner and changes the decision appropriately.

For the wider committee’s responsibilities, use the [enterprise AI agent buying criteria](https://devrev.ai/blog/enterprise-ai-agent-buying-criteria). Keep this memo focused on the commitment Finance is being asked to make now.

These questions separate a useful pilot result from a defensible investment claim. Apply the same scope and evidence rules to each answer.

A CFO guide to AI agents is useful only if it changes the next decision. It should let an owner say yes, not yet, or no.

Choose one pilot. Ask Operations to sign off on accepted work and Finance to sign off on the cash assumptions. Put unresolved conditions beside the requested commitment.

Then take that memo to the review. A forecast earns attention; a commitment needs evidence.

## FAQ

### How do you calculate ROI for an AI agent?

Any AI agent ROI figure starts with a defined investment period, comparable baseline, and evidenced benefits. Account for setup, operation, verification, and rework without counting a cost twice. Decide whether the calculation measures economic value or incremental cash. Only then calculate a return using consistent benefit and investment definitions that your finance team accepts.

### Is cost per accepted outcome the same as TCO?

No. Total cost of ownership describes expenditure across a defined scope and period. Cost per accepted outcome relates a chosen cost total to qualifying results from that same scope. It adds a delivery denominator, not evidence of positive returns. State whether setup costs are included before comparing two figures.

### Why don’t hours saved automatically become cash savings?

Saved hours can create capacity without reducing expenditure. Cash savings require evidence that spending changed against a credible baseline, such as a reduced service contract. Avoided hiring also needs a supported counterfactual. Keep service improvements and capacity benefits visible, but don’t describe them as cash Finance has already recovered.

### What if the agent produces no accepted outcomes?

Report the incurred cost and the attempted, rejected, or unfinished work. A cost-per-accepted-outcome calculation has no valid denominator when nothing qualifies; the result is not zero. Review whether a bounded experiment can resolve the failure, or apply the stop condition agreed before further spending on the same workflow.