AI agents in financial services: 2026 evidence before action

Evaluate AI agents in financial services with a workflow dossier covering evidence, authority, approval boundaries, and release decisions.

Updated

10 min read

The reviewer has a proposed client message, but not yet a defensible reason to send it.

An onboarding agent has drafted a request to a new client, asking for a document it believes is missing. The draft cites an attachment already on file. Nothing has been sent. Before this becomes a customer contact, someone has to answer a plain question: does the evidence behind this action hold, and does the agent have the authority to take it?

That question is the whole subject of this page. Not whether AI agents in financial services are impressive, but whether one specific action can survive review.

As of September 2026, that’s the decision most regulated teams are actually stuck on.

TL;DR

  • Review the evidence and authority behind one specific action before you widen an agent’s scope. Reliable assistance is a valid endpoint on its own.
  • Institution policy and legal requirements are different sources of control. Label them separately, because a policy choice is not a legal mandate.
  • A stopped workflow can leave completed actions behind, so reconciliation needs a named owner, not just a stop button.

What are AI agents in financial services?

AI agents in financial services are software systems that use information and tools to assist with or perform tasks across financial workflows. They may retrieve internal guidance, prepare a client response, or initiate a configured action. The institution decides their permitted scope, data access, review conditions, and operating rules for each use case.

That definition is deliberately narrow on one point: scope is a design choice, set by the institution, not a fixed property of the technology. An agent that only reads a policy document and an agent that drafts a message to a client sit at very different risk levels. The work of a compliance-first design is deciding where on that range a given task belongs, and proving it can stay there.

The reviewer starts with the proposed action

Most writing about agents in banking starts with the category and its promise. A reviewer starts somewhere much smaller: one action, sitting in a queue, waiting for a decision. In our onboarding case, that action is a drafted email asking the client for an identity document the agent flagged as missing.

The useful discipline is to work backwards from the action to its justification. Claimed authority, then evidence, then control, then the release decision. If any link is weak, hold the action until someone can resolve the gap. This is the opposite of the demo instinct, which shows the finished output and asks you to trust the path behind it.

What authority is actually being claimed?

Separate two things that get blurred constantly: capability and permission. The agent can compose a message and can reach a send tool. Neither fact grants it the right to act. Authority means an identity, a permitted operation, a specific resource, and a condition under which the operation is allowed, all checked at the moment of execution.

In this case, proposed institution policy permits the agent to analyze the client file and draft a response, but requires a human reviewer to confirm before anything is sent. That is a control the institution chose. It is not a universal legal rule about onboarding, and it should not be written as one.

Read-only retrieval isn’t exempt either. Reading a client record still creates access and disclosure exposure, so the same authority question applies before a single field is retrieved. To keep this separation clean, tie the acting identity to a scoped grant rather than a general account, the way non-human identity frames agent principals.

Which evidence supports this client request?

The draft rests on a claim: a required document is missing. That claim is only reviewable if the agent can show its working. Not hidden reasoning, but a defined evidence packet: the client identifier, the source system it checked, the document version and its retrieval time, the effective policy date that made the document required, the proposed message text, and the specific discrepancy it found.

Assemble those fields and a reviewer can check the real question. Is this the right client? Is this the current policy? Was the document genuinely absent, or attached under a name the agent didn’t match? Record the decisions and actions taken, apply the institution’s retention and access rules, and keep sensitive content to the minimum the review needs.

If the evidence is incomplete, the correct output isn’t a softer draft. It’s a hold, or a narrower task. Grounding this retrieval in a permission-aware context layer is the job AI knowledge management describes, and it’s why evidence quality is a knowledge problem before it is a model problem.

The workflow control dossier

An onboarding case gives us a method, not a universal rule. The same review questions can travel to other regulated workflows without flattening their different obligations. The table below applies the method across six workflows. Read the middle columns as proposed institution controls, which each institution sets for itself, and read the release owner as the accountable human, not an automated gate.

WorkflowProposed action and evidenceProposed institution controlRelease owner and boundary
KYC and AML supportAssemble identity discrepancies with source records and match rationaleHold uncertain matches for specialist reviewAML operations; no automated customer classification assumed
Fraud investigation supportSummarize linked alerts, timestamps, and conflicting evidenceRestrict disclosure; separate the summary from any account restrictionFraud lead; freeze authority is specified separately
Client serviceDraft a fee explanation using effective terms and verified account contextRoute unsupported exceptions and hardship requests to a personService owner; no universal human-transfer rule is claimed
Client onboardingDraft a missing-document request tied to the correct client and policy versionReviewer confirms recipient, evidence, and approved contentOnboarding lead; completion is not account approval
Regulatory reporting preparationAssemble draft fields with lineage, calculation version, and missing inputsKeep preparation separate from filing authorizationReporting owner; applicable filing duties require scoped review
Internal knowledgeRetrieve current policy with source permissions and effective dateEnforce access; mark uncertain or superseded guidanceKnowledge owner; read-only is not risk-free

The pattern the table makes visible is simple. The workflow determines the evidence to gather and the boundary where a human decides, not a single approval checkbox reused everywhere. An onboarding draft, a fraud summary, and a draft regulatory filing carry different exposures and different owners. A generic approval step flattens that difference and hides the risk it was supposed to catch.

What these sources do not regulate

This is the part that most category pages get wrong, and where a financial-services page earns its legal review. Naming a regulator is not the same as citing a rule that applies to your agent. Consider the actual scope of these examples.

The Federal Reserve’s revised model risk guidance, issued with SR 26-2 on April 17, 2026, supersedes SR 11-7 and SR 21-8. Its own footnote 3 states that generative and agentic AI models “are not within the scope of this guidance.”

So don’t apply that document’s validation regime to an agent automatically. Nor should the exclusion be read as a statement that these systems are unregulated. The same guidance says a banking organization’s own risk-management practices should determine appropriate controls for tools it does not cover. It is guidance, not prescriptive agent law, and it is described as most relevant to Fed-regulated banks with more than $30 billion in total assets.

On adverse-action notices, be equally careful. CFPB Circulars 2022-03 and 2023-03 were withdrawn on May 12, 2025, so they are not current guidance. The current authority is Regulation B, 12 CFR 1002.9, which governs adverse-action notifications and reason-giving for covered credit decisions. It requires a specific statement of principal reasons; a generic reference to internal policy or a failing score is insufficient.

That is a narrow US credit rule, with its own procedures and exceptions. It is not a universal requirement for a human signature on every service response, and it should never be stretched into one.

On politically exposed persons, an August 21, 2020 US interagency statement addressed risk-based due diligence. It noted that PEP status alone does not trigger special additional steps under the rule it discussed. That is a historical, US, bank-focused statement. It is not a current global AML survey, and it does not create an automatic human-escalation rule for every PEP match. When your page needs global AI-governance framing, defer it to EU AI Act compliance rather than treating EU rules as universal financial-services law.

Approval is not the end of the control

A green checkmark from a reviewer feels like the finish line. It is closer to the midpoint. The action must still execute under the approved conditions, and the workflow must remain intelligible if only part of it completes.

Check authority again at execution

Approval records a decision inside a workflow. It doesn’t, by itself, hand the agent access to a tool, an account, or a data source. Between approval and execution, records can change and conditions can drift. So the system should recheck the acting identity, the permitted operation, the current recipient, and the payload before it acts, including any limit on the approver’s own authority.

If the underlying record changed materially since review, that isn’t a small variance to wave through. It reopens the review.

Reconcile partial actions before retrying

Now the failure that matters. The approved message sends, but the follow-on step that updates the client’s case status fails. A naive retry re-sends the message. The correct response is to stop new actions, inspect what completed, and assign repair before anything runs again.

A stop control prevents further activity, but it can’t unsend a message or reverse every completed transaction. A partially completed workflow therefore needs someone to own the cleanup: identify completed steps, prevent duplicate effects on retry, and decide what the client is told.

If recovery is possible, distinguish an authorized compensating action from manual repair. Designing that boundary is exactly what AI agent guardrails enforce at runtime, and it’s where transaction-level examples belong, kept secondary to the control itself.

Expand only after the evidence holds

The adoption sequence follows from everything above. Start where the agent retrieves and drafts. Test it against reviewed cases. Only then authorize a bounded action, and only after the evidence and approval design have actually passed. Assistance is a legitimate place to stop. There’s no dollar tier that forces a team to hand over more autonomy than its evidence supports.

Good tests are the ones designed to fail loudly. They reject evidence tied to the wrong client, a stale policy version, and an unauthorized recipient. They confirm that a changed record reopens review, and that a partial execution creates visible repair work rather than a silent retry.

If you measure anything, measure four things: evidence completeness, how often reviewers disagree with the agent, unauthorized action attempts, and unresolved partial states, each with a defined denominator and window. Promise the count, not a percentage you haven’t earned.

This is a reasonable point to introduce Computer, by DevRev, and only here, after the evidence problem is on the table. Computer is an AI resolution and intelligence-and-action platform that works alongside existing systems; it is not a CRM and not a replacement for the specialists who own these decisions. Where it is relevant, Computer Memory can hold permission-respecting, connected context for a workflow, subject to the access rules the institution already enforces.

DevRev positions Computer Agent Studio as a way to configure the task and its review points. The dossier’s practical outcome is modest: show the reviewer the right client record, current policy version, proposed message, and permitted next step together. Product and legal owners must confirm this language and its application to any real financial-services deployment before publication.

Your next release decision should fit inside one dossier someone is willing to own. Take the onboarding action in front of you, write down the evidence that would let it survive review, and name the release owner who accepts or rejects that scope. If you can’t fill in all three, you’ve found your real starting point.

Frequently Asked Questions

DEVREV

See Computer work for you

Your AI teammate that finds answers, takes action, and gets work done across every tool.