---
Title: "AI agents in financial services: 2026 evidence before action"
Url: "https://devrev.ai/blog/ai-agents-financial-services"
Published: "2026-09-23"
Last Updated: "2026-09-23"
Author: "Nivedita Bharathi"
Category: "Blog, Computer"
Excerpt: "Evaluate AI agents in financial services with a workflow dossier covering evidence, authority, approval boundaries, and release decisions."
Reading Time: 10
---

# AI agents in financial services: 2026 evidence before action

The reviewer has a proposed client message, but not yet a defensible reason to send it.

An onboarding agent has drafted a request to a new client, asking for a document it believes is missing. The draft cites an attachment already on file. Nothing has been sent. Before this becomes a customer contact, someone has to answer a plain question: does the evidence behind this action hold, and does the agent have the authority to take it?

That question is the whole subject of this page. Not whether AI agents in financial services are impressive, but whether one specific action can survive review. 

As of September 2026, that’s the decision most regulated teams are actually stuck on.

## TL;DR

- Review the evidence and authority behind one specific action before you widen an agent’s scope. Reliable assistance is a valid endpoint on its own.
- Institution policy and legal requirements are different sources of control. Label them separately, because a policy choice is not a legal mandate.
- A stopped workflow can leave completed actions behind, so reconciliation needs a named owner, not just a stop button.

## What are AI agents in financial services?

AI agents in financial services are software systems that use information and tools to assist with or perform tasks across financial workflows. They may retrieve internal guidance, prepare a client response, or initiate a configured action. The institution decides their permitted scope, data access, review conditions, and operating rules for each use case.

That definition is deliberately narrow on one point: scope is a design choice, set by the institution, not a fixed property of the technology. An agent that only reads a policy document and an agent that drafts a message to a client sit at very different risk levels. The work of a compliance-first design is deciding where on that range a given task belongs, and proving it can stay there.

## The reviewer starts with the proposed action

Most writing about agents in banking starts with the category and its promise. A reviewer starts somewhere much smaller: one action, sitting in a queue, waiting for a decision. In our onboarding case, that action is a drafted email asking the client for an identity document the agent flagged as missing.

The useful discipline is to work backwards from the action to its justification. Claimed authority, then evidence, then control, then the release decision. If any link is weak, hold the action until someone can resolve the gap. This is the opposite of the demo instinct, which shows the finished output and asks you to trust the path behind it.

### What authority is actually being claimed?

Separate two things that get blurred constantly: capability and permission. The agent can compose a message and can reach a send tool. Neither fact grants it the right to act. Authority means an identity, a permitted operation, a specific resource, and a condition under which the operation is allowed, all checked at the moment of execution.

In this case, proposed institution policy permits the agent to analyze the client file and draft a response, but requires a human reviewer to confirm before anything is sent. That is a control the institution chose. It is not a universal legal rule about onboarding, and it should not be written as one.

Read-only retrieval isn’t exempt either. Reading a client record still creates access and disclosure exposure, so the same authority question applies before a single field is retrieved. To keep this separation clean, tie the acting identity to a scoped grant rather than a general account, the way [non-human identity](https://devrev.ai/blog/non-human-identity) frames agent principals.

### Which evidence supports this client request?

The draft rests on a claim: a required document is missing. That claim is only reviewable if the agent can show its working. Not hidden reasoning, but a defined evidence packet: the client identifier, the source system it checked, the document version and its retrieval time, the effective policy date that made the document required, the proposed message text, and the specific discrepancy it found.

Assemble those fields and a reviewer can check the real question. Is this the right client? Is this the current policy? Was the document genuinely absent, or attached under a name the agent didn’t match? Record the decisions and actions taken, apply the institution’s retention and access rules, and keep sensitive content to the minimum the review needs.

If the evidence is incomplete, the correct output isn’t a softer draft. It’s a hold, or a narrower task. Grounding this retrieval in a permission-aware context layer is the job [AI knowledge management](https://devrev.ai/blog/ai-knowledge-management) describes, and it’s why evidence quality is a knowledge problem before it is a model problem.

## The workflow control dossier

An onboarding case gives us a method, not a universal rule. The same review questions can travel to other regulated workflows without flattening their different obligations. The table below applies the method across six workflows. Read the middle columns as proposed institution controls, which each institution sets for itself, and read the release owner as the accountable human, not an automated gate.

| Workflow | Proposed action and evidence | Proposed institution control | Release owner and boundary |
| --- | --- | --- | --- |
| KYC and AML support | Assemble identity discrepancies with source records and match rationale | Hold uncertain matches for specialist review | AML operations; no automated customer classification assumed |
| Fraud investigation support | Summarize linked alerts, timestamps, and conflicting evidence | Restrict disclosure; separate the summary from any account restriction | Fraud lead; freeze authority is specified separately |
| Client service | Draft a fee explanation using effective terms and verified account context | Route unsupported exceptions and hardship requests to a person | Service owner; no universal human-transfer rule is claimed |
| Client onboarding | Draft a missing-document request tied to the correct client and policy version | Reviewer confirms recipient, evidence, and approved content | Onboarding lead; completion is not account approval |
| Regulatory reporting preparation | Assemble draft fields with lineage, calculation version, and missing inputs | Keep preparation separate from filing authorization | Reporting owner; applicable filing duties require scoped review |
| Internal knowledge | Retrieve current policy with source permissions and effective date | Enforce access; mark uncertain or superseded guidance | Knowledge owner; read-only is not risk-free |

The pattern the table makes visible is simple. The workflow determines the evidence to gather and the boundary where a human decides, not a single approval checkbox reused everywhere. An onboarding draft, a fraud summary, and a draft regulatory filing carry different exposures and different owners. A generic approval step flattens that difference and hides the risk it was supposed to catch.

### What these sources do not regulate

This is the part that most category pages get wrong, and where a financial-services page earns its legal review. Naming a regulator is not the same as citing a rule that applies to your agent. Consider the actual scope of these examples.

The Federal Reserve’s revised model risk guidance, issued with [SR 26-2](https://www.federalreserve.gov/supervisionreg/srletters/SR2602.htm) on April 17, 2026, supersedes SR 11-7 and SR 21-8. Its own [footnote 3](https://www.federalreserve.gov/frrs/guidance/supervisory-guidance-on-model-risk-management.htm) states that generative and agentic AI models “are not within the scope of this guidance.”

So don’t apply that document’s validation regime to an agent automatically. Nor should the exclusion be read as a statement that these systems are unregulated. The same guidance says a banking organization’s own risk-management practices should determine appropriate controls for tools it does not cover. It is guidance, not prescriptive agent law, and it is described as most relevant to Fed-regulated banks with more than $30 billion in total assets.

On adverse-action notices, be equally careful. CFPB [Circulars 2022-03 and 2023-03](https://www.consumerfinance.gov/compliance/guidance/withdrawn-guidance/) were withdrawn on May 12, 2025, so they are not current guidance. The current authority is [Regulation B, 12 CFR 1002.9](https://www.consumerfinance.gov/rules-policy/regulations/1002/9/), which governs adverse-action notifications and reason-giving for covered credit decisions. It requires a specific statement of principal reasons; a generic reference to internal policy or a failing score is insufficient.

That is a narrow US credit rule, with its own procedures and exceptions. It is not a universal requirement for a human signature on every service response, and it should never be stretched into one.

On politically exposed persons, an [August 21, 2020 US interagency statement](https://www.fincen.gov/news/news-releases/agencies-issue-statement-bank-secrecy-act-due-diligence-requirements-customers) addressed risk-based due diligence. It noted that PEP status alone does not trigger special additional steps under the rule it discussed. That is a historical, US, bank-focused statement. It is not a current global AML survey, and it does not create an automatic human-escalation rule for every PEP match. When your page needs global AI-governance framing, defer it to [EU AI Act compliance](https://devrev.ai/blog/eu-ai-act-compliance) rather than treating EU rules as universal financial-services law.

## Approval is not the end of the control

A green checkmark from a reviewer feels like the finish line. It is closer to the midpoint. The action must still execute under the approved conditions, and the workflow must remain intelligible if only part of it completes.

### Check authority again at execution

Approval records a decision inside a workflow. It doesn’t, by itself, hand the agent access to a tool, an account, or a data source. Between approval and execution, records can change and conditions can drift. So the system should recheck the acting identity, the permitted operation, the current recipient, and the payload before it acts, including any limit on the approver’s own authority.

If the underlying record changed materially since review, that isn’t a small variance to wave through. It reopens the review.

### Reconcile partial actions before retrying

Now the failure that matters. The approved message sends, but the follow-on step that updates the client’s case status fails. A naive retry re-sends the message. The correct response is to stop new actions, inspect what completed, and assign repair before anything runs again.

A stop control prevents further activity, but it can’t unsend a message or reverse every completed transaction. A partially completed workflow therefore needs someone to own the cleanup: identify completed steps, prevent duplicate effects on retry, and decide what the client is told. 

If recovery is possible, distinguish an authorized compensating action from manual repair. Designing that boundary is exactly what [AI agent guardrails](https://devrev.ai/blog/ai-agent-guardrails) enforce at runtime, and it’s where transaction-level examples belong, kept secondary to the control itself.

## Expand only after the evidence holds

The adoption sequence follows from everything above. Start where the agent retrieves and drafts. Test it against reviewed cases. Only then authorize a bounded action, and only after the evidence and approval design have actually passed. Assistance is a legitimate place to stop. There’s no dollar tier that forces a team to hand over more autonomy than its evidence supports.

Good tests are the ones designed to fail loudly. They reject evidence tied to the wrong client, a stale policy version, and an unauthorized recipient. They confirm that a changed record reopens review, and that a partial execution creates visible repair work rather than a silent retry.

If you measure anything, measure four things: evidence completeness, how often reviewers disagree with the agent, unauthorized action attempts, and unresolved partial states, each with a defined denominator and window. Promise the count, not a percentage you haven’t earned.

This is a reasonable point to introduce **Computer, by DevRev**, and only here, after the evidence problem is on the table. Computer is an AI resolution and intelligence-and-action platform that works alongside existing systems; it is not a CRM and not a replacement for the specialists who own these decisions. Where it is relevant, Computer Memory can hold permission-respecting, connected context for a workflow, subject to the access rules the institution already enforces.

DevRev positions Computer Agent Studio as a way to configure the task and its review points. The dossier’s practical outcome is modest: show the reviewer the right client record, current policy version, proposed message, and permitted next step together. Product and legal owners must confirm this language and its application to any real financial-services deployment before publication.

Your next release decision should fit inside one dossier someone is willing to own. Take the onboarding action in front of you, write down the evidence that would let it survive review, and name the release owner who accepts or rejects that scope. If you can’t fill in all three, you’ve found your real starting point.

## FAQ

### Does human approval give an agent permission to act?

Human approval and technical authorization are different controls. An approval records a decision within an institution’s workflow; it does not grant access to another account, tool, or data source. The system should check the acting identity, the allowed operation, and current conditions again before execution, including any limits on the approver’s own authority.

### Does SR 26-2 directly govern generative and agentic AI models?

No. The Federal Reserve’s revised model risk guidance, issued with SR 26-2 in April 2026, expressly excludes generative and agentic AI models from its scope. That exclusion does not mean those systems are unregulated. Institutions still need to determine the applicable legal duties and appropriate governance for their specific activities, systems, and risks.

### Are read-only financial workflows automatically safe?

No. Reading client records can still expose confidential information or disclose it to an unauthorized recipient. A read-only agent needs appropriately scoped access, source controls, and a defined handling policy. Review requirements depend on the institution’s workflow and applicable obligations, not simply on whether the agent can write to a system.

### What happens if an approved workflow completes only partly?

A partially completed workflow needs reconciliation before a retry. The institution should identify which actions succeeded, which failed, and whether repeating a request could duplicate an effect. A stop control prevents further activity but does not undo every completed action. Assign an owner for repair, authorized compensation, and customer communication.