Model risk management governs what a model predicts. It does not govern what an agent does.
Banks already validate models. Agents break that frame: they act rather than score, hold tool permissions and wallets, and compose new plans continuously. The regimes below all still apply — they just were not written for something that acts on its own initiative.
Get an Agent Trust Gap BriefThe regimes an agent programme actually meets
| Regime | Where agents strain it |
|---|---|
| DORA EU financial entities | ICT risk management, third-party register, resilience testing. An agent calling an external tool is ICT third-party dependency — most registers have not caught up. |
| OSFI Guideline E-23 Canada, federally regulated FIs | Model risk management, expanded to AI and ML. In force 1 May 2027. Third-party model accountability is explicit. |
| SR 11-7 US Federal Reserve | The model risk foundation. Agents strain it because the 'model' now plans, calls tools and acts — validation was designed for something that scores. |
| NYDFS Part 500 New York DFS | Cybersecurity, access controls, incident reporting. Non-human identity and revocation land squarely here. |
| MAS FEAT / Veritas Singapore | Fairness, ethics, accountability, transparency. Singapore's Jan 2026 framework is reportedly the only one addressing autonomous agents directly. |
| APRA CPS 230 / 234 Australia | Operational risk and information security, including service-provider management. |
| PSD2 / SCA EU payments | Strong customer authentication. Agent-initiated payments raise the question of who is authenticating, and AP2's mandates exist to answer it. |
| Basel operational resilience Global | Critical operations must survive disruption. An agent that cannot be stopped is an operational-resilience finding, not just a security one. |
Verified 2026-08-06. Interpretation for planning, not a compliance determination and not legal advice. Full crosswalk →
Why model risk management does not cover agents
Every bank already has an MRM function, and the reasonable first instinct is to route agents through it. That works for the model. It does not work for the agent.
A model scores. An agent acts.
Validation asks whether outputs are accurate and stable. It does not ask whether an action was authorized, because models did not take actions.
Inventory is per-model, not per-agent
The model inventory has one entry. The agent may hold twelve tool permissions, a wallet and a delegation from a named human. None of that appears.
Validation is periodic; agents act continuously
An annual validation cycle cannot govern something that composes a new plan every few seconds.
Explainability is not attribution
Knowing why a model produced a score is different from proving which agent acted, under whose authority, against which exact payload.
The practical answer is not to replace MRM. It is to add the layer it was never built to hold: identity per agent, a gate between decision and effect, and a receipt that survives an examiner asking who authorized this.
Where agents touch regulated processes
Credit and adverse action
If an agent influences a decline, someone must explain it and evidence the basis. Reconstructing that from logs after the fact is slow and rarely conclusive.
KYC, AML, sanctions
An agent that screens or clears alerts is inside a supervised control. Its actions need the same evidentiary standard as a human analyst's.
Payments and treasury
Agent-initiated payments raise authorization and SCA questions. A cap that refuses beats a throttle that slows — see wallet governance.
Trading and markets
Speed is the aggravating factor. Circuit breakers per effect class and measured kill-switch time-to-effect are the controls that matter.
Client communication
Anything an agent sends to a client is a supervised communication with retention and review obligations attached.
Regulatory reporting
If an agent touches a submission, the evidence chain has to survive an examiner working backwards from the filing.
What an examiner will ask
- Show me your agent inventory, and the named human accountable for each.
- Show me an action this agent took, and the authorization that permitted it.
- Show me an action it was refused. (Most estates cannot — refusals go unrecorded, so the record contains only successes.)
- Demonstrate revocation. How long did it take to take effect, measured?
- Reconstruct this specific incident end to end.
- Show me the spend ceiling, and what happens at it.
- Show me where your documentation disagrees with the running system.
Question three is the one that separates a governed estate from an instrumented one. Refusals are evidence. A system that records only what succeeded has no record.
Run the gap briefCommon questions
Does SR 11-7 apply to AI agents?
In supervised institutions, expect examiners to treat decision-making agents within model risk expectations: inventory, validation, ongoing monitoring.
What does DORA require of agent deployments?
ICT risk management, incident reporting within defined timelines, and oversight of third-party providers — which includes agent platforms touching important functions.
What is OSFI E-23's significance?
It applies from May 2027 to federally regulated Canadian institutions with a broad model definition and lifecycle expectations.
Are agents in scope of model inventory?
If they inform or take decisions, assume yes. Assuming exclusion because the system is called an agent is not a defensible position.
What evidence do examiners ask for?
Inventory, risk rating with reasoning, validation, monitoring, and the ability to explain a specific decision after the fact.