ApexClaw
Home › Industries › BFSI
Industries

Model risk management governs what a model predicts. It does not govern what an agent does.

Banks already validate models. Agents break that frame: they act rather than score, hold tool permissions and wallets, and compose new plans continuously. The regimes below all still apply — they just were not written for something that acts on its own initiative.

Get an Agent Trust Gap Brief

The regimes an agent programme actually meets

RegimeWhere agents strain it
DORA
EU financial entities
ICT risk management, third-party register, resilience testing. An agent calling an external tool is ICT third-party dependency — most registers have not caught up.
OSFI Guideline E-23
Canada, federally regulated FIs
Model risk management, expanded to AI and ML. In force 1 May 2027. Third-party model accountability is explicit.
SR 11-7
US Federal Reserve
The model risk foundation. Agents strain it because the 'model' now plans, calls tools and acts — validation was designed for something that scores.
NYDFS Part 500
New York DFS
Cybersecurity, access controls, incident reporting. Non-human identity and revocation land squarely here.
MAS FEAT / Veritas
Singapore
Fairness, ethics, accountability, transparency. Singapore's Jan 2026 framework is reportedly the only one addressing autonomous agents directly.
APRA CPS 230 / 234
Australia
Operational risk and information security, including service-provider management.
PSD2 / SCA
EU payments
Strong customer authentication. Agent-initiated payments raise the question of who is authenticating, and AP2's mandates exist to answer it.
Basel operational resilience
Global
Critical operations must survive disruption. An agent that cannot be stopped is an operational-resilience finding, not just a security one.

Verified 2026-08-06. Interpretation for planning, not a compliance determination and not legal advice. Full crosswalk →

Why model risk management does not cover agents

Every bank already has an MRM function, and the reasonable first instinct is to route agents through it. That works for the model. It does not work for the agent.

A model scores. An agent acts.

Validation asks whether outputs are accurate and stable. It does not ask whether an action was authorized, because models did not take actions.

Inventory is per-model, not per-agent

The model inventory has one entry. The agent may hold twelve tool permissions, a wallet and a delegation from a named human. None of that appears.

Validation is periodic; agents act continuously

An annual validation cycle cannot govern something that composes a new plan every few seconds.

Explainability is not attribution

Knowing why a model produced a score is different from proving which agent acted, under whose authority, against which exact payload.

The practical answer is not to replace MRM. It is to add the layer it was never built to hold: identity per agent, a gate between decision and effect, and a receipt that survives an examiner asking who authorized this.

Where agents touch regulated processes

Credit and adverse action

If an agent influences a decline, someone must explain it and evidence the basis. Reconstructing that from logs after the fact is slow and rarely conclusive.

KYC, AML, sanctions

An agent that screens or clears alerts is inside a supervised control. Its actions need the same evidentiary standard as a human analyst's.

Payments and treasury

Agent-initiated payments raise authorization and SCA questions. A cap that refuses beats a throttle that slows — see wallet governance.

Trading and markets

Speed is the aggravating factor. Circuit breakers per effect class and measured kill-switch time-to-effect are the controls that matter.

Client communication

Anything an agent sends to a client is a supervised communication with retention and review obligations attached.

Regulatory reporting

If an agent touches a submission, the evidence chain has to survive an examiner working backwards from the filing.

What an examiner will ask

  1. Show me your agent inventory, and the named human accountable for each.
  2. Show me an action this agent took, and the authorization that permitted it.
  3. Show me an action it was refused. (Most estates cannot — refusals go unrecorded, so the record contains only successes.)
  4. Demonstrate revocation. How long did it take to take effect, measured?
  5. Reconstruct this specific incident end to end.
  6. Show me the spend ceiling, and what happens at it.
  7. Show me where your documentation disagrees with the running system.

Question three is the one that separates a governed estate from an instrumented one. Refusals are evidence. A system that records only what succeeded has no record.

Run the gap brief

Common questions

Does SR 11-7 apply to AI agents?

In supervised institutions, expect examiners to treat decision-making agents within model risk expectations: inventory, validation, ongoing monitoring.

What does DORA require of agent deployments?

ICT risk management, incident reporting within defined timelines, and oversight of third-party providers — which includes agent platforms touching important functions.

What is OSFI E-23's significance?

It applies from May 2027 to federally regulated Canadian institutions with a broad model definition and lifecycle expectations.

Are agents in scope of model inventory?

If they inform or take decisions, assume yes. Assuming exclusion because the system is called an agent is not a defensible position.

What evidence do examiners ask for?

Inventory, risk rating with reasoning, validation, monitoring, and the ability to explain a specific decision after the fact.