When US banking regulators revised their interagency model risk management guidance in April 2026, the most important sentence was an exclusion: generative and agentic AI systems are outside its scope. The Federal Reserve's Vice Chair for Supervision explained the logic plainly — rapidly evolving technologies may require a different approach, and stretching a framework built for credit-scoring regressions over a large language model helps no one. The candour is welcome. The consequence is a governance vacuum, formally acknowledged, at exactly the moment adoption is accelerating.
What should a bank or insurer do while the technology-specific framework is drafted? The wrong answer is nothing. The institutions that thrive in guidance gaps are the ones that build their own interim standard, document it, and are ready to defend it — because when formal guidance lands, it invariably resembles what the best-governed firms were already doing.
Why the old framework genuinely does not fit Classical model risk management assumes a model with defined inputs, a stable estimation methodology, and testable outputs. Generative systems violate every assumption: the input space is unbounded natural language, behaviour shifts with every provider update, outputs are non-deterministic, and "validation" in the classical sense — backtesting against outcomes — has no direct analogue for a system that drafts emails, summarizes documents, or plans multi-step actions. Agentic systems add a second-order problem: the unit of risk is no longer a prediction but a chain of actions, each conditioned on the last.
That does not mean nothing transfers. The three pillars of classical MRM — effective challenge, independent validation, and lifecycle documentation — translate directly; only their techniques change. Effective challenge becomes adversarial testing and red-teaming. Validation becomes evaluation suites run before deployment and continuously after, tracking behaviour drift across provider updates. Documentation becomes system cards: intended use, prohibited use, known failure modes, escalation paths.
An interim standard in five commitments A defensible interim framework fits on one page. Every generative or agentic system is inventoried and risk-tiered by the harm its worst plausible output could cause. Every system above the lowest tier has a named accountable owner, documented behavioural limits, and human checkpoints at consequential actions. Every deployment is preceded by scenario-based adversarial testing, repeated on a schedule and after every material provider update. Kill-switch procedures exist and have been exercised. And incidents — including near-misses — are logged and reviewed with the same discipline as operational-risk events.
The Financial Stability Board's consultation on AI sound practices and the US regulators' planned request for information will eventually harden into expectations. Institutions running the interim playbook will recognize most of what arrives. The rest will discover that "the guidance didn't cover it" has never once worked as an examination defence.
Related reading - [Every bank exam is now an AI exam](/insights/ai-bank-examinations-occ-fed-jonas-osman-abdelghafour) - [The EU AI Act's August 2026 milestone](/insights/eu-ai-act-august-2026-banks-insurers-jonas-osman-abdelghafour) - [When AI governance meets cyber defense](/insights/ai-cybersecurity-nist-profile-banking-jonas-osman-abdelghafour)
See also Model Risk, Banking Risk and Board risk governance.
About the author Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.
Frequently asked questions
Why the old framework genuinely does not fit?
Classical model risk management assumes a model with defined inputs, a stable estimation methodology, and testable outputs. Generative systems violate every assumption: the input space is unbounded natural language, behaviour shifts with every provider update, outputs are non-deterministic, and "validation" in the classical sense — backtesting against outcomes — has no direct analogue for a system that drafts emails, summarizes documents, or plans multi-step actions. Agentic systems add a sec...
What should risk leaders know about an interim standard in five commitments?
A defensible interim framework fits on one page. Every generative or agentic system is inventoried and risk-tiered by the harm its worst plausible output could cause. Every system above the lowest tier has a named accountable owner, documented behavioural limits, and human checkpoints at consequential actions. Every deployment is preceded by scenario-based adversarial testing, repeated on a schedule and after every material provider update. Kill-switch procedures exist and have been exercised...
What should risk leaders know about related reading?
See also [Model Risk](/expertise/model-risk), [Banking Risk](/expertise/banking-risk) and [Board risk governance](/governance).
What should risk leaders know about about the author?
Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.