Your model-risk framework validated your credit scorecards for years, and it passed every examination. It has a hole in it, and generative AI is walking straight through it. That framework — the discipline built on the supervisory letter known as SR 11-7 — assumed a model that sits still: stable, inspectable, backtestable, and owned by the bank. A third-party foundation model is none of those.

This article lays out the replacement: The GenAI Assurance Grid, five components that govern generative and agentic AI by assuring the system around the model rather than validating a core you are not allowed to open. It reads as a build order with one control mechanic running through the middle — AI proposes, the orchestration layer disposes — and it carries one diagnostic, the Assurance-Gap Read, that names which of the five rings is currently the weakest for the system in front of you.

Why this matters now

Three calendars moved at once, and together they end the option to wait.

In the United States, the Federal Reserve, the Office of the Comptroller of the Currency, and the Federal Deposit Insurance Corporation issued SR 26-2 on the seventeenth of April, twenty twenty-six, superseding the fifteen-year-old SR 11-7 and SR 21-8 and moving to materiality-based oversight — most directly for Fed-regulated banks above roughly thirty billion dollars in assets. Its attachment states plainly that generative and agentic artificial intelligence are not within the scope of the guidance, while making equally clear that institutions must still apply risk-management and governance practices to what it does not cover. Out of scope of this guidance is not out of governance.

In the European Union, Article 50 of the EU AI Act begins to apply on the second of August, twenty twenty-six, adding an upstream transparency duty: machine-readable marking of synthetic content, and disclosure that a customer is dealing with an AI system — scoped by use case rather than a blanket watermark on every interaction.

In Vietnam, the AI Law 134/2025 and the Personal Data Protection Law 91/2025 are in force in twenty twenty-six, and the State Bank of Vietnam has a draft circular on artificial intelligence in banking out for consultation. Treat the finance-sector grace window and any penalty figure as reported and pending primary-source confirmation until the underlying texts clear — the in-force dates alone already make this a dated obligation rather than a whitepaper topic.

Read together, the three calendars say one thing: the regulator did not deregulate generative AI. It told you the old rulebook does not cover this kind of model, and handed the assurance problem back to you.

The problem: a framework built for one kind of model meets another entirely

Classical model-risk management, built on the SR 11-7 lineage, makes four quiet assumptions. A model is stable — a fixed specification you can pin down. It is inspectable — you own the code, the weights, the training data. It is backtestable — you can score its predictions against history. And it is passive — it outputs a number, and a human or a hard-coded rule takes the action.

A generative model, reached through a third-party interface, breaks all four, and each break costs you a specific control. It is non-deterministic and prompt-sensitive, so there is no single ground truth to validate against. It is closed, so there are no weights or training data to examine. It drifts on silent vendor updates, so a workflow that passed on Monday can fail by Friday. And as an agent it stops being passive — it plans, it calls tools, it changes state — so you have no identity or entitlement control over what is effectively a synthetic employee.

The classical toolkit therefore does not fail because your validators are weak. It fails because there is no stable specification to pin down, nothing to open, and no clean history to test against. Conceptual soundness, backtesting, and outcomes analysis do not map cleanly onto this kind of model — each has to be supplemented by system-level evaluation, not discarded.

Most banks deploy generative AI faster than they can assure it, then reach for the model-risk playbook they already own — exactly the wrong tool for a model that is closed, drifting, and not theirs. A quieter failure sits underneath. Much of this AI never reaches the risk team at all: it arrives embedded inside vendor software as a helpful feature and bypasses model-risk intake entirely. What is not inventoried cannot be tiered, contained, or challenged.

The framework: assure the system, not the weights

Here is the reframe, in one sentence.

You cannot govern generative AI by validating the model alone; you govern it by assuring the system around it.

Stop staring into the sealed core you are not allowed to open, and wrap control around it instead — control you do own. What data the system can reach. What tools it may call. What it is allowed to do before a human is in the loop. Whether you would notice if it silently got worse. And who answers when a vendor's model fails your customer. Assure the system, not the weights.

The strongest objection to this framework is a good one, and it deserves an honest answer. A generative model reached through an interface is really a non-deterministic piece of software — better governed through information-technology risk, cybersecurity, and third-party risk than through the model-risk lineage. That objection is largely correct, and SR 26-2 implicitly concedes it by carving generative AI out of model-risk scope. It does not weaken the grid; it tells you where the grid's centre of gravity has to sit — the orchestration layer and the agent's entitlements, not the model's math. Much of what follows is IT, cyber, and vendor-management discipline wearing a model-risk hat. That is the point, not a flaw.

The mechanic that holds the whole structure together is a separation, not a technology.

AI proposes the action the orchestration layer disposes of it = a decision you can defend

The model produces a proposal; a separate control plane decides whether that action may reach the customer or the ledger. Intelligence and authority, kept apart on purpose.

The five components — a build order with a return arc

Read the grid as five control rings around a sealed core, not as a checklist to complete left to right. Each ring has one job, one failure mode, and one diagnostic question that tells you whether it holds.

COMPONENT 1

Inventory — the AI System Registry

What is not inventoried cannot be governed. One register of every AI system — built, bought, or embedded inside a vendor feature you already pay for — fed by network and data-loss-prevention discovery and application-programming-interface monitoring rather than by a survey nobody answers, and kept separate from the quantitative-model inventory SR 26-2 still governs. No generative tool enters the bank without an intake record; shadow use surfaces from monitoring, not from good intentions. This is largely an IT-procurement and third-party-risk discipline, which is exactly why it is usually the ring that is missing. Diagnostic: could you list every generative and agentic system running in the bank today, including the ones buried inside vendor features?

COMPONENT 2

Tier — materiality, autonomy, and entitlements

Tier on entitlements, not on intention. Classify every registered system on three axes at once: how material its decisions are, how autonomous it is, and — critically — what tools and transaction limits it has actually been granted. The tier then sets validation depth, challenge intensity, monitoring cadence, and oversight mode. Tiering gives false comfort when it is set once at intake and never revisited: a low-risk summarizer becomes a high-risk actor the moment it is handed a payment permission. When a system gains a tool, re-tier it before it acts — the autonomy is the risk, not the label the use case started with. Diagnostic: have you tiered this system by what it can decide and which tools it may call, not just by what it was built to do?

COMPONENT 3

Orchestrate — the control plane

The load-bearing ring. Put a control layer between the model and the systems of record: input and output guardrails, retrieval entitlements, role-based tool permissions, transaction limits, human-in-the-loop for material actions, a kill-switch and rollback, and an audit-ready record of every prompt, tool call, decision, approval, and outcome. Functionally this is your new model-risk environment for a non-deterministic system — coded into architecture, not written into a policy memo. It fails in both directions: skip the ring and an autonomous agent reaches the ledger with no working brake; over-engineer it and latency pushes users back toward unsanctioned shadow tools. Diagnostic: when the model acts, does a control layer decide whether the action is allowed, or does it reach the customer and the ledger directly?

COMPONENT 4

Evaluate — continuous, not point-in-time

Test the pipeline, not the bare model. Replace the once-a-year validation report with continuous, system-level evaluation: groundedness and faithfulness, context recall, answer relevancy, prompt-robustness, red-teaming for jailbreak and prompt injection, and drift monitoring for silent vendor updates. When a vendor updates the model, re-evaluation should trigger automatically — the unit under test is the pipeline of retrieval, prompts, and guardrails, not the bare foundation model in isolation. Without a counterfactual baseline you cannot tell whether a system that passed on Monday has quietly got worse by Friday. The tooling here — retrieval-evaluation and model-as-a-judge libraries such as Ragas and TruLens — is emerging practice, useful today, not yet a settled industry standard. Diagnostic: do you test the system continuously in production, or did you validate it once and file the report?

COMPONENT 5

Challenge vendors — accountability without access

You answer for a model you cannot open. The bank stays accountable, so effective challenge has to be redefined: adversarial testing of the bank's own system boundary — retrieval, prompts, guardrails, jailbreak resilience — plus disciplined reliance on vendor model and system cards, contractual audit rights and change-notification clauses, and a concentration register tracking how much of the bank now rests on one frontier provider. Rubber-stamping a vendor's own attestation is not assurance; it is concentration risk wearing the costume of assurance. Diagnostic: if a closed vendor model failed your customer, could you show independent challenge, or only the vendor's own paperwork?

The return arc matters as much as the build order. What Evaluate learns in production changes the tier; what the tier says changes the entitlements the orchestration layer enforces; what vendor challenge uncovers changes what you are willing to register at all. A grid read strictly left to right, once, is a checklist — and a checklist is what the old framework already was.

The Assurance-Gap Read: which ring binds, by use case

The grid is not a checklist to complete all at once. For any one system a single ring is the weakest, and which one it is changes by use case. The tier decides how hard each ring has to hold.

TierTrigger — materiality × autonomy × entitlementsEvaluate depthChallenge / approvalOrchestrate / monitorOversight mode
ProhibitedManipulative, deceptive, or social-scoring uses barred by lawReject at intakeBanned
HighMaterial decision — credit underwriting or AML — and agentic tool execution, or sensitive dataRed-team, jailbreak, fairness, retrieval-faithfulnessIndependent second-line challengeContinuous drift monitoring; strict entitlements; immutable audit trailHuman-in-the-loop plus a hard kill-switch
MediumAdvisory or customer-facing, with an AI-labelling dutyGolden-set benchmarking; context-recall and relevancy checksSecond-line review of guardrails; vendor model cardsTopic-drift guardrails; hallucination checksHuman-on-the-loop
LowInternal productivity, low autonomyBaseline sanity checks; no statistical backtestBusiness-line self-attestation; automatic registry intakeData-loss-prevention on outbound queriesHuman-in-command

The top row is the cheapest control in the table: a use case barred by law is rejected at intake, before anyone builds an evaluation harness for it.

Below it, the read is quick. When a system can move money, the binding rings are usually Tier and Orchestrate. When it is customer-facing, they are Orchestrate and Challenge. When it is an internal copilot, they are Inventory and Evaluate. Name the system, find its weakest ring, and fix that ring first.

Run the Assurance-Gap Read on your own portfolio.The full diagnostic — the tiering matrix, the five diagnostic questions, and a ninety-day action sequence — is laid out in a free five-page playbook.
Download the GenAI Assurance Grid

Proof: the hole is in the system, not the math

The Apple Card, operated with Goldman Sachs, is the case every risk committee should study. When customers alleged gender bias in credit-limit decisions, New York regulators investigated and found no evidence of intentional bias — the model was not the villain. The Consumer Financial Protection Bureau issued an order of more than eighty-nine million dollars anyway, because the dispute-handling workflow could not properly route and investigate tens of thousands of customer complaints. A sound model, a missing orchestration and oversight layer: this thesis in a single case.

An older caution, from outside banking and outside generative AI, illustrates the third ring specifically. Years ago a large property company let an automated pricing model drive real purchasing decisions at scale, with no effective human control and no working kill-switch when the market turned; losses ran into the hundreds of millions and the business unit was shut down. It is not a bank case and not a generative-AI case — treat it strictly as a warning about what happens when an autonomous system can commit a firm with no working brake, and nothing more.

How to apply this — 5 steps

Read this as a build order, not a big-bang program. Mapped onto a ninety-day sequence, steps one and three land in the first month, steps two and four in the second, and step five in the third.

  1. Stand up the AI System Registry first. Discover shadow AI from network and data-loss-prevention monitoring; register everything generative and agentic, vendor-embedded features included.
  2. Tier by entitlements, not by intention. Cross materiality, autonomy, and tool access. Fast-track the low tier for momentum, and reserve the full regime for anything that decides or transacts.
  3. Build the orchestration layer before agents touch state. Route every model and agent call through a control plane with guardrails, entitlement-checked tool access, transaction limits, a kill-switch, and an audit-ready log.
  4. Replace annual validation with continuous evaluation. Run a golden set in production, and re-evaluate automatically on every vendor model update.
  5. Redefine effective challenge for models you cannot open. Test your own boundary, secure audit rights and change-notice clauses, and keep a concentration register on your frontier providers.

The board action underneath all five is one sentence: stand up the registry and the orchestration layer before any agent touches customers, money, or a regulated decision.

Risks and caveats

This framework does not argue the regulator exited generative AI, and it does not argue validation is finished. SR 26-2 moves generative and agentic AI outside one guidance's scope while explicitly keeping governance expected; classical methods do not map cleanly and must be supplemented, not discarded. The evaluation tooling referenced above — Ragas, TruLens, and similar model-as-a-judge libraries — is cited as emerging practice, not a settled standard. Vietnam legal specifics, the finance-sector grace window and any penalty figure, are reported and pending primary-source confirmation; treat them as directional, not exact, until the underlying texts are verified. The property-company pricing case is not a bank case and not a generative-AI case, and is not offered as banking precedent. And the objection that most of this belongs to IT, cyber, and third-party risk rather than model risk is largely correct — conceded on purpose, because it is what tells you where the control belongs.

This is independent thought leadership, not affiliated with any current or past employer, and not a substitute for your own legal and regulatory review. Vietnamese bank and market references anywhere in this series use publicly disclosed data only.

Where this fits in the series

This framework extends prior weeks rather than re-deriving them. The 4 Stone Guardians governance gate is the pre-deploy checkpoint; this grid is the ongoing assurance of a live, drifting system once that gate has been passed. The Agent Army operating model explains why Orchestrate and tier-on-entitlements bind together the moment autonomy enters the picture — the control plane described here is what that agent operating model runs on top of. The AI-Ready Bank's five-layer data-readiness map is a dependency this grid assumes rather than re-derives: the grid evaluates the retrieval store's faithfulness, not the five data layers underneath it. And the AI-Augmented Bank's five-pillar operating model is the parent architecture — this grid is the inside of its Trust & Governance pillar, specialised to second-line model-risk and validation discipline.

Before your next risk committee

Your model-risk framework has a GenAI hole, and you do not close it by validating harder — the model is closed, drifting, and not yours to open. You close it by assuring the system around it: see it with an AI System Registry, size it by entitlements, contain it with an orchestration layer, watch it with continuous evaluation, and own the vendor risk through disciplined challenge.

Every week, we send one architecture-grade framework like this one to leaders governing AI in banking, in The AI Architect Letter — one issue a week, no hype. The free five-page GenAI Assurance Grid playbook gives you the Assurance-Gap Read to run this quarter.

Get the 5-page GenAI Assurance Grid playbook.All five components in build order, the materiality-tiering matrix, the Assurance-Gap Read — one question per ring — and a ninety-day action sequence.
Get the playbook
Prefer the briefing on video?Watch _Video 11: Stop Validating the Model. Assure the System._ — the five components argued end to end with the evidence.
Watch the briefing

Sources & note. Regulatory references: US Federal Reserve / OCC / FDIC SR 26-2, issued 17 April 2026, superseding SR 11-7 and SR 21-8 and moving to materiality-based oversight, most directly for Fed-regulated banks above approximately $30bn in assets — its attachment places generative and agentic AI "not within the scope" of the guidance while requiring institutions to apply risk-management and governance practices to what it does not cover, so it is outside one guidance's scope and explicitly not outside governance; EU AI Act (Regulation 2024/1689) Article 50, transparency duties beginning to apply 2 August 2026 — machine-readable marking of synthetic content and disclosure of AI interaction, scoped by use case rather than a blanket watermark; Vietnam AI Law 134/2025/QH15 and PDPL Law 91/2025/QH15, in force in 2026, plus a State Bank of Vietnam draft circular on AI in banking out for consultation — the finance-sector grace window and every penalty figure are reported and pending primary-source confirmation, directional rather than exact until the underlying texts are verified. Case evidence — Apple Card / Goldman Sachs: New York regulators investigated allegations of gender bias in credit-limit decisions and found no evidence of intentional bias, so the model was cleared; the CFPB order of more than $89M was for the dispute-handling workflow, which could not properly route and investigate tens of thousands of customer complaints. The pricing-model case is a large property company, from outside banking and outside generative AI, cited strictly as a warning about an autonomous system committing a firm with no working brake — not a bank case, not a GenAI case, and not banking precedent. Hedges that are load-bearing: classical validation methods (conceptual soundness, backtesting, outcomes analysis) do not map cleanly onto a non-deterministic, closed, drifting model and must be supplemented by system-level evaluation, not discarded; the evaluation tooling named here — retrieval-evaluation and model-as-a-judge libraries such as Ragas and TruLens — is emerging practice, useful today and not yet a settled industry standard; the objection that a generative model reached through an interface is better governed through IT, cyber, and third-party risk than through the model-risk lineage is largely correct and is conceded on purpose, because it locates the control in the running system rather than the model's math; and the tiering matrix is a control design, not a legal classification — the Prohibited row tracks uses barred by law, and the exact scope of high-risk obligations for banking remains subject to the primary texts. Vietnamese bank and market references use publicly disclosed data only; this is independent thought leadership, not affiliated with any current or past employer, and not a substitute for your own legal and regulatory review.