Every health system deploying clinical AI ends up in the same room: a cross-functional committee chartered to approve, escalate, and block deployments. In practice it approves and escalates, and blocking almost never happens. In most organisations the committee convenes monthly while models update weekly, so it cannot keep pace with what it governs.
Banking solved this decades ago: in 2011 the Federal Reserve and the OCC issued joint model risk management guidance. This article is a transfer audit, element by element: what maps cleanly, what breaks, and what the breaks mean for how you build governance around the integration problem it sits inside.
Does governance routing scale better than a standalone AI committee?
Usually, yes. A standalone committee is advisory and throughput-limited: it approves and escalates more than it blocks. Governance routing flips this by embedding approval gates and role-based access into the delivery lifecycle, so decisions are made where production is gated. It makes governance a property of the system around the model rather than a standing meeting.
Practical guidance keeps describing the same first deliverable: a committee charter and RACI with authority to approve, escalate, and block deployments. Routing turns that authority into gates embedded in the work itself: pre-development checks, pre-deployment model and data checks, and post-deployment monitoring.
At programme level, a multi-agent system or an AI systems programme replaces the standing committee with routed, tiered decision pathways, each tier owning a named gate and sign-off. Effective challenge survives in both models, but only routing turns it into a gate you must pass rather than a discussion item. It is the same instinct as treating compliance as architecture, with the control built into the pipeline itself.
What does banking model risk management (SR 11-7) teach clinical AI governance?
The control spine transfers almost whole: model inventory, independent validation, ongoing monitoring, and documented decision rights. HIPAA’s privacy logic does not transfer alongside it. NIST’s AI RMF is the shared vocabulary that lets both halves sit in one frame.
SR 11-7’s requirements are well documented: robust development, independent validation with effective challenge, and governance. Its three validation pillars (conceptual soundness, outcomes analysis, ongoing monitoring) remain the intellectual core, and model inventory is the first and most portable artefact: you cannot govern what you cannot list, and the inventory is where data integration shows up first, because every entry names the data the model consumes and where it came from.
In April 2026 the regulators issued SR 26-2, which supersedes SR 11-7, narrows what counts as a model, and tiers oversight by materiality (exposure plus purpose), while carving out generative and agentic AI. Treat SR 11-7 as the spine of transferable practice; SR 26-2 is the current rulebook.
HIPAA’s administrative, physical, and technical safeguards govern protected health information, a data-privacy control family that protects records and does not establish whether a model is clinically sound. Conflating the two is how health systems end up with privacy checklists standing in for validation. NIST’s AI Risk Management Framework lays out four functions (Govern, Map, Measure, Manage) for holding both control families in one frame, and it is where your deployment controls get tested under load.
So the spine transfers cleanly except for the pillar that assumes an objective “correct” to validate against. That is where the transfer breaks.
Why is clinical ground truth subjective, and why does that make validation hard?
Because medicine has no shared, absolute definition of “correct.” A benchmark is scored against what the physician did that day rather than what was right. When correctness is privately held and contested, validation turns into an evidence-and-review process rather than a threshold.
This is the Oracle Problem, described in a16z’s analysis of Protege data. The practical way to verify a model was right is to follow the patient forward over some horizon, and choosing that horizon is itself subjective. One illustration: patient characteristics explain only 3.4% of the variation in partial-versus-total knee replacement, and adding surgeon identity lifts it to 14.8%; across 15 procedures, physician identity accounts for between 7% and 77% of the explained variation. The knee figure sits inside that wider range.
A model can disagree with the recorded action and still be clinically defensible; it can also agree for the wrong reason, and a static benchmark cannot tell the two apart. Outcomes analysis, the most strained pillar, becomes evidence-and-review: how the model behaves on edge cases, how bias was tested, and whether a clinician can verify the reasoning. This is where the data that feeds the contested labels matters, and where hallucination shows up as a compliance problem. Bias mitigation becomes structural, and subtler failure modes like task-hacking will not be caught by a naive eval.
How do you tell whether an AI governance function can actually enforce decisions?
Once validation becomes a review process, what makes governance real is whether anyone must listen. So ask: can it block a deployment? A governance function only counts if it holds a decision right that gates production: the power to block, stop, or kill. A function that can only advise produces routing theatre: meetings and charters that enforce nothing. Charter language alone does not change that.
Grant Thornton frames authority around who can say no. Boards should know which decisions are autonomous, human-in-the-loop, or human-approved, and who set those tiers. The concrete test is whether the decision rights are bound to the gate: deploy approval, rollback or kill switch, and incident response. A governance body that can only log a concern ends up ceremonial, which is why the gate lives in deployment.
The evidence is thin in practice. Only 20% of business leaders say their organisation has an AI-specific incident response playbook with tested owners. Even banking’s own guidance is now formally non-enforceable. Enforcement reduces to a tested mechanism that can actually stop a release, and a kill switch nobody has rehearsed pulling is a checklist item rather than a control.
Conclusion
Banking hands clinical AI a control spine that transfers almost whole, but its cleanest joint seizes up on subjective ground truth. What keeps the whole apparatus honest is the power to stop a deployment.
That leaves you with a simpler model of governance than the one you inherited. The spine carries over; HIPAA’s privacy logic does not. Accepting subjective ground truth turns validation into evidence-and-review and bias mitigation into a structural obligation. Governance ends up as a property of the delivery pipeline, and the one question that matters for any governance function is whether it can block a deployment. If it cannot, it will only ever advise. To place that question inside the wider picture, start with the cluster overview.
Frequently Asked Questions
Where can I find NIST AI RMF guidance for governing AI systems?
The NIST AI Risk Management Framework is published by the US National Institute of Standards and Technology, and the accompanying AI RMF Playbook and Generative AI Profile are available free from the NIST Trustworthy and Responsible AI site. Treat it as the shared vocabulary that lets banking’s model-risk controls and HIPAA’s privacy safeguards sit in one governance frame. Start with the four core functions: Govern, Map, Measure, and Manage.
Does SR 11-7 still apply after SR 26-2?
Not as the current rulebook. SR 26-2, issued in April 2026, superseded SR 11-7, narrowed its scope, and tiers oversight by model materiality, so both exposure and purpose now determine how much scrutiny a model receives. SR 11-7 still works as the spine of transferable practice: model inventory, independent validation, monitoring, and documented decision rights. Use it as the conceptual foundation, not the live compliance text.
Why does SR 26-2 exclude generative and agentic AI?
Regulators drew the line because generative and agentic systems do not behave like the deterministic models SR 11-7 was built to validate. Their outputs are probabilistic, their behaviour shifts with prompts and context, and they can act across systems, which a static inventory and a fixed validation date cannot capture. The exclusion is a signal, not a reprieve: these systems still need governance, just not the classic model-risk template.
What actually belongs in a clinical AI model inventory?
At minimum, each entry should record the model’s purpose and clinical use case, its owner and sponsor, the data it consumes (including PHI exposure), its version and deployment point, its validation evidence, and its review date. Banking treats the inventory as the first control because you cannot govern what you cannot list. For clinical AI, add the human oversight role and escalation path tied to each model.
What is effective challenge, and why does it matter?
Effective challenge is the requirement that someone with the expertise, standing, and incentive can question a model’s assumptions, data, and outputs and have those objections matter. It is not a review stamp or a sign-off; it is structured scepticism with enough authority to change the outcome. Banking builds it into independent validation. In clinical AI, it only counts if the challenger can delay or block a deployment, not merely log a concern.
What is the difference between independent validation and routine testing?
Routine testing checks whether a model works as built, usually against the developer’s own benchmarks and by the people who built it. Independent validation goes further: a separate party with no stake in the outcome re-examines conceptual soundness, reproduces outcomes on unseen data, and stress-tests the assumptions. The independence is the point. Self-reported accuracy is marketing; independent challenge is evidence.
How often should a clinical AI model be revalidated?
There is no single interval, and a fixed annual review is usually the wrong answer. Revalidation should be triggered by material change: a model update, a shift in the patient population or care pathway, new data sources, or drifting performance against monitored outcomes. Banking pairs scheduled monitoring with event-based reviews, and clinical AI should do the same. Set a maximum interval, then revalidate earlier when drift signals appear.
What does a kill switch actually look like in a clinical AI system?
A kill switch is a tested mechanism that can halt or roll back a model’s influence on care, not just a clause in a policy. In practice it means a documented owner, a way to disable the model or revert to the prior workflow within a defined time, and an incident response playbook that names who triggers it and what clinicians do meanwhile. If nobody has practised pulling it, it is not a control.
Who should own clinical AI governance?
No single function can own it alone, but the decision right should sit with someone accountable for patient outcomes, not with IT or procurement. Banking separates the model owner, who is accountable for performance, from independent validation, which reports elsewhere to preserve challenge. In a health system, that means clinical leadership holds accountability, with data, compliance, and informatics as routed contributors rather than figureheads on an advisory committee.
Can a small health system adopt banking’s governance model without a large team?
Yes, because the transferable spine is lightweight: a model inventory, a documented validation step, monitoring, and named decision rights. What a small system cannot copy cheaply is independent validation at scale, so it typically shares validators across a network or buys independent review. Scale the effort to model materiality. Govern the few models that can harm patients deeply; handle the rest with lighter tiers.
What happens if a clinical AI model makes a wrong recommendation?
A wrong recommendation is not automatically a governance failure, because clinical ground truth is contested and a model can disagree with the recorded action yet still be defensible. What matters is whether the system caught the problem and whether a human could override it. Governance covers the response: detection through monitoring, a clinician’s authority to reject the output, and an incident review that feeds back into validation.
Is a high accuracy score enough to approve a clinical AI tool?
No. A single accuracy number hides which cases the model gets wrong and cannot distinguish a defensible disagreement from a lucky guess, because medicine has no shared, absolute definition of correct. Approval should rest on evidence and review: how the model behaves on edge cases, how bias was tested, and whether clinicians can verify its reasoning. Treat the score as one input, never the verdict.