Model capability used to be the hard part of healthcare AI. It isn’t any more. The binding constraint is now the system around the model: the data, the governance, the compliance controls, and the deployment choices that decide whether any of it survives the clinical load your business runs.
Heidi, the Australian-founded clinical scribe, puts the split plainly. “The model is maybe 20% of the system,” says co-founder and CTO Yu Liu, “and the data architecture is what determines whether the other 80% holds up under clinical load.” Heidi runs 2.7 million patient interactions a week across more than 190 countries.
The numbers agree. A 2025 HFMA survey of 233 health systems found 88% use AI, but only 18% have mature governance and a fully formed AI strategy. IDC and Lenovo found that 88% of AI proofs of concept never reach widescale deployment. Adoption has outpaced governance.
This series explains two failure shapes. The first is the pilot trap, an operating-model problem where a successful demo never becomes a governed capability. The second is the integration failure that carries direct liability: the Xsolis breach exposed 1,396,519 people across seven health systems. The four articles below map those two failure shapes onto four problems: the operating model, governance, compliance, and deployment.
In This Series
- Why Data Integration Is the Binding Constraint for Healthcare AI
- What Banking Model Risk Management Teaches Clinical AI Governance
- Treating Compliance as Architecture in Clinical AI Systems
- Clinical AI Deployment and Vendor Choices That Hold Up Under Load
Why is data integration the binding constraint for healthcare AI?
Model capability has commoditised; your real constraint is the operational surface around it. Heidi’s CTO puts the model at “maybe 20% of the system,” and the other 80% is data architecture that has to hold under real clinical load. With roughly 88% of AI pilots never reaching production and only 18% of health systems holding mature governance, the distance between a working demo and a governed, integrated capability is where outcomes are actually decided.
The pilot trap is an operating-model failure. A pilot succeeds on sponsor energy and a clean, hand-picked dataset. Production meets messy live feeds with no owner for data contracts and no definition of “done” beyond the demo. Underneath it sits the FHIR R4 foundation: legacy HL7, CCD, and X12 feeds reconciled into a shared, typed model, which is what AWS HealthLake does with its R4 store. Streaming versus batch is a topology decision driven by clinical latency and auditability. Streaming is where session-isolation and context-leakage risks surface, which is why live delivery paths like Zerobus Ingest matter.
After the data is sorted come governing clinical AI models and compliance requirements for clinical AI. The full argument is in Why Data Integration Is the Binding Constraint for Healthcare AI.
What does banking model risk management teach clinical AI governance?
Banking’s SR 11-7 guidance is the closest mature analogue for clinical AI governance. Its spine (a model inventory, independent validation, and ongoing drift monitoring) transfers. What does not is HIPAA’s data-safeguard logic, a different control family. The harder obstacle is the Oracle Problem: patient characteristics explain only 3.4% of variation in a knee-replacement decision until surgeon identity lifts it to 14.8%.
Those controls map onto a clinical AI estate, with the NIST AI Risk Management Framework as the shared vocabulary that lets banking-style model risk and healthcare compliance sit in one frame. Governance routing, where decisions sit in the systems that gate production, scales better than a standalone advisory committee. Because ground truth is contested, validation becomes an evidence-and-review process.
Governance needs a governed data estate, which is why data integration comes first. What Banking Model Risk Management Teaches Clinical AI Governance carries the comparison.
Why does compliance need to be architecture, not paperwork?
A signed BAA is only a baseline, not a safeguard. In a clinical setting a hallucination is a false statement about a patient, so it is governed as an accuracy and disclosure violation rather than a logged quality defect. Compliance that holds under load is designed in: PHI de-identification at the trust boundary, immutable audit logs, and session isolation that stops one patient’s context bleeding into another’s. These are properties of the system, not promises in a contract.
Grounding makes hallucination a controllable property: retrieval-augmented generation bound to source records, plus grounding checks in blocking mode. PHI de-identification at the trust boundary (Safe Harbor’s 18 identifiers, the minimum necessary standard applied to prompts) keeps ePHI out of the inference path. Audit completeness means every PHI-touching event captured at resource level, held immutably, and session-isolated. Residency becomes jurisdiction-by-design across GDPR, UK GDPR, and the EU AI Act, pinned to permitted regions through options like AWS GovCloud.
Liability makes this concrete. The Xsolis breach exposed 1,396,519 people across seven health systems. Treating Compliance as Architecture in Clinical AI Systems has the full picture.
Which deployment and vendor choices hold up under clinical load?
Your hosting choice is a trade between control and operational burden. On-prem or VPC-isolated deployment maximises control; BAA-covered cloud APIs such as Amazon Bedrock or Azure OpenAI Service balance it; external model APIs generally fail the clinical bar without extra controls. Weigh whether a vendor’s architecture survives real load, with isolated sessions, pinned regions, no cross-tenant context bleed, and no dependence on a single external endpoint, and whether a BAA actually covers every PHI flow.
RAG grounding gives auditable, source-linked accuracy you can update without retraining; fine-tuning bakes weights in and complicates audit, so grounding is the safer clinical default. Alongside a BAA, read the signals that matter: SOC 2 Type II, audit rights, and whether interoperability paths through Redox, Epic or Cerner reintroduce unmanaged PHI flows. And understand the HIPAA-versus-BAA distinction: the regulation versus the contract that makes a vendor permissible.
Clinical AI Deployment and Vendor Choices That Hold Up Under Load covers the full comparison.
Resource Hub: Healthcare AI’s Integration Problem Deep Dives
The four articles interlock in sequence. The data foundation makes governance meaningful. Governance becomes enforceable only where compliance is built in. Deployment and vendor decisions are where all three land in practice. Read them in order, or jump to your current constraint.
Understanding the constraint
- Why Data Integration Is the Binding Constraint for Healthcare AI: the 20/80 split, why the pilot trap is an operating-model failure, and the FHIR foundation. About 12 minutes.
- What Banking Model Risk Management Teaches Clinical AI Governance: SR 11-7’s transferable controls, the Oracle Problem, and governance routing. About 11 minutes.
Building systems that hold up
- Treating Compliance as Architecture in Clinical AI Systems: hallucination as a compliance violation, PHI at the trust boundary, immutable audit, and residency. About 13 minutes.
- Clinical AI Deployment and Vendor Choices That Hold Up Under Load: isolation models, RAG versus fine-tuning, and reading a vendor’s HIPAA and BAA posture. About 11 minutes.
Where to start
The order depends on your situation. The health systems that scale AI will be the ones that turn pilots into governed operating capabilities.
Frequently Asked Questions
Is it true that a more capable AI model would finally get our pilots into production?
No. Model capability is no longer the bottleneck, so a better model does not close the gap between a demo and production. Heidi’s CTO sizes the model at roughly 20% of the system; the other 80% is data architecture and governance that must hold under real clinical load. Upgrade the model if you like, but fix the operating model around it first.
How do I tell whether we have a pilot trap or an integration failure?
A pilot trap is an operating-model failure where a successful demo never becomes a governed capability; an integration failure is a technical breakdown that carries hard liability, like the Xsolis breach that exposed 1,396,519 people. If sponsorship fades and no one owns the data contracts, you have the first. If live feeds or PHI controls break, you have the second.
What happens if protected health information leaks into a model prompt?
Treat it as a compliance incident, not a logging glitch. In a clinical setting a false statement about a patient is an accuracy and disclosure violation, and leaked PHI compounds it. Controls belong at the trust boundary: de-identify the 18 Safe Harbor identifiers, apply the minimum necessary standard to prompts, and keep ePHI out of the inference path entirely.
How long should a clinical AI pilot take before it reaches governed production?
There is no fixed clock, and the deadline should not be the demo. Move a pilot into production only when it has an owner, defined data contracts, and a governed home in the systems that gate deployment. Teams that treat the demo as the finish line stall; teams that define done as governed capability move faster overall.
Is fine-tuning ever worth it for clinical AI, or should we always use RAG?
Sometimes, but RAG is the safer default for clinical work. RAG grounds answers in source records you can audit and update without retraining, which keeps hallucination a controllable property. Fine-tuning bakes behaviour into the weights, so when a guideline changes you must retrain and re-evidence the model. Reserve fine-tuning for narrow, stable tasks.
Does signing a BAA make a vendor HIPAA compliant?
No. A BAA is a contract that makes a vendor permissible, not proof the vendor is safe, and it is not the same thing as the HIPAA regulation itself. Treat it as a baseline you then verify against architecture: isolated sessions, pinned regions, immutable audit at resource level, and no cross-tenant context bleed. Paper does not hold under clinical load.
What is the difference between streaming and batch ingestion for clinical data?
It is a topology decision, not a preference. Batch ingestion suits retrospective analytics and bulk reconciliation of legacy HL7, CCD, and X12 feeds into the FHIR R4 model. Streaming suits latency-sensitive clinical workflows where a delayed signal changes care. Choose based on clinical latency and auditability requirements, then reconcile both into the same typed, governed data estate.
How often should we revalidate a clinical AI model after it goes live?
Revalidation is continuous, not a one-off. SR 11-7’s transferable spine includes ongoing drift monitoring, and clinical AI needs the same: watch for performance decay as populations, coding, and guidelines shift, and re-run independent review on a defined cadence or when drift crosses a threshold. A single accuracy figure at launch proves nothing about month twelve.
Is on-premises hosting always the safest option for clinical AI?
No. On-premises or VPC-isolated hosting maximises control but adds serious operational burden, and control you cannot staff is not safety. BAA-covered cloud APIs such as Amazon Bedrock or Azure OpenAI Service balance control with manageability. The real test is whether the architecture survives load: isolated sessions, pinned regions, and no reliance on a single external endpoint.
What is a model inventory, and why does every clinical AI estate need one?
A model inventory is the register of every AI model in use, what each one does, who owns it, and how it is validated and monitored. Banking’s SR 11-7 treats it as the spine of model risk management, and it transfers cleanly to clinical AI. Without one you cannot prove oversight, route governance, or answer a regulator’s most basic question.
Can a small clinic follow the same governance playbook as a large health system?
The principles are the same, but the scale of effort differs. A small clinic still needs an owner for AI, defined data contracts, a model inventory, and audit that captures PHI-touching events. What changes is how much you build versus buy: a BAA-covered cloud service can carry controls a small team could never run in-house. Start with the data foundation.
What does it mean to run grounding checks in blocking mode?
It means the system refuses to answer when grounding fails, rather than answering anyway. In blocking mode, a grounding check compares the model’s output against the retrieved source records and halts any response that is not supported by them. That turns hallucination from an unpredictable defect into a controllable, auditable property of the pipeline.