Many clinical AI pilots fall over at governance review: the first time the committee asks where the PHI flows, the answer is a consumer API with no BAA and no plan for the EHR.
You’ve probably watched this happen. The demo performs well, then stalls at governance review or under real concurrency. The hosting decision and the PHI boundary are what go untreated. Clinical AI has to reach EHR data, which forces a hosting choice before you pick a model (the integration problem at the heart of healthcare AI).
So here’s the frame. Three hosting legs: on-prem/VPC-isolated, BAA-covered cloud APIs, and external model APIs. Two accuracy paths: RAG grounding versus fine-tuning. One gate: the BAA. By the end you can rule options in or out before a POC, on the strength of a test for what holds up under load and under contract.
RAG grounding or fine-tuning: which is safer for clinical AI accuracy?
Grounding first. RAG pulls verified sources and attaches a grounding score, so answers stay source-linked, auditable and updatable without retraining. Fine-tuning bakes knowledge into the weights, where it can’t be inspected, updated or debiased.
Bedrock’s knowledge bases ingest clinical documents, generate vector embeddings and retrieve relevant chunks at query time, then check the answer against those chunks. Below a configured threshold, the response is blocked before it reaches your application.
Fine-tuning has a place: domain-focused tuning of small models can lift performance on narrow clinical tasks. But domain-adapted weights can’t be source-linked, so validation is harder and bias accumulates. In clinical work the architecture around the model matters more than the model, and fine-tuning leaves you without that auditable architecture.
The vector store is where PHI leaks. If access control stops at the model API, the document store itself isn’t controlled at all: the index becomes an unmanaged PHI store.
The verdict: grounding is the default, and fine-tuning is reserved for a bounded, validated, de-identified task under human review. Treat ungrounded statements as compliance violations (compliance as architecture).
Where either strategy runs is the next decision, and EHR integration forces it before you choose a model.
On-prem/VPC-isolated, BAA-covered cloud APIs, or external model APIs: which actually holds up under clinical load?
The three hosting paths trade control against operational burden, and the right pick depends on what your team can actually run.
VPC-isolated regulated deployment is the maximum-control option: PHI stays inside your perimeter and no provider sees a patient record. BAA-covered cloud APIs such as Amazon Bedrock and Microsoft Azure OpenAI Service balance model quality against a vendor-managed baseline and contractual PHI coverage. External model APIs from OpenAI, Anthropic or Google fail the bar on consumer tiers, which carry no BAA and no residency guarantees.
The verdict: an external model API without a signed BAA is out, the BAA-covered cloud API is the default for most clinical workloads, and on-prem/VPC isolation earns its cost only when residency or GPU control demands it.
The hidden cost is operations. Self-hosting means GPU and VRAM sizing, a serving layer, ECC memory, confidential computing and patching, and cloud per-token billing scales unfavourably past roughly 100,000 clinical inferences a month. Residency decides viability: AWS Regions and AWS GovCloud pin where PHI can live, and if a cloud path can’t satisfy residency, on-prem/VPC is the only option left. A model gateway with Model Context Protocol as the isolation and audit boundary lets you switch models behind one controlled surface.
The EHR surface forces this choice before you pick a model (data integration is the binding constraint). Whichever leg you pick, the same question applies to any vendor’s data architecture under load.
Will a clinical AI vendor’s data architecture hold up under real clinical load?
“Holds up under load” is a compound test you can run on every vendor: isolated per-patient sessions, pinned regions, and no dependence on a single external endpoint.
Load means concurrency, tail latency, throughput and uptime, measured at 10, 50 and 100 concurrent users. A FastAPI-versus-Triton benchmark shows why this matters: FastAPI handles single requests with lower overhead, but Triton’s dynamic batching nearly doubled throughput on the same GPU at marginal added latency.
Probe four failure modes: cross-tenant context bleed, where one patient’s context surfaces in another session; missing session isolation; dependence on one external endpoint; and a production tier missing the feature the pilot used. Bleed tends to show up at concurrency, not in a single-user demo.
ECC GPU memory detects and corrects single-bit errors, so corruption fails loudly instead of silently reaching a clinical decision. Demand Third-Party/Vendor AI Risk with audit rights, not a polished reference list. TrueFoundry reports that Medtronic, Innovaccer, Aviva, Siemens Healthineers and ResMed have passed compliance and security review, and that Innovaccer runs roughly 17 million clinical inference requests a month inside AWS GovCloud. Epic and Cerner data via Redox must not reintroduce unmanaged PHI flows.
None of that architecture matters until the contract permits touching PHI.
What should you look for in a vendor’s HIPAA and BAA posture?
HIPAA is the regulation; a Business Associate Agreement is the contract that makes a vendor permissible to touch PHI. There is no “HIPAA-certified” vendor, and no signed BAA means disqualification.
HHS does not endorse any certification. What can be compliant is a deployment: a specific model, on specific infrastructure, under contract.
Then check scope, not existence. The BAA has to cover the exact tier, model and feature in production, including zero data retention terms. OpenAI signs a BAA only for sales-managed Enterprise or Edu, not the self-serve tiers (OpenAI enterprise privacy). Anthropic’s BAA covers the first-party API and a HIPAA-ready enterprise plan, not the consumer tiers (Anthropic trust center). A consumer tier carries no BAA at all, so PHI on it is a violation.
Follow the chain. Audit rights must extend to subcontractors and sub-processors, a common failure point. A BAA covers the vendor’s side on the surfaces it names; workforce training, access controls and risk analysis stay yours.
Weigh SOC2 Type II, audit rights and regional controls as supporting signals beside the BAA. And keep Epic and Cerner paths via Redox, plus any MCP wiring, inside the BAA’s scope.
Conclusion
Under real concurrency, accuracy, hosting and compliance collapse into one decision.
The default that falls out: a BAA-covered cloud API (Amazon Bedrock or Azure OpenAI Service) behind a mediation and audit boundary, grounded with RAG.
Take the compound test to every vendor and ask for p95, throughput and uptime evidence at 10, 50 and 100 concurrent users.
Let the gate end evaluations early: no signed BAA covering the exact tier, model and sub-processor chain means disqualification.
Control survives only if you staff, audit and renew it, and that is what holds up under real clinical load.
Frequently Asked Questions
What actually counts as PHI when we deploy clinical AI?
Anything that identifies a patient or could reasonably be used to re-identify them counts as PHI, which in practice means your prompts, retrieved documents, vector embeddings, logs and audit trails as well as the obvious fields. The safe assumption is that every artefact touching a real patient record is in scope, so map each storage and transit point before you pick a hosting path.
Do we still need a BAA if the AI only works on de-identified data?
If the data is genuinely de-identified under the HIPAA Privacy Rule, either by safe harbour or expert determination, it is no longer PHI and a BAA is not strictly required. In practice the expert determination is hard to satisfy for clinical data, and the moment you re-link identifiers or feed the output back into a patient record you need one, so treat it as required.
Can we use a frontier model from OpenAI, Anthropic or Google in clinical use?
Yes, but not by going direct to their standard consumer or pay-as-you-go APIs, which carry no BAA and no residency guarantees. You reach the same frontier models through a BAA-covered cloud endpoint such as Amazon Bedrock, Azure OpenAI Service or a comparable enterprise surface, where the contract, retention terms and region pinning are explicit. The model matters less than the contract it runs under.
How do I tell whether a RAG grounding score means anything?
A grounding score is only worth trusting if it measures something real, such as retrieval relevance or whether the answer is supported by the cited source, and if it is computed and logged on every response. Ask whether it can block an answer when it falls below a threshold, because a score that only decorates the interface tells you nothing about clinical safety.
Is on-prem or VPC-isolated hosting actually cheaper over three years?
Rarely, unless you already run GPU infrastructure at scale. The licence saving is real, but you inherit GPU and VRAM sizing, a serving layer, ECC memory, confidential computing, patching and the headcount to operate all of it, which is control you must staff rather than a one-off saving. For most clinical workloads the BAA-covered cloud API is cheaper to run and easier to audit.
What p95 latency should clinical AI hit to be usable at the point of care?
For anything a clinician waits on mid-consultation, aim for a p95 under roughly two seconds, and treat anything beyond three to four seconds as a workflow problem rather than a minor delay. Batch documentation and overnight review can tolerate far more, so measure p95 at 10, 50 and 100 concurrent users and weigh it against throughput, because dynamic batching trades tail latency for throughput.
Does signing a BAA make the vendor responsible if PHI is breached?
No. The BAA makes the vendor a business associate with defined obligations, including safeguarding PHI and notifying you of a breach, but the covered entity stays accountable for HIPAA compliance overall. What the agreement buys you is enforceability and a clear chain of responsibility, which is why scope, sub-processors and audit rights matter as much as the signature itself.
What happens if our clinical AI vendor is acquired or swaps its model provider mid-contract?
That is exactly why the contract needs change-of-control and sub-processor clauses with advance notification, and why you want a mediation layer in front of the model. If the vendor is acquired or silently repoints to a different provider, you need the right to re-approve the change and the architectural ability to move, otherwise your BAA scope and validation evidence quietly stop matching what is deployed.
Can we switch model providers later without rebuilding the integration?
You can, if the model sits behind a gateway or a Model Context Protocol boundary rather than being wired directly into your application. The gateway becomes the single controlled surface you audit and repoint, so a provider swap is a routing change rather than a rebuild. You still need to revalidate prompts, grounding behaviour and the BAA for the new model, so budget for that work.
If we ground answers with RAG, do we still need human-in-the-loop review?
For any output with clinical consequence, yes. Grounding makes an answer source-linked and auditable, and guardrails in blocking mode can stop a response that fails its check, but neither replaces clinician sign-off on a decision that affects a patient. RAG raises the floor on trustworthiness; it does not transfer accountability away from the person acting on the output.
What does “no cross-tenant context bleed” look like in practice?
It means one customer’s prompts, retrieved chunks or cached context can never surface in another customer’s session, even under heavy concurrency. Concretely, you want isolated per-patient sessions, tenant-scoped retrieval indexes and no shared caches that quietly hold state across accounts. Ask the vendor to demonstrate it under load, because bleed tends to appear at concurrency, not in a single-user demo.