Guides › Data sovereignty
Data Sovereignty for Enterprise AI: Keeping Sensitive Data Under Your Control
Data sovereignty, for an organization running AI, means your sensitive data stays under your control, your policies and your jurisdiction even while the models doing the work belong to someone else. The term used to be about where a database sat. Enterprise AI changed the question, because the most capable models now run outside your boundary and every prompt, retrieval and fine-tune is a decision about what crosses it. This guide defines the term for an AI deployment, lays out the four routes sensitive data takes out of your environment, explains why vendor terms alone don't close them, and describes what a sovereign AI deployment looks like in practice.
What data sovereignty means when the model is someone else's
"Data sovereignty" carries two meanings, and most arguments about it come from mixing them. The jurisdictional meaning is about law. Data is subject to the laws of the country where it is collected, stored or processed, and a provider headquartered elsewhere may be reachable by a foreign government even when the servers are local. The operational meaning is about control. Who can read a record, which systems process it, how long copies persist, and whether anyone outside your organization can learn from it. Regulators care about both. Security teams can only build for the second one, and it's the one AI changed.
Before generative AI, keeping sensitive data sovereign in the operational sense was mostly a hosting problem. Pick a region, encrypt at rest, restrict access, and the data stayed put. A frontier model breaks that pattern in one move. The model isn't in your environment, and it can't answer a question about your data without being shown some of it. Every useful AI workflow is therefore a transfer, and the practical question becomes which transfers you allow, what they contain, and whether you can prove it afterward.
That is the working definition this guide uses. An AI deployment is sovereign when your organization decides, and can demonstrate, what sensitive data reaches any third party, in what form, for which user, and for how long it persists. The jurisdictional questions still matter, and the data sovereignty vs data residency page covers them. The rest of this cluster is mostly about control, because that's where the exposure sits.
Four routes sensitive data takes out of your boundary
Each route below moves data to a third-party model provider, and each has a different owner inside your organization, which is why none of them tends to be governed end to end.
The first route gets the most attention and is the least controllable by engineering. The second is the one most organizations underestimate, because a connector feels like a configuration setting and behaves like a bulk export. A single answer can carry a dozen document chunks to the provider, scoped by whatever permissions the connector happened to inherit, and the person who asked never sees what was sent. The routes a knowledge base leaks through apply here in full, with a third party on the receiving end.
Why contract terms and settings aren't the control
The enterprise tiers of the major providers now commit, by default, not to train on customer inputs and outputs, and most offer a retention window with a zero-data-retention option for qualifying workloads. Those terms are worth having and worth negotiating. They're also commitments about what the provider will do with data after it arrives. Nothing in them reduces what arrives.
Three gaps follow. The terms attach to the account, so an employee pasting a customer file into a personal consumer account is covered by none of them, and on consumer tiers training on inputs can be the default rather than the exception. The terms are unverifiable from your side; you can't inspect a provider's retention pipeline, so compliance rests on audit reports and trust. And the terms say nothing about what your own connectors chose to send, which is the route with the highest volume and the least visibility.
What a sovereign AI deployment looks like in practice
The pattern that holds up is to leave the data where it lives and bring a scoped, substituted view of it to the model. Five properties define it.
- The data stays put. Source systems keep the records, the permissions and the audit trail they already have. Nothing is migrated into a new store to make it AI-ready.
- Retrieval is scoped to the person asking. A query runs against what that identity is entitled to see, with entitlements enforced at query time rather than inherited once at index time. See RAG permissions.
- Real values never reach the provider. Before retrieved content crosses to a third-party model, names, account numbers and other identifiers are replaced with stand-ins. The model reasons on the stand-ins; the real values are restored in the reader's own window, gated on their entitlement. The third-party platforms page walks through the mechanics.
- Every retrieval is recorded. Which chunks each query returned, for whom, and whether the pattern looks like extraction, keyed to opaque IDs so the record itself contains no content. See knowledge base telemetry.
- Deployment follows the obligation. The same controls run as a cloud service, inside your private cloud, or as an offline bundle in a disconnected enclave with no outbound dependency. See sovereign AI deployment.
Two things fall out of this design. The frontier model is still in the loop, so you keep its capability without self-hosting. And "what did we send to the provider" becomes a question with an exact answer, which is the question an auditor, a regulator or a customer will eventually ask.
Regulations that invoke sovereignty, and what each one asks
Most regulations never use the phrase. They set obligations that a sovereignty posture satisfies. A short map follows, with the usual caution that obligations depend on your sector and your counsel.
The framework crosswalk maps data-layer failure modes across MITRE ATLAS, OWASP, NIST AI RMF, ISO/IEC 42001 and the EU AI Act, and is the place to go if you need a finding expressed in a specific vocabulary.
Where Hardshell sits
Hardshell is the secure data foundation between an organization's data and its AI platforms. It connects to existing systems without migrating them, inherits their permissions, scopes every retrieval to the person asking, substitutes stand-ins for real identifiers before anything crosses to a model provider, and records every retrieval keyed to opaque IDs. Exposure testing measures what a knowledge base or training set can leak before a model goes live, and dataset hardening reduces it. It deploys as a cloud service, inside a customer's private cloud, or as an offline bundle for disconnected environments.
Frequently asked questions
Is data sovereignty the same as data residency?
No. Residency is where data is stored. Sovereignty is whose laws govern it and who controls what happens to it, including at inference time, when an AI query can move data to a model served somewhere else. A regional hosting choice settles residency and leaves most of the sovereignty question open. The comparison page covers the distinction and the regulations behind it.
Does an enterprise plan from OpenAI, Anthropic, Microsoft or Google make our data sovereign?
It improves the terms considerably. No training on your data by default, a defined retention window, often a zero-data-retention option and regional storage. It doesn't change what your users and connectors send, can't be verified from your side, and doesn't cover employees using personal accounts. Treat it as necessary, and pair it with controls on what crosses.
Can we get there by self-hosting an open-weight model?
Yes, for workloads where an open-weight model is good enough and you can carry the serving cost. It's the cleanest answer to the jurisdictional question. It doesn't by itself solve the operational one, because a self-hosted assistant over an over-permissioned knowledge base still shows people documents they were never cleared to see. The deployment page compares the options.
What about zero-data-retention APIs?
Zero data retention means the provider discards prompts and completions after processing rather than holding them for a monitoring window. It narrows the retention route and is worth requesting where eligible. Data still crosses, is still processed, and is still visible to the provider's systems at inference time, so it's a reduction in persistence rather than in exposure.
How do we find out what has already left?
Start with the connectors, since they're the highest-volume route and usually the least logged. Retrieval telemetry on the knowledge base shows which documents are being returned and to whom; the self-test turns that into one number. For prompts, your provider's admin console and your DLP logs are the available record. For fine-tunes, the training manifests are the record, and whatever went into the weights is there to stay.