Guides › Data sovereignty

Data Sovereignty for Enterprise AI: Keeping Sensitive Data Under Your Control

Data sovereignty, for an organization running AI, means your sensitive data stays under your control, your policies and your jurisdiction even while the models doing the work belong to someone else. The term used to be about where a database sat. Enterprise AI changed the question, because the most capable models now run outside your boundary and every prompt, retrieval and fine-tune is a decision about what crosses it. This guide defines the term for an AI deployment, lays out the four routes sensitive data takes out of your environment, explains why vendor terms alone don't close them, and describes what a sovereign AI deployment looks like in practice.

At a glance
  • What it means. Control over who can see, process, retain and learn from your sensitive data, wherever the model runs. Residency (where data is stored) and localization (a legal requirement to keep it in-country) are narrower questions, covered on the sovereignty vs residency page.
  • Why AI reopened it. Prompts, retrieval context, fine-tuning uploads and provider logs all move data across a boundary that storage-era residency controls never covered.
  • What vendor terms cover. Training exclusions, retention windows and zero-data-retention options govern what a provider does after data arrives. They don't change what arrives.
  • What the control looks like. Data stays where it lives, retrieval is scoped to the person asking, real identifiers are replaced before anything reaches a third-party model, and every retrieval is recorded. Deployment follows your obligations: cloud, private cloud, on-prem or disconnected.

What data sovereignty means when the model is someone else's

"Data sovereignty" carries two meanings, and most arguments about it come from mixing them. The jurisdictional meaning is about law. Data is subject to the laws of the country where it is collected, stored or processed, and a provider headquartered elsewhere may be reachable by a foreign government even when the servers are local. The operational meaning is about control. Who can read a record, which systems process it, how long copies persist, and whether anyone outside your organization can learn from it. Regulators care about both. Security teams can only build for the second one, and it's the one AI changed.

Before generative AI, keeping sensitive data sovereign in the operational sense was mostly a hosting problem. Pick a region, encrypt at rest, restrict access, and the data stayed put. A frontier model breaks that pattern in one move. The model isn't in your environment, and it can't answer a question about your data without being shown some of it. Every useful AI workflow is therefore a transfer, and the practical question becomes which transfers you allow, what they contain, and whether you can prove it afterward.

That is the working definition this guide uses. An AI deployment is sovereign when your organization decides, and can demonstrate, what sensitive data reaches any third party, in what form, for which user, and for how long it persists. The jurisdictional questions still matter, and the data sovereignty vs data residency page covers them. The rest of this cluster is mostly about control, because that's where the exposure sits.

Four routes sensitive data takes out of your boundary

Each route below moves data to a third-party model provider, and each has a different owner inside your organization, which is why none of them tends to be governed end to end.

RouteWhat crossesWho usually controls itWhere it shows up
Prompts and pasted contentWhatever an employee types or uploads, including documents, spreadsheets and screenshotsIndividual users, under acceptable-use policyShadow AI on personal accounts; consumer tiers where training on inputs can be the default
Retrieval contextDocument chunks an AI knowledge base or connector returns to the model as context for an answerApplication and data teamsCopilots, assistants and agents wired to SharePoint, Drive, ticketing, CRM and file shares
Fine-tuning and training uploadsEntire datasets copied into a provider's infrastructure and, after training, encoded in model weightsML and platform teamsCustom models, fine-tunes, batch jobs
Provider retention and loggingPrompts, outputs and context the provider holds for abuse monitoring, support or debugging, visible to its subprocessorsProcurement and legal, through the contractDefault retention windows, zero-data-retention eligibility, subprocessor lists

The first route gets the most attention and is the least controllable by engineering. The second is the one most organizations underestimate, because a connector feels like a configuration setting and behaves like a bulk export. A single answer can carry a dozen document chunks to the provider, scoped by whatever permissions the connector happened to inherit, and the person who asked never sees what was sent. The routes a knowledge base leaks through apply here in full, with a third party on the receiving end.

Why contract terms and settings aren't the control

The enterprise tiers of the major providers now commit, by default, not to train on customer inputs and outputs, and most offer a retention window with a zero-data-retention option for qualifying workloads. Those terms are worth having and worth negotiating. They're also commitments about what the provider will do with data after it arrives. Nothing in them reduces what arrives.

Three gaps follow. The terms attach to the account, so an employee pasting a customer file into a personal consumer account is covered by none of them, and on consumer tiers training on inputs can be the default rather than the exception. The terms are unverifiable from your side; you can't inspect a provider's retention pipeline, so compliance rests on audit reports and trust. And the terms say nothing about what your own connectors chose to send, which is the route with the highest volume and the least visibility.

A setting you can't verify is a promise. Sovereignty is decided before data crosses the boundary, in what you allow to cross and in what form. Everything after that point is contract management.

What a sovereign AI deployment looks like in practice

The pattern that holds up is to leave the data where it lives and bring a scoped, substituted view of it to the model. Five properties define it.

  • The data stays put. Source systems keep the records, the permissions and the audit trail they already have. Nothing is migrated into a new store to make it AI-ready.
  • Retrieval is scoped to the person asking. A query runs against what that identity is entitled to see, with entitlements enforced at query time rather than inherited once at index time. See RAG permissions.
  • Real values never reach the provider. Before retrieved content crosses to a third-party model, names, account numbers and other identifiers are replaced with stand-ins. The model reasons on the stand-ins; the real values are restored in the reader's own window, gated on their entitlement. The third-party platforms page walks through the mechanics.
  • Every retrieval is recorded. Which chunks each query returned, for whom, and whether the pattern looks like extraction, keyed to opaque IDs so the record itself contains no content. See knowledge base telemetry.
  • Deployment follows the obligation. The same controls run as a cloud service, inside your private cloud, or as an offline bundle in a disconnected enclave with no outbound dependency. See sovereign AI deployment.

Two things fall out of this design. The frontier model is still in the loop, so you keep its capability without self-hosting. And "what did we send to the provider" becomes a question with an exact answer, which is the question an auditor, a regulator or a customer will eventually ask.

Regulations that invoke sovereignty, and what each one asks

Most regulations never use the phrase. They set obligations that a sovereignty posture satisfies. A short map follows, with the usual caution that obligations depend on your sector and your counsel.

RegimeWhat it asks of an AI deploymentWhere the sovereignty question lands
GDPR (EU)Lawful basis and processor terms for personal data; Chapter V conditions on transfers outside the EEAA US-headquartered model provider is a processor and a transfer, regardless of region selection
EU AI Act, as amended by Regulation (EU) 2026/1744Data governance, logging and documentation for high-risk systems (Annex III obligations apply from 2 December 2027); transparency duties already in forceDemonstrating what data a system used and what left it. Many internal assistants will not be high-risk
HIPAA (US)A business associate agreement with any vendor that receives PHI; minimum necessary disclosureRetrieval context carrying PHI to a provider is a disclosure; substitution reduces what is disclosed
GLBA and state privacy laws (US)Safeguards for nonpublic personal information; vendor oversightConnector-fed assistants over customer records
NIST SP 800-171 and CMMC (US defense supply chain)Protection of controlled unclassified information in nonfederal systems, including where it flowsCUI in prompts or retrieval context reaching a commercial AI service
ITAR and EAR (US export control)Controls on export-controlled technical data, including access by foreign personsTechnical data reaching a cloud service where foreign access can't be excluded

The framework crosswalk maps data-layer failure modes across MITRE ATLAS, OWASP, NIST AI RMF, ISO/IEC 42001 and the EU AI Act, and is the place to go if you need a finding expressed in a specific vocabulary.

Where Hardshell sits

Hardshell is the secure data foundation between an organization's data and its AI platforms. It connects to existing systems without migrating them, inherits their permissions, scopes every retrieval to the person asking, substitutes stand-ins for real identifiers before anything crosses to a model provider, and records every retrieval keyed to opaque IDs. Exposure testing measures what a knowledge base or training set can leak before a model goes live, and dataset hardening reduces it. It deploys as a cloud service, inside a customer's private cloud, or as an offline bundle for disconnected environments.

Frequently asked questions

Is data sovereignty the same as data residency?

No. Residency is where data is stored. Sovereignty is whose laws govern it and who controls what happens to it, including at inference time, when an AI query can move data to a model served somewhere else. A regional hosting choice settles residency and leaves most of the sovereignty question open. The comparison page covers the distinction and the regulations behind it.

Does an enterprise plan from OpenAI, Anthropic, Microsoft or Google make our data sovereign?

It improves the terms considerably. No training on your data by default, a defined retention window, often a zero-data-retention option and regional storage. It doesn't change what your users and connectors send, can't be verified from your side, and doesn't cover employees using personal accounts. Treat it as necessary, and pair it with controls on what crosses.

Can we get there by self-hosting an open-weight model?

Yes, for workloads where an open-weight model is good enough and you can carry the serving cost. It's the cleanest answer to the jurisdictional question. It doesn't by itself solve the operational one, because a self-hosted assistant over an over-permissioned knowledge base still shows people documents they were never cleared to see. The deployment page compares the options.

What about zero-data-retention APIs?

Zero data retention means the provider discards prompts and completions after processing rather than holding them for a monitoring window. It narrows the retention route and is worth requesting where eligible. Data still crosses, is still processed, and is still visible to the provider's systems at inference time, so it's a reduction in persistence rather than in exposure.

How do we find out what has already left?

Start with the connectors, since they're the highest-volume route and usually the least logged. Retrieval telemetry on the knowledge base shows which documents are being returned and to whom; the self-test turns that into one number. For prompts, your provider's admin console and your DLP logs are the available record. For fine-tunes, the training manifests are the record, and whatever went into the weights is there to stay.

Sources

OpenAI, Business data privacy, security, and compliance, accessed October 2026. Anthropic, Commercial Terms of Service, accessed October 2026. Microsoft and Google publish comparable statements for Microsoft 365 Copilot and Gemini for Google Workspace; check the current versions.

Regulation (EU) 2016/679 (GDPR), Chapter V. Regulation (EU) 2024/1689 (AI Act) as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), in force 27 July 2026.

NIST, SP 800-171 Rev. 3, Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations, May 2024. 45 CFR 164.502(b) (HIPAA minimum necessary standard). 22 CFR Parts 120-130 (ITAR); 15 CFR Parts 730-774 (EAR).