Guides › Sovereignty vs residency
Data Sovereignty vs Data Residency: What Each One Means for Enterprise AI
Residency is where data is stored. Sovereignty is whose laws govern it and who controls what happens to it. Localization is a legal requirement to keep it in-country. The three get used interchangeably, and AI makes the confusion expensive, because a model query moves data at inference time in ways a storage region was never designed to govern. This page defines each term, explains why AI breaks the residency answer, maps the regulations that are usually invoked, and lists the questions to ask a vendor.
Three terms, three questions
The ordinary trap is to answer the residency question and present it as the sovereignty answer. A hyperscaler region in Frankfurt or Sydney settles where the bytes sit. If the provider is headquartered in the United States, the CLOUD Act lets US authorities compel production of data the provider controls wherever it is stored, and EU data protection authorities have said a CLOUD Act request is not by itself a lawful basis for a transfer under GDPR. That tension is structural, and no region setting resolves it. The hyperscalers' European sovereign cloud offerings, including the AWS European Sovereign Cloud launched in January 2026, try to answer it with separate corporate entities and EU-resident operations. The European Commission's proposed Cloud and AI Development Act, published in June 2026, would formalize tiers of sovereignty assurance that explicitly weigh exposure to third-country law. It isn't law yet.
Why AI breaks the residency answer
Residency controls were designed for data at rest. An AI system touches data in three motions that residency never covered.
- Inference. Every prompt and every piece of retrieved context is processed by the model wherever the model is served. Regional storage of your tenant says nothing about regional inference unless the provider commits to it separately, and several providers offer regional storage with inference that can route elsewhere.
- Retrieval. A connector or knowledge base pulls documents from a store in one jurisdiction and hands them to a model as context. The documents never changed residency. Copies of them crossed the boundary in the query path. See the third-party platforms page.
- Training and fine-tuning. Data copied into a provider's training pipeline is encoded into weights that are served from wherever the model runs, and can't be deleted from the model afterward. The training data leakage guide covers what can be recovered from weights.
There is a fourth motion worth naming. Provider telemetry, abuse monitoring and support tooling sit with the provider's own infrastructure and subprocessors, and retention terms rather than residency settings govern them.
The regulations usually invoked, and which question each one asks
Two things are true at once. None of these regimes mentions retrieval context, and every one of them applies to it, because retrieval context is a copy of regulated data crossing a boundary. If you need a given failure mode expressed in a specific framework's language, the crosswalk covers MITRE ATLAS, OWASP, NIST AI RMF, ISO/IEC 42001 and the EU AI Act.
Questions to ask a vendor about AI data sovereignty
- Where is inference performed for our tenant, and is that committed in writing or only the storage region?
- Which model providers are subprocessors, where are they domiciled, and what law can reach them?
- What crosses to the model on each request: the prompt only, or retrieved context, files and tool results as well?
- Can we see a per-request record of what was sent and for whom?
- Are real identifiers removed or replaced before the boundary, or sent in the clear?
- What is the retention window at the provider, and can it be set to zero for our workloads?
- Can the service run inside our boundary, on our private cloud or disconnected, with no outbound dependency?
- If we leave, what is deleted, what is certified deleted, and what was already encoded into a model?
Where Hardshell sits
Hardshell keeps data where it lives and brings a scoped, substituted view of it to the model, so the residency of the source systems is preserved and what crosses to a provider is recorded per request. It deploys as a cloud service, inside a customer's private cloud, or as an offline bundle, which lets the architecture follow whichever of the three questions your regulator is asking. The main data sovereignty guide covers the control model in full.
Frequently asked questions
If our data is stored in an EU region, is it sovereign?
Stored in the EU, yes. Sovereign, only if the provider can't be compelled by a third country and your application doesn't send copies elsewhere at query time. Region selection answers the first question and neither of the other two.
Does the EU AI Act require data residency?
No. It imposes governance, logging and documentation duties, and after the 2026 amendment the Annex III high-risk obligations apply from 2 December 2027. Residency obligations for personal data come from GDPR's transfer rules, and sector rules may add more.
Is a self-hosted model the only way to be fully sovereign?
It's the cleanest answer to the jurisdictional question for workloads where an open-weight model is sufficient. For workloads that need a frontier model, the alternative is to control what crosses. Scoped retrieval, substitution before the boundary, and a record of every request. See sovereign AI deployment.
We're a US company with no EU data. Does any of this apply?
The sovereignty question still applies in the operational sense. HIPAA, GLBA, state privacy laws, ITAR and CUI rules all constrain what can reach a third-party service and who can see it, and none of them care which continent the servers are on. The main guide covers the US regimes.