Guides › Third-party AI platforms
How to Use ChatGPT, Claude, Copilot and Gemini Without Sending Sensitive Data to the Provider
Blocking the tools doesn't work, and buying the enterprise tier is only part of the answer. This guide covers what actually crosses to a model provider when your team uses a frontier AI platform, the four ways organizations try to control it, and the design that keeps the model in the loop while real identifiers stay inside your boundary.
What actually crosses the provider boundary
When someone uses an AI assistant over company data, the provider receives more than the question. The table lists the channels, in rough order of how much attention each one gets, which is close to the reverse of how much data each one carries.
Most governance effort goes to the first row. The third row is where the volume is. One answer grounded on an internal corpus can send ten or twenty document chunks to the provider, chosen by a similarity function with no notion of sensitivity, scoped by whatever permissions the connector inherited, and invisible to the person who asked. The AI knowledge base security guide covers why retrieval behaves like an access decision. Here the point is that the recipient is outside your organization.
Four ways to control it, and what each one costs
The first three are familiar. The fourth is the one this page is about, because it's the only option that keeps a frontier model in the workflow while taking real identifiers out of the provider's hands.
How substitution works
The design has four steps, and the order matters.
- Connect. The data layer connects to the systems your team already uses (SharePoint, Drive, ticketing, CRM, file shares) and inherits their permissions. Nothing is copied into a new store.
- Retrieve in scope. When someone asks a question in ChatGPT or Claude, the connector searches only what that person is cleared to see, with entitlements checked at query time. The RAG permissions guide explains why index-time permissions aren't enough.
- Substitute outbound. Before retrieved content crosses to the model provider, every real identifier is replaced with a consistent stand-in. The model receives "CUST-A4417 is 62 days past due on $184,220.00" and reasons on it without ever holding a customer's name.
- Restore inbound. The answer comes back with stand-ins in it. In the reader's own window, the stand-ins resolve to the real values, if and only if that reader holds the entitlement. A reader without it sees the stand-in and nothing else.
Every retrieval and every restoration attempt is recorded, keyed to opaque IDs, so the log itself contains no sensitive content. A pattern of denied restorations is its own signal, and the telemetry guide covers what else the record can show.
What this doesn't solve
Scope stated plainly, because a control that overclaims is worse than no control.
- Typed prompts. What a person types or pastes into the assistant goes to the provider as typed. Substitution governs retrieved content. Acceptable-use policy, training and DLP still own the prompt box.
- Personal accounts. An employee using a consumer account on a personal device is outside every enterprise control, including this one. The fix is to make the sanctioned path the easier path.
- Inference-time visibility. The provider's systems still process the substituted content. Zero-data-retention terms limit how long it persists; they don't make it invisible.
- Memorized public data. If sensitive information about your organization is already in a model's training data from public sources, no boundary control retrieves it.
A checklist for your next AI vendor review
- Which of our channels (prompts, uploads, connectors, agents, memory) reach the provider, and which team owns each?
- Is the provider a subprocessor under our data processing terms, and does our transfer basis cover it?
- What is the default retention window, and are we eligible for zero data retention?
- Is training on our inputs and outputs excluded in the contract itself, rather than only on a policy page the provider can edit?
- What permissions does each connector inherit, and when were they last reconciled with the source systems?
- Can we produce a record of which documents reached the provider for a given user and day?
- Are real identifiers replaced before content crosses the boundary, or does the provider receive them in the clear?
- Which users are on personal accounts, and why is the sanctioned path harder than the unsanctioned one?
Where Hardshell sits
Hardshell is the data layer in the fourth option. It appears as a connector inside ChatGPT and Claude, connects to the systems a team already runs, inherits their permissions, scopes retrieval to the caller, substitutes stand-ins for real identifiers before the provider boundary, restores real values only for readers who hold the entitlement, and records every retrieval. The same foundation runs exposure testing and dataset hardening ahead of deployment. The homepage shows the four steps in motion.
Frequently asked questions
Does ChatGPT Enterprise or Claude for Enterprise train on our data?
Per each provider's published terms as of October 2026, business and API tiers don't use customer inputs or outputs for training by default, and retention windows with zero-data-retention options apply to qualifying workloads. Consumer tiers differ. On some of them training on conversations is the default or an opt-in that also extends retention. Confirm the current terms in the contract rather than the policy page.
Is this the same as redaction or masking?
Redaction removes information and the reader loses it. Substitution replaces identifiers with consistent stand-ins so the model can still reason about them, then restores the real values for an entitled reader. The answer is as useful as it would have been in the clear. The provider just never held the identifiers.
Does this work with Microsoft Copilot and Gemini?
The pattern is the same for any assistant that retrieves from your data. Scope the retrieval, substitute before the boundary, restore for the reader. Platform-specific connector support varies, so ask about the platforms you run.
What about agents that call tools?
Tool results are a retrieval channel with a different name. Anything an agent fetches on a user's behalf goes to the model as context, so the same scoping and substitution apply, and agentic retrieval widens the surface because the model decides what to fetch. See indirect prompt injection for the attacker's side of that.