Guides › Third-party AI platforms

How to Use ChatGPT, Claude, Copilot and Gemini Without Sending Sensitive Data to the Provider

Blocking the tools doesn't work, and buying the enterprise tier is only part of the answer. This guide covers what actually crosses to a model provider when your team uses a frontier AI platform, the four ways organizations try to control it, and the design that keeps the model in the loop while real identifiers stay inside your boundary.

At a glance
  • What crosses. Typed prompts, uploaded files, retrieved context from connectors, tool results and, on some products, saved memory. Retrieval context is the largest and least visible stream.
  • What enterprise terms fix. Training exclusion and a defined retention window. They don't reduce what is sent.
  • The four options. Block, contract, self-host, or change what the model receives. Only the last keeps frontier capability and reduces exposure at the same time.
  • The design. Retrieval scoped to the caller, real values replaced with stand-ins before the provider boundary, real values restored only in the reader's window, every retrieval recorded.

What actually crosses the provider boundary

When someone uses an AI assistant over company data, the provider receives more than the question. The table lists the channels, in rough order of how much attention each one gets, which is close to the reverse of how much data each one carries.

ChannelWhat the provider receivesWho decides
Typed promptThe question, plus anything pasted into itThe user
Uploaded filesFull documents, spreadsheets and images, including content the user never readThe user
Retrieval contextChunks of internal documents returned by a knowledge base or connector to ground the answer; several per query, selected by similarity rather than sensitivityThe application team that built or enabled the connector
Tool and agent resultsOutput of any system an agent calls on the user's behalf, such as a CRM lookup or a database queryThe agent's builder
Memory and historyPrior conversations the product keeps to personalize later ones, on products that offer itAdmin settings, where they exist
MetadataIdentity, timestamps, usage telemetryThe provider's terms

Most governance effort goes to the first row. The third row is where the volume is. One answer grounded on an internal corpus can send ten or twenty document chunks to the provider, chosen by a similarity function with no notion of sensitivity, scoped by whatever permissions the connector inherited, and invisible to the person who asked. The AI knowledge base security guide covers why retrieval behaves like an access decision. Here the point is that the recipient is outside your organization.

Four ways to control it, and what each one costs

ApproachKeeps frontier modelsReduces what crossesCovers personal accountsWhat it costs
Block the tools (DLP rules, network policy, banned apps)NoOnly on managed devicesNo. Usage moves to phones and home machinesLost productivity, no visibility
Rely on enterprise terms (no training, retention window, zero data retention)YesNoNoContract negotiation; unverifiable from your side
Self-host an open-weight modelNoYes, entirelyNoServing cost, capability gap, you own the stack
Change what the model receives (scoped retrieval and substitution)YesYes, for everything that flows through the connectorPartially; see scope belowA data layer between your systems and the model

The first three are familiar. The fourth is the one this page is about, because it's the only option that keeps a frontier model in the workflow while taking real identifiers out of the provider's hands.

How substitution works

The design has four steps, and the order matters.

  • Connect. The data layer connects to the systems your team already uses (SharePoint, Drive, ticketing, CRM, file shares) and inherits their permissions. Nothing is copied into a new store.
  • Retrieve in scope. When someone asks a question in ChatGPT or Claude, the connector searches only what that person is cleared to see, with entitlements checked at query time. The RAG permissions guide explains why index-time permissions aren't enough.
  • Substitute outbound. Before retrieved content crosses to the model provider, every real identifier is replaced with a consistent stand-in. The model receives "CUST-A4417 is 62 days past due on $184,220.00" and reasons on it without ever holding a customer's name.
  • Restore inbound. The answer comes back with stand-ins in it. In the reader's own window, the stand-ins resolve to the real values, if and only if that reader holds the entitlement. A reader without it sees the stand-in and nothing else.

Every retrieval and every restoration attempt is recorded, keyed to opaque IDs, so the log itself contains no sensitive content. A pattern of denied restorations is its own signal, and the telemetry guide covers what else the record can show.

What the provider sees

Same question, two views. Your analyst's window reads: "Marcus T. Whitmore is the largest East-region exposure in the Q3 aging, $184,220.00 open, 62 days past due. Account owner is Dana Okafor." The model provider's logs hold: "CUST-A4417 is the largest East-region exposure in the Q3 aging, $184,220.00 open, 62 days past due. Account owner is OWNER-2210." The figures and the reasoning are intact. The identifiers never left.

What this doesn't solve

Scope stated plainly, because a control that overclaims is worse than no control.

  • Typed prompts. What a person types or pastes into the assistant goes to the provider as typed. Substitution governs retrieved content. Acceptable-use policy, training and DLP still own the prompt box.
  • Personal accounts. An employee using a consumer account on a personal device is outside every enterprise control, including this one. The fix is to make the sanctioned path the easier path.
  • Inference-time visibility. The provider's systems still process the substituted content. Zero-data-retention terms limit how long it persists; they don't make it invisible.
  • Memorized public data. If sensitive information about your organization is already in a model's training data from public sources, no boundary control retrieves it.

A checklist for your next AI vendor review

  • Which of our channels (prompts, uploads, connectors, agents, memory) reach the provider, and which team owns each?
  • Is the provider a subprocessor under our data processing terms, and does our transfer basis cover it?
  • What is the default retention window, and are we eligible for zero data retention?
  • Is training on our inputs and outputs excluded in the contract itself, rather than only on a policy page the provider can edit?
  • What permissions does each connector inherit, and when were they last reconciled with the source systems?
  • Can we produce a record of which documents reached the provider for a given user and day?
  • Are real identifiers replaced before content crosses the boundary, or does the provider receive them in the clear?
  • Which users are on personal accounts, and why is the sanctioned path harder than the unsanctioned one?

Where Hardshell sits

Hardshell is the data layer in the fourth option. It appears as a connector inside ChatGPT and Claude, connects to the systems a team already runs, inherits their permissions, scopes retrieval to the caller, substitutes stand-ins for real identifiers before the provider boundary, restores real values only for readers who hold the entitlement, and records every retrieval. The same foundation runs exposure testing and dataset hardening ahead of deployment. The homepage shows the four steps in motion.

Frequently asked questions

Does ChatGPT Enterprise or Claude for Enterprise train on our data?

Per each provider's published terms as of October 2026, business and API tiers don't use customer inputs or outputs for training by default, and retention windows with zero-data-retention options apply to qualifying workloads. Consumer tiers differ. On some of them training on conversations is the default or an opt-in that also extends retention. Confirm the current terms in the contract rather than the policy page.

Is this the same as redaction or masking?

Redaction removes information and the reader loses it. Substitution replaces identifiers with consistent stand-ins so the model can still reason about them, then restores the real values for an entitled reader. The answer is as useful as it would have been in the clear. The provider just never held the identifiers.

Does this work with Microsoft Copilot and Gemini?

The pattern is the same for any assistant that retrieves from your data. Scope the retrieval, substitute before the boundary, restore for the reader. Platform-specific connector support varies, so ask about the platforms you run.

What about agents that call tools?

Tool results are a retrieval channel with a different name. Anything an agent fetches on a user's behalf goes to the model as context, so the same scoping and substitution apply, and agentic retrieval widens the surface because the model decides what to fetch. See indirect prompt injection for the attacker's side of that.

Sources

OpenAI, Business data privacy, security, and compliance, accessed October 2026: no training on business or API data by default; data retention controls and a zero-data-retention option for qualifying organizations; regional data residency for eligible customers.

Anthropic, Commercial Terms of Service and Privacy Center, accessed October 2026: no training on commercial customer content unless the customer opts in; zero data retention available for qualifying use.

Microsoft and Google publish comparable data-handling statements for Microsoft 365 Copilot and Gemini for Google Workspace. Terms change; confirm the version in force when you sign.