What ChatGPT, Claude and Gemini actually do with what you paste

Published 2026-07-29

Every few months someone claims that a chatbot “reads everything you type” and someone else replies that it “forgets immediately”. Both are wrong in the same way: they answer a question about a product with a slogan. The real answer depends on which plan you are on, which settings you have touched, and what a court has said this quarter.

Here is what the providers stated as of 29 July 2026, and — more usefully — how to work out the answer yourself when this page goes stale.

The four questions that matter

  1. Training. Is your text used to improve the model?
  2. Retention. How long is it stored after you delete it?
  3. Human review. Can a person read it?
  4. Legal hold. Can a court freeze all of the above?

They are independent. Opting out of training does not shorten retention. Deleting a conversation does not necessarily delete the copy held for abuse review. And a preservation order overrides all three.

Training: the consumer/business split

The single most important thing to know is that consumer plans and business plans are governed by different terms at every major provider.

OpenAI states that ChatGPT improves by training on your conversations unless you opt out, and that ChatGPT Business, Enterprise, Edu and API usage are excluded from training by default. Temporary Chat conversations do not appear in history, do not create memories, and are not used for training. The opt-out for a personal account lives in Settings → Data Controls.

Anthropic changed course in August 2025. Under the consumer terms published on 28 August 2025, conversations and coding sessions from Free, Pro and Max plans are used for training unless the user opts out, and the retention period for data covered by that consent extends to five years — a significant change from the previous 30-day default. Existing users were asked to choose by 28 September 2025. Claude for Work (Team and Enterprise) and API usage sit under separate commercial terms and are excluded.

Google has historically kept a similar shape for Gemini: consumer conversations may be reviewed by humans and used to improve the service, with a settings page controlling activity storage, while Workspace and Cloud usage is governed by the enterprise agreement.

The pattern is consistent enough to use as a default assumption: if you are not paying through a business agreement, assume your text trains the model unless you have personally changed the setting.

Retention: deletion is not always deletion

Providers typically keep a copy of deleted conversations for a short window for abuse and safety review — 30 days is the common figure. That is unremarkable and is disclosed in the policies.

What is less obvious is that retention can be extended by forces outside the product. In the New York Times copyright litigation, OpenAI was ordered to preserve consumer output logs, including chats users had deleted. That obligation ended on 26 September 2025, after which OpenAI returned to its standard practice of deleting conversations and Temporary Chats within about 30 days — but data preserved under the order still exists, and logs tied to accounts specifically flagged in the case remain retained. In July 2026 the publishers moved for sanctions, arguing deletion continued despite the order.

The lesson is not that any particular company behaved badly. It is structural: a chat log is a business record of a company that can be sued. Plan for that, the same way you would for email.

Human review

Every major provider allows some human review of conversations, usually for safety, abuse detection and quality evaluation, and usually on a sampled basis. Enterprise agreements narrow this; consumer terms generally permit it. If your mental model is “only a machine sees it”, correct that model — the probability is low, but it is not zero, and for genuinely confidential material a low probability of a stranger reading it is still a disclosure.

What this means in practice

Do not rely on a setting you have not personally checked. Defaults change, and they changed in 2025 for at least one major provider in the direction of more training, not less. Go into the account you actually use and look.

Do not assume a work email means a business plan. Plenty of people sign up for a personal account with a company address. The terms follow the plan, not the domain.

Assume durability for anything sensitive. Not because deletion is a lie, but because retention windows, legal holds and backups all outlive the button in the interface.

Separate the two questions. “Can I use AI for this task?” and “does the AI need the identifying details?” are different. The second one usually has a better answer than the first.

The cheap mitigation

Most of the risk in the four questions above attaches to identifiable content. A conversation about [NAME_1]’s refund for [CARD_1] is retained, trained on and legally frozen exactly like any other conversation — and none of that matters much, because nothing in it identifies anyone.

That is the whole argument for scrubbing before the paste rather than negotiating with the policy after it. Paste your text here, replace the identifiers with stable placeholders, get your answer, and restore the real values in the reply. The mapping stays in your browser tab because this site has no server that could hold it — which, unlike a privacy policy, is not something anyone can change next quarter.

How to re-check these facts yourself

Policies move. When you need the current answer:

If those four checks take you more than fifteen minutes, that is itself an answer about how much certainty you have.