Is it safe to use ChatGPT with confidential data?
Not with raw confidential data — but yes if you anonymize it first. Whatever you paste is sent to the provider’s servers, stored, and on consumer accounts used to train the model. This guide lays out the real risks with named incidents, gives you a quick risky-vs-okay decision aid, and shows the safe way to still get AI’s help on sensitive files.
What happens when you paste confidential data
Your text is transmitted to the provider and stored on its servers, and on consumer accounts (Free, Go, Plus, Pro) it’s used to train the models by default. (Source: OpenAI, accessed July 2026.) We cover exactly how retention, training and deletion work in does ChatGPT store your data? — here the question is simpler: given that, is it safe?
The real risks, ranked
- Accidental leaks (“shadow AI”). Staff paste confidential material into personal accounts, invisibly to IT.
- Model-training exposure. On consumer accounts your text trains the model by default.
- Provider & subprocessor breaches. The provider — or one of its vendors — can be hacked.
- Discoverability & legal holds. Prompts can be retained by court order and produced as evidence.
- Accidental public exposure. A wrong setting or shared link can put a conversation on the open web.
It already happened — named incidents
- Samsung (2023). Within weeks of allowing ChatGPT, engineers pasted confidential source code and internal meeting notes into it; Samsung banned generative AI on company devices.
- The March 2023 ChatGPT bug. A flaw in the redis-py library let some users see others’ chat titles and first messages, and exposed payment details — name, email, billing address, card expiry and the last four digits (never the full card) — for 1.2% of ChatGPT Plus subscribers active in a ~9-hour window. (Source: OpenAI.)
- Shared chats on Google (mid-2025). A “make this chat discoverable” toggle caused shared ChatGPT conversations to be indexed by search engines; OpenAI killed the feature within days as too risky. (Source: TechCrunch.)
- The Mixpanel subprocessor breach (Nov 2025). A breach of OpenAI’s analytics vendor exposed some API customers’ names, emails, approximate location and organization IDs — proof that even a secure provider’s vendors are an attack surface. (Source: OpenAI.)
This isn’t rare — the shadow-AI numbers
Occasional pasting adds up to a systemic leak. A security analysis of 1.6 million workers found about 11% of what employees paste into ChatGPT is confidential, and that fewer than 1% of staff cause 80% of the leaks. (Source: Cyberhaven.) A 2025 report found 68% of employees use free-tier AI through personal accounts, and 57% enter sensitive data; (Source: Menlo Security); another found generative AI is now the #1 channel for corporate-to-personal data exfiltration. (Source: LayerX.)
When is it risky vs. relatively okay?
| What you want to put in | Verdict |
|---|---|
| Names, client or patient details, any PII | Don’t paste raw — anonymize first |
| Trade secrets, source code, unreleased plans | Don’t paste |
| Privileged or regulated data (legal, health, financial) | Don’t paste — high duty of care |
| Information that is already public | Usually fine |
| Generic questions with no real data in them | Fine |
| An anonymized copy with placeholders instead of real values | Fine — this is the safe way |
The “anonymize first / fine” rows above all describe the same move: strip the identifiers locally, then paste. That’s exactly what Occlira does — anonymize on your machine, work on the copy, and restore the originals when you’re done.
A three-question gut check before you paste:
- Would a breach of this embarrass you or breach a duty (to a client, patient or employer)?
- Could you do the task just as well on an anonymized version?
- Are you on a consumer account (trained on by default) or a business tier?
Does a business tier make it safe?
It helps with one risk: Business, Enterprise, Edu and the API aren’t trained on by default and come with a data-processing agreement — see the full tier breakdown in does ChatGPT store your data? But the data still leaves your device and can still be exposed in a subprocessor breach, and a DPA doesn’t by itself make the processing lawful — that’s covered in use AI without breaking the GDPR. A business tier reduces risk; it doesn’t remove it.
The safe workflow
The rule is data minimization — don’t send what the task doesn’t need:
- Anonymize or remove the confidential parts before anything is sent.
- Run the AI tool on the anonymized copy.
- Prefer a no-training tier, and never paste credentials or regulated data.
- Keep a human in the loop and verify the output before you rely on it.
How Occlira makes it safe
Occlira is that first step. It detects and removes personal data from documents, spreadsheets, email, audio and images on your own computer — no cloud, no account. You anonymize a file locally, run any AI tool on the anonymized copy, then restore the real values on your machine. That neutralizes every risk above: nothing sensitive to paste by mistake, to train on, to leak in a breach, or to subpoena.
One honest caveat: Occlira’s anonymization is reversible, which in GDPR terms is pseudonymization — strong local minimization, but the mapping (kept on your device) is still personal data, so treat it as control, not a loophole. The step-by-step is in anonymize before ChatGPT.
Frequently asked questions
Not in its raw form. Whatever you paste is transmitted to and stored by the provider, and on consumer accounts it trains the model by default — plus it can be exposed in a breach or preserved by a court order. It is safe, though, if you anonymize the confidential parts first and run the AI on the placeholder version.
Yes. A March 2023 bug briefly exposed some users’ chat titles and payment details, in November 2025 a breach of OpenAI’s analytics vendor Mixpanel exposed some API customers’ contact details, and in mid-2025 a “make discoverable” toggle let shared chats be indexed by Google before OpenAI removed it. Each shows the data leaves your control once it’s sent.
In 2023, within weeks of allowing it, Samsung engineers pasted confidential source code and internal meeting notes into ChatGPT. Samsung responded by banning generative-AI tools on company devices — the classic “shadow AI” cautionary tale.
Not normally — but it has happened. The 2023 bug showed other users’ chat titles, and shared links were briefly indexed by search engines. More routinely, staff who review flagged content and, in litigation, opposing parties can end up seeing conversations. Treat anything you paste as leaving your control.
It lowers one risk: Business, Enterprise, Edu and the API aren’t trained on by default and come with a data-processing agreement. But the data still leaves your device and can still be exposed in a subprocessor breach — as the Mixpanel incident showed. It reduces risk; it doesn’t remove it.
Shadow AI is employees using personal AI accounts for work, outside IT’s visibility or controls. It’s widespread — reports find most AI use runs through personal accounts and a large share includes sensitive data — which is why occasional pasting adds up to a systemic leak.
Client or patient personal data, trade secrets and source code, privileged or regulated material, credentials, and anything whose exposure would breach a duty or embarrass you. If you need AI’s help with such a document, anonymize it first.
Remove the confidential parts before they’re sent: anonymize the file locally with Occlira, run any AI tool on the anonymized copy, then restore the real values on your own machine. There’s then nothing sensitive to leak, train on, or subpoena.
Use AI on sensitive files, safely
Anonymize locally, then use any AI tool on the copy. Free for 14 days on Windows and macOS.
More: what is PII? · for law firms · redact Word, PDF & Excel · how your data is handled