How to anonymize documents before pasting them into ChatGPT
The safest way to use AI on a sensitive file is to remove the personal data first, run the AI on an anonymized copy, and restore the real values afterwards — all on your own computer. Occlira does exactly this for documents, spreadsheets, email and audio, so nothing private ever leaves your machine.
Why pasting raw data is riskier than it looks
It feels harmless — you copy a paragraph from a contract or a patient note into a chat box to get a quick draft. But that text is sent to the provider’s servers, and it doesn’t always stop there. A security analysis by Cyberhaven found that about 11% of everything employees paste into ChatGPT is confidential company data, and the share of sensitive material flowing into AI tools has been climbing sharply as usage grows. (Source: Cyberhaven Labs.)
This is not hypothetical. Within roughly three weeks of allowing ChatGPT, Samsung engineers leaked confidential material three separate times — semiconductor source code, defect-detection routines and an internal meeting transcript — pasted in to fix bugs and write up notes. The company banned generative AI on its devices soon after. (Source: Forbes.)
Three things people get wrong about ChatGPT and privacy
The reassurances people rely on don’t hold up as well as they think:
- “I’ll just delete the chat.” Deleting is a request, not a guarantee. In May 2025 a US court ordered OpenAI to preserve ChatGPT logs — including chats users had already deleted — as evidence in the New York Times case. The order was later narrowed, but the lesson stands: once your text is on someone else’s server, you no longer decide when it’s gone. (Source: OpenAI.)
- “It’s just for me, it won’t be used to train anything.” On Free, Plus and Pro accounts, model training is on by default — you have to switch off “Improve the model for everyone” in Data Controls. Most workplace ChatGPT use runs through exactly these personal accounts. (Source: OpenAI Help Center.)
- “We signed a business agreement, so we’re covered.” An enterprise plan or a data-processing agreement reduces exposure, but the personal data still leaves your control and lands on a third party’s infrastructure — subject to their security, their subprocessors and orders like the one above. For anyone bound by professional secrecy, that’s a different risk from keeping the data on your own machine.
Each of these is an argument for the same fix: don’t send the personal data in the first place — which is exactly what Occlira’s anonymize → use-AI → restore workflow does, on your own machine. Read more in does ChatGPT store your data?
The safe workflow: anonymize → use AI → restore
- Open your file in Occlira. Drop in the document, spreadsheet, email or audio file you were about to send to an AI tool. Nothing uploads — detection runs on your own machine, so the file never leaves your computer even to be scanned.
- Review what gets removed. Occlira highlights every name, email, phone number, ID, financial detail and more. You keep or redact each item in a quick review step, and can select any extra text the detector missed and mark it too. You stay in control of exactly what is replaced.
- Anonymize. Real values are swapped for consistent, neutral placeholders — <PERSON_1>, <IBAN_1>, <DATE_1>. Because the same person keeps the same tag throughout, the text still reads coherently. A mapping of each placeholder to its original value is saved locally so the change is reversible.
- Use the AI tool. Paste or upload the anonymized version into ChatGPT, Claude, Gemini or Copilot and do your work. The model reasons over the placeholders just as well as over real names, so you get a usable answer — draft, summary, analysis — without the personal data ever reaching the provider.
- Restore the originals. Bring the AI’s output back into Occlira and restore the real values from your local mapping, so your finished document is complete and accurate. The restore happens entirely on your device.
What to anonymize before pasting — a quick checklist
The goal is to strip anything that identifies a person, directly or in combination. Work through this list — it maps to the categories Occlira detects automatically (40+ types in total):
- Names of people — clients, patients, employees, opposing parties, witnesses
- Contact details — email addresses, phone numbers, postal addresses
- Government and national IDs — SSN, NINO, passport, tax ID, Codice Fiscale, DNI, INSEE
- Financial identifiers — IBANs, account and card numbers, salary figures
- Health data — diagnoses, medical record numbers, anything that reveals a condition
- Dates that pin down a person — dates of birth, admission or hearing dates
- Case, matter, invoice and reference numbers that map back to a real file
- Quasi-identifiers — a rare job title plus a city, or any detail unique enough to re-identify someone on its own
What Occlira removes
Names, email addresses, phone numbers, postal addresses, IBANs and card numbers, national IDs (SSN, NINO, VAT, passports and more), dates of birth, IP addresses and medical record numbers — across Word, PDF, Excel, email and plain text, plus audio recordings. Detection runs locally with an AI model backed by precise pattern rules, and you confirm every item in the review step, so nothing is redacted — or missed — without your say-so.
Is anonymizing enough to stay compliant?
It removes the hardest part of the problem. When you paste the anonymized version, no identifiers reach the provider — so there is no personal data for a third party to store, train on, or be ordered to hand over. That is the data-minimization principle regulators expect you to apply before using any external tool. Use AI without breaking the GDPR goes into the practical pattern.
Two honest caveats. First, the mapping that lets you restore the originals is itself personal data — but it stays on your device, so your processing is minimal and under your control, not scattered across a vendor’s cloud. Second, EU regulators set a high bar for calling anything truly “anonymous”: the EDPB’s Opinion 28/2024 stresses a case-by-case assessment. Removing identifiers before you share is one strong control — not a substitute for your organisation’s full obligations.
Reversible by design: pseudonymization, not a black box
Because Occlira keeps a local mapping, its anonymization is reversible — in GDPR terms,
that’s pseudonymization. This is the practical bit: the placeholders are consistent (the same person is
<PERSON_1> everywhere), so the AI’s answer still lines up with your real data when you restore
it. You get the usefulness of working on the real document without the real document ever being exposed.
Does it work with Claude, Gemini and Copilot?
Yes. The approach is tool-agnostic. Occlira produces an anonymized copy you can use with any AI assistant; because the placeholders are consistent, you still get a useful answer, and you restore the real values on your own machine when you are done.
Frequently asked questions
Treat anything you paste into a cloud AI tool as leaving your control. On personal accounts your chats can be used to train the model unless you opt out, and a 2025 US court order showed that even deleted chats can be preserved. The safe approach is to remove the personal data first, work with the AI on the placeholder version, and restore the real values afterwards on your own machine.
No. Models work fine on consistent placeholders such as <PERSON_1> or <IBAN_1> — the same entity keeps the same tag, so the text stays coherent. You get a usable draft, summary or analysis, then restore the real values locally so the final document is complete.
Not from the placeholders alone — a tag like <PERSON_1> carries no personal information, and the mapping back to the real name never leaves your computer. Just avoid leaving obviously identifying context (a unique job-title-plus-location combination, for example) that a reader could resolve without the mapping.
On Free, Plus and Pro accounts, model training is on by default — you have to turn off “Improve the model for everyone” in Data Controls to stop it. Temporary Chat and business products (Team, Enterprise, the API) are not used for training. But the cleanest protection is simply not to send the personal data at all.
Not necessarily. In May 2025 a US court ordered OpenAI to preserve ChatGPT logs — including chats users had deleted — as evidence in the New York Times case. The scope was later narrowed, but it’s a concrete reminder that once your text is on another company’s servers, “delete” is a request, not a guarantee.
It removes the hardest part of the problem: if no names, IDs or other identifiers are sent, there is no personal data for a third party to store, train on or be compelled to disclose, which is exactly the data-minimization principle regulators expect. The local mapping is itself personal data, but it stays on your device — so you keep processing to a minimum and under your own control. It is one strong control, not a substitute for your organisation’s full compliance obligations.
Yes. The workflow is tool-agnostic: Occlira produces an anonymized file you can use with any AI assistant, then you restore the originals locally.
Same principle, and it matters more: files carry hidden metadata, tracked changes and comments on top of the visible text. Occlira anonymizes the document itself — including that hidden data — so the copy you upload is clean, not just the part you can see.
Nowhere. Detection, anonymization, the placeholder mapping and restore all run on your computer. Your files and the personal data in them never reach us or any cloud service.
Keep sensitive files private with AI
Anonymize once, use any AI tool, restore locally. Free for 14 days on Windows and macOS.
More: redact personal data from audio · local vs cloud PII redaction · for law firms · how your data is handled