Answers

Is DeepSeek safe to use? Data in China, training, and what never to paste (2026)

Published 3 September 2026 · Updated 4 September 2026 · Occlira team

Is DeepSeek safe to use? For a throwaway question, it is about as safe as any free chatbot; for anything confidential, it is not. Its privacy policy says the company collects, processes and stores your personal data in the People’s Republic of China. Your prompts, uploads and chat history feed model training unless you switch that setting off. Its Terms of Use put every dispute under PRC law and cap the company’s liability at the greater of what you paid or $100. The models are a different matter: the weights are MIT-licensed, and a model run on your own machine keeps nothing you type. Below: the app and the model, separately, with the documents quoted as they read today.

Who this is from. Occlira is a desktop app that replaces the names, IDs and addresses in a file with placeholders on your own computer, so the copy you paste into DeepSeek carries placeholders instead of the details it found. This guide is what we tell our own users. How it works ↓ · Free 14-day trial

One correction to the usual advice. Much of the coverage still says DeepSeek trains on everything with no way out; on the documents live today that is half right. The privacy policy of 10 February 2026 lists, among your rights, “the right to opt-out of using your Personal Data for training our models or optimizing our technologies”. The Terms of 27 March 2026 name the switch: “Improve the model for everyone”. What no setting changes is the location or the jurisdiction: the data sits in the PRC, and PRC law governs disputes.

Jump to: where your data goes · training and opt-out · PRC law and the terms · who banned it · running it locally · using it without the names · FAQ

Where does DeepSeek store your data? What the DeepSeek privacy policy says

DeepSeek does not bury it. Its privacy policy, last updated 10 February 2026, states: “To provide you with our services, we directly collect, process and store your Personal Data in People’s Republic of China.” The sentence immediately before it is the warning that leads there: “The Personal Data we collect from you may be stored on a server located outside of the country where you live.” (Source: DeepSeek Privacy Policy, accessed 3 September 2026.)

What travels there is more than the text of your prompts. The policy’s own categories cover account details; User Input — text and voice prompts, uploaded files and chat history; IP-based location; cookies on the web; payment data for open-platform customers; and data collected automatically. That last category “includes your device model, operating system, IP address, device identifiers and system language”, alongside log data on “the features you use and the actions you take”.

What the policy does not say matters as much. It sets no retention period: data is kept as long as necessary to provide the services and for the other purposes it lists. The period varies with the amount, type and sensitivity of the data, and with legal requirements. It promises appropriate safeguards for transfers out of certain countries and carries a supplement for the EEA, the UK and Switzerland — but the destination in the main text is not conditional: China is where the service runs.

Does DeepSeek train on your data — and can you opt out?

Yes, and yes. The policy lists as a purpose of processing: “To improve and develop the Services and to train and improve our technology, such as our machine learning models and algorithms.” Among your rights it lists “the right to opt-out of using your Personal Data for training our models or optimizing our technologies”. The Terms of Use of 27 March 2026 name the control that implements it: inputs and outputs can be used to operate and improve the services unless you have disabled the option called “Improve the model for everyone”. Because the permission runs until you disable it, it is opt-out, not opt-in. (Sources: Privacy Policy; Terms of Use, both accessed 3 September 2026.)

Be clear about what the switch buys you. It changes what your data is used for: it does not delete anything, does not shorten a retention period the policy never states, and does not move a byte out of the PRC. Neither document says where the setting sits in the app, and we did not test the UI. Nor do they say whether an opt-out reaches conversations already used for training — assume it does not, because a model that has trained on your text cannot untrain on request.

Who else can reach it: PRC law, and what the Terms sign you up to

This is where the country of storage stops being trivia. Article 7 of China’s National Intelligence Law provides that “all organizations and citizens shall support, assist, and cooperate with national intelligence efforts in accordance with law, and shall protect national intelligence work secrets they are aware of.” How far it reaches in practice is contested among specialists — but a company assessing a processor cannot assume the article is decorative. (Source: English translation as reproduced by Wikipedia’s article on the law, which also documents the disagreement among specialists; accessed 3 September 2026.)

The Data Security Law cuts the other way, from your side of the transaction. Article 36 provides that “domestic organizations and individuals must not provide data stored within the mainland territory of the PRC to the justice or law enforcement institutions of foreign countries without the approval of the competent authorities of the PRC.” Any route you would use at home to compel disclosure or deletion runs into that provision. (Source: English translation as reproduced by Wikipedia’s article on the law, accessed 3 September 2026.)

The Terms of Use point the same way. Disputes: “The establishment, execution, interpretation, and resolution of disputes under these Terms shall be governed by the laws of the People’s Republic of China in the mainland.” Recourse: the service is provided “on an ‘as is’ and ‘as available’ basis”, aggregate liability is capped at “the greater of the amount you paid for the service… or one hundred dollars ($100)”, and “you should NOT treat any Outputs as professional advice”. The one clause that runs in the user’s favour is on ownership — “we assign any rights, title, and interests — if any — in the Outputs of the Services to you.” (Source: DeepSeek Terms of Use, 27 March 2026, accessed 3 September 2026. The liability clauses are capitalised in the original; we quote them in lower case.)

What went wrong in early 2025: an open database and an app sending data unencrypted

The findings that shaped DeepSeek’s reputation are both over a year old. In late January 2025 Wiz Research reported “a publicly accessible ClickHouse database belonging to DeepSeek, which allows full control over database operations, including the ability to access internal data”, exposing over a million lines of log streams. The log table held “plaintext logs, including Chat History, API Keys, backend details, and operational metadata”. Wiz found it with “straightforward reconnaissance techniques” that identified “around 30 internet-facing subdomains”. It disclosed the exposure, and says DeepSeek “promptly secured” it. (Source: Wiz Research, 29 January 2025.)

A week later NowSecure published an analysis of the iOS app. It reported that the app “transmits sensitive data over the internet without encryption, making it vulnerable to interception and manipulation”, and that it “uses outdated Triple DES encryption, reuses initialization vectors, and hardcodes encryption keys, violating best security practices”. User data, it added, “is transmitted to servers controlled by ByteDance, raising concerns over government access and compliance risks”. (Source: NowSecure, 6 February 2025.)

The database was closed after disclosure; we found no published retest of the iOS app and no separate DeepSeek breach dated 2026.

Which countries and regulators have banned or restricted DeepSeek?

Country / regulatorWhat was doneStatus on 3 September 2026
Italy — Garante per la protezione dei dati personaliOn 30 January 2025 the authority ordered an urgent limitation on the processing of Italian users’ data by the Hangzhou and Beijing DeepSeek entities. It had found their answers to its questions “del tutto insufficiente” (entirely insufficient), and it opened an investigation alongside the order.Open. We found no confirmed fine and no formal closure as of 3 September 2026.
Czech Republic — NÚKIB (national cyber and information security agency)A formal warning of 10 July 2025 “regarding a cybersecurity threat associated with the use of products, applications, solutions, websites, and web services… provided by the company DeepSeek”. It gives three grounds: inadequate protection of transmitted data, collection extensive enough to allow de-anonymization, and DeepSeek’s subjection to PRC legal frameworks that give state authorities access to data.In force; no withdrawal found.
Australia — Department of Home AffairsA direction under the Protective Security Policy Framework, signed 4 February 2025, required existing installations to be removed and blocked future access on non-corporate Commonwealth systems and devices. Secretary Stephanie Foster: “…the use of DeepSeek products, applications and web services poses an unacceptable level of security risk to the Australian Government.”No amendment found; treat as still in force.
Taiwan — Ministry of Digital AffairsOn 2 February 2025 the ministry said government agencies and critical infrastructure “should not use DeepSeek, because it ‘endangers national information security’”. Reported scope: public-sector staff, public schools, state-owned enterprises and critical-infrastructure operators.In force; no change found.
Netherlands — government, plus a warning from the APCivil servants were banned from using the app in February 2025 over the risk of sensitive information reaching Chinese servers. The data-protection authority told the public that “people would do well to ask themselves whether they are really willing to upload sensitive information about themselves and others”.In force for civil servants; no ban on the public.
Germany — Berlin Commissioner for Data ProtectionAsked DeepSeek in May 2025 to withdraw its apps from the German stores. When it did not, the authority reported the app to Apple and Google on 27 June 2025 as unlawful content under Article 16 of the Digital Services Act.Unresolved. Both companies refused to block the app, according to a June 2026 report; it remains available in the German stores.
South Korea — PIPCDownloads were pulled from the Korean app stores in early 2025 while the commission examined the company’s cross-border transfers.Downloads resumed on 28 April 2025, after DeepSeek published a revised, Korean-language version of its policy.
United States — New York StateOn 10 February 2025 Governor Hochul announced a statewide ban prohibiting the DeepSeek application from being downloaded on ITS-managed government devices and networks.In force. We looked for comparable bans in other US states and federal agencies and could not confirm any from a primary source on 3 September 2026.

Sources: Garante press release, 30 January 2025; NÚKIB warning, 10 July 2025; TechXplore on the Australian direction, 4 February 2025; Taipei Times, 2 February 2025; DutchNews, 6 February 2025; Berlin Commissioner for Data Protection, 27 June 2025, and iphone-ticker.de on the refusal, June 2026; The Korea Herald, 28 April 2025; Office of the Governor of New York, 10 February 2025. All checked 3 September 2026.

Almost all of this is about government devices. Italy is the exception, where the Garante’s order reached the processing of ordinary users’ data; none of the other measures stop the public from installing the app, so reading “DeepSeek is banned in country X” as “you may not use it” is usually wrong. And a regulator wanting an app gone does not make it go: Berlin asked DeepSeek to withdraw, then used the DSA route to Apple and Google; by a June 2026 report both had refused. We found no EU-wide ban as of 3 September 2026.

DeepSeek’s API and “business” tier: what the terms actually say

With OpenAI and Anthropic, moving to the API or a business plan changes the data terms. With DeepSeek, as far as the published documents go, it does not. The Open Platform Terms of Service (released 22 April 2026, effective 29 April 2026) state plainly that “the Open Platform Service and the chat service use the same account”. They are generous on what you may do with the output — you “may apply the Inputs and Outputs of the Services to a wide range of use cases, including… training other models (such as model distillation)”. (Source: DeepSeek Open Platform Terms of Service, accessed 3 September 2026.)

The gap is on the other side of the ledger. As of 3 September 2026 we found no clause in those terms addressing whether API inputs and outputs are used for training, no retention period, and no separate platform privacy policy. On the documents available, the general privacy policy above governs API customers too: same storage location, same training default, same silence on retention. For a business the answer to “can we put this in the stack” is: not a distinct privacy regime, and any personal data sent through it is a transfer to a third country your own rules must cover.

DeepSeek vs ChatGPT vs Claude

ServiceWhere the provider says the data is storedTrained on by default?How long it is kept
DeepSeek — app, website and APIThe People’s Republic of China: the policy says DeepSeek “directly collect[s], process[es] and store[s] your Personal Data in People’s Republic of China”Yes, unless you switch off “Improve the model for everyone”No period stated: as long as necessary for the purposes set out in the policy
ChatGPT — Free, Go, Plus, ProOpenAI’s own servers and its cloud providers, mainly in the United States (this row follows our own sourced page, checked 30 August 2026)Yes, unless you switch off “Improve the model for everyone”Until you delete a chat, then purged within 30 days, with exceptions
Claude — Free, Pro, MaxNot stated on the Anthropic pages cited hereSet by the model-training choice Anthropic required every consumer user to make by 8 October 2025Five years if you allow training, 30 days if you decline; incognito chats are “not used to improve Claude, even if you have enabled Model Improvement”

Sources: DeepSeek’s privacy policy and terms; for ChatGPT, the OpenAI documents quoted on our own does ChatGPT save your data page (checked 30 August 2026); for Claude, Anthropic’s consumer-terms update and its model-training article, plus its API data-retention page for the API sentence quoted below (accessed 3 September 2026).

The second row decides most of it. All three keep what you send, the consumer tier of each is governed by a training setting, and none of them is private from the provider; what differs is the legal system around the copy. Anthropic states that on the API “retained data is never used for model training without your express permission”, and OpenAI does not train on business and API data by default; DeepSeek’s API runs on the same account and the same policy as its chat. One wording trap: DeepSeek and OpenAI give the training switch the same name, “Improve the model for everyone”, so an opt-out guide written for one chatbot reads as though it works for the other. The detail on the other two: is Claude safe? and Is ChatGPT private?

Running DeepSeek locally: the model is not the app

Everything above is about a service. The models are a separate object, and the licence is unusually liberal: the official model card states that “this code repository and the model weights are licensed under the MIT License”. You can download them from Hugging Face, run them through Ollama or LM Studio, or rent them inside someone else’s cloud. Microsoft announced in January 2025 that DeepSeek-R1 “is now available via a serverless endpoint through the model catalog in Azure AI Foundry”, and AWS offers it on Bedrock, where “customers retain full control over their data and can set up safeguards such as Amazon Bedrock Guardrails”. (Sources: DeepSeek-R1 model card; Microsoft Azure blog, 29 January 2025; AWS. Neither vendor draws the contrast with DeepSeek’s own app; that follows from who runs the model and where.)

Run the weights on your own hardware, with the runtime working offline, and the privacy question largely dissolves: there is no vendor policy left to trust, because the text stays on the machine. Three things it does not fix.

  • A distill is not R1. The full model has 671B parameters, 37B of them activated. The versions that fit on ordinary hardware are distills — “DeepSeek-R1-Distill models are fine-tuned based on open-source models, using samples generated by DeepSeek-R1” — Qwen-based at 1.5B, 7B, 14B and 32B, Llama-based at 8B and 70B. Test the hosted service, deploy a 7B build, and you are not running what you tested.
  • Content behaviour travels with the weights. In early 2025 Enkrypt AI put 300 geopolitical questions to the hosted model and reported that “the censorship rate for DeepSeek-Chat soared to 88%, meaning nearly 9 out of 10 questions on certain sensitive incidents were effectively refused”, with 114 of 125 China-related answers tilted toward the Chinese perspective. The Dispatch got “Sorry, that’s beyond my current scope. Let’s talk about something else” on Tiananmen Square. Both tests ran against the hosted service and we found no 2026 refresh, but whatever behaviour was trained into the weights travels with them, while service-side filters may not. (Sources: Enkrypt AI, 30 January 2025; The Dispatch, 12 February 2025.)
  • You inherit the security work. A self-hosted endpoint means your patching, your access control and your logs — work the hosted service was doing for you.

What never to paste into DeepSeek

Sort by consequence rather than by category: if sending the same text by email to an unknown company abroad would be a reportable event, it is not a prompt. That rules out client and patient files, employee and candidate records, ID and account numbers, unreleased financials, contracts under NDA, credentials and source code containing them, and anything covered by professional secrecy or a data-processing agreement you signed with someone else.

Which leaves the awkward middle: the document you need help with that happens to have names in it. Every control in this article belongs to DeepSeek — the training toggle, the retention it does not specify, the jurisdiction. The one that stays with you is what goes into the box.

How Occlira keeps the identifiers out of the prompt

Occlira is a desktop app for Windows and macOS (Apple Silicon) that does the removal on your own computer: detection runs on-device, and the app reaches the network only to activate the licence (via Polar), download its model and check for updates. Open a document, spreadsheet, email, PDF, scan, photo or audio recording and it flags the personal data in it — names, addresses, phone numbers, ID and account numbers, dates — with a confidence score on each, so you decide what goes; anything it missed you can add by selecting it. What you get back is an anonymized copy of the file.

Occlira Deanonymize screen: an anonymized .docx selected, the session auto-detected from the file metadata, a Restore real data button, and a warning that the restored file will contain real personal data and should stay on this machine.
The way back for chats without an extension: save the reply as a file, run Deanonymize, and the real values are restored on your own machine.

From there the route is the same for every assistant, DeepSeek included: press “Copy anonymized text” and paste it into the chat, or upload the anonymized file with your prompt. You work with the placeholders, the reply comes back with them still in it, and Deanonymize restores the real values in one click — save the reply as a file, run it through, and you get a *_restored copy on your own machine. Word and Excel files stay Word and Excel files, formatting intact, and lose their hidden layer (comments, tracked changes, document properties); PDFs, emails and scans come back as anonymized .txt files. Values become consistent placeholders such as <PERSON_1>, and the mapping that reverses them stays on your device, readable only by your account and deleted after seven days by default. The connectors that do this inside the tool itself — the browser extension, the Claude Desktop connector, the Word add-in — cover ChatGPT, Claude, Gemini and Word, not DeepSeek, which is why here you work on the copy the app gives you.

The honest limits. Because that mapping exists, this is pseudonymization, not anonymization: it reduces what you expose, it does not take the data outside the GDPR, and it does not by itself make a service that stores data in the PRC compliant. The rest of the text still goes there, so read what is left before you send it. It helps; it does not guarantee compliance. Step by step: anonymize text before ChatGPT, which works the same way for any chatbot.

Frequently asked questions

For a throwaway question, about as safe as any free chatbot; for anything confidential, no. The privacy policy says DeepSeek collects, processes and stores your personal data in the People’s Republic of China; prompts, uploads and chat history train its models unless you switch off “Improve the model for everyone”; and the Terms of Use put disputes under PRC law and cap liability at the greater of what you paid or $100.

In China. The privacy policy states: “To provide you with our services, we directly collect, process and store your Personal Data in People’s Republic of China.” What goes there: your account details, your inputs (text and voice prompts, uploaded files, chat history), IP-based location and automatically collected device data. The policy sets no fixed retention period.

It trains by default, and there is an opt-out — the claim repeated online that there is none does not match the current documents. The privacy policy lists “the right to opt-out of using your Personal Data for training our models or optimizing our technologies”, and the Terms of Use name the control: “Improve the model for everyone”. Opting out changes what your data is used for, not where it sits.

Mostly on government devices rather than for the public. Australia removed it from Commonwealth systems (4 February 2025), Taiwan barred it for public-sector bodies (2 February 2025), the Netherlands banned it for civil servants, the Czech NÚKIB warned against it (10 July 2025), and New York banned it on state devices (10 February 2025). Italy went furthest: the Garante ordered a limitation on processing Italian users’ data.

As far as the published documents show, its API falls under the same privacy policy as the consumer chat. The Open Platform Terms of Service (effective 29 April 2026) state that “the Open Platform Service and the chat service use the same account”, and we found no clause on training, no retention period and no separate platform policy. If a DPA, professional secrecy or sector rules bind you, treat it as a transfer to a third country and check it first.

Running the weights yourself, with the runtime offline, keeps what you type on your machine — there is no vendor policy to trust. The model card states that “this code repository and the model weights are licensed under the MIT License”. But most “DeepSeek-R1” builds that fit on a laptop are distills fine-tuned from Qwen or Llama, not R1 itself, and studies from early 2025 found heavy refusal on China-related topics.

One publicly documented exposure, in early 2025. On 29 January 2025 Wiz Research reported a publicly accessible DeepSeek database holding “plaintext logs, including Chat History, API Keys, backend details, and operational metadata”; DeepSeek secured it after disclosure. We found no separate breach dated 2026.

No source reviewed for this page describes DeepSeek as malware. What researchers documented was weak engineering: in February 2025 NowSecure reported that the iOS app “transmits sensitive data over the internet without encryption”, with hardcoded keys and outdated Triple DES. That analysis is over a year old and we found no published retest.

On the point most people mean — where the data ends up and under whose law — no. All three keep what you send, and their consumer tiers are governed by a training setting. DeepSeek names the PRC as the storage location and PRC law as the governing law; OpenAI and Anthropic offer business and API tiers that are not trained on by default. That is a different risk, not privacy.

Take the identifiers out before the text reaches it. Occlira runs the detection on your own computer: open the file, review what it flagged, and the values become consistent placeholders such as <PERSON_1>. Press “Copy anonymized text” and paste that into DeepSeek, or upload the anonymized file with your prompt; then save the reply as a file and run it through Deanonymize, which writes a restored copy locally in one click. Because the mapping stays on your device, this is pseudonymization, not anonymization — and the rest of the text still goes to servers in China.

Keep the names out of the chat

Anonymize locally, then use whichever assistant you like on the placeholder copy. Free for 14 days on Windows and macOS (Apple Silicon); one-time licence from €149 per seat.

More: Is ChatGPT private? · is Grok private? · shadow AI at work · local vs cloud redaction · what is PII?