What is PII? Types, examples, and how it differs from personal data
PII (personally identifiable information) is any information that can identify a specific person — on its own, like a passport number, or combined with other data, like a ZIP code plus a date of birth. The US standard, NIST SP 800-122, frames it in two parts: information that distinguishes or traces an identity, plus any other information linked or linkable to that person.
What counts as PII? The official definition
There is no single closed list — what counts depends on the jurisdiction. The most widely-cited US definition comes from NIST Special Publication 800-122, which splits PII into (1) information that can be used to distinguish or trace an individual’s identity — a name, Social Security number, date and place of birth, mother’s maiden name or biometric records — and (2) any other information that is linked or linkable to an individual, such as medical, educational, financial or employment data. (Source: NIST SP 800-122.) That “linked or linkable” test is the key idea: data doesn’t have to name someone to be PII — it just has to help identify them.
Types of PII, with examples
NIST groups PII into these example categories (it’s illustrative, not exhaustive):
| Category | Examples |
|---|---|
| Names | Full name, maiden name, mother’s maiden name, alias |
| Personal ID numbers | SSN, passport, driver’s licence, taxpayer ID, patient ID, financial or credit-card number |
| Address information | Street address or email address |
| Asset / online identifiers | IP address, MAC address, cookie and device identifiers |
| Telephone numbers | Mobile, landline, fax |
| Personal characteristics | Photo of the face, fingerprints, retina scan, voice signature, other biometrics |
| Owned-property identifiers | Vehicle registration or title number, device serial number |
Categories from NIST SP 800-122 §2.2. Occlira detects 40+ types like these across documents, spreadsheets, email, audio and images.
Direct identifiers, indirect identifiers and quasi-identifiers
Privacy teams classify PII by how it identifies someone — the split behind most classification schemes:
- Direct identifiers single someone out on their own — a passport or SSN, an email address, a phone number.
- Indirect (quasi-) identifiers don’t identify anyone alone but do in combination. NIST’s own example: a list of credit scores identifies no one, but add age, address and gender and the individuals become identifiable. (Source: NISTIR 8053.)
The classic demonstration is Latanya Sweeney’s, who re-identified a US governor’s “de-identified” medical record using only date of birth, ZIP code and sex — and estimated that the same three fields uniquely identify roughly 87% of the US population (a later replication put it near 63%). (Source: Golle 2006, citing Sweeney.) That’s why removing the obvious names isn’t enough: leftover quasi-identifiers can re-identify people, so good redaction has to catch both.
PII vs personal data vs PHI
| Term | Framework | Scope | Example |
|---|---|---|---|
| PII | US practice (e.g. NIST SP 800-122) | Information that distinguishes or traces an individual, plus anything linked or linkable to them | SSN, name, email, IP address |
| Personal data | EU/UK GDPR, Art. 4(1) | Any information relating to an identified or identifiable person, directly or indirectly — explicitly includes online identifiers | Name, location data, cookie ID, IP address |
| PHI | US HIPAA | Health information tied to one of 18 identifiers, held by a covered entity | A diagnosis next to a name or medical-record number |
The three overlap heavily, and the practical task is identical — find the identifiers and remove them. The main difference is breadth: GDPR Art. 4(1) covers anyone identifiable “directly or indirectly,” and Recital 30 spells out that online identifiers — IP addresses, cookie IDs, RFID tags — are personal data. That makes the GDPR notably broader than the classic US notion of PII.
Sensitive vs non-sensitive PII (and GDPR special categories)
Not all PII carries the same risk, so it’s useful to split it in two:
- Sensitive PII can cause real harm if exposed and warrants the strongest protection — national IDs, financial and payment data, health records, biometrics and precise location.
- Non-sensitive PII is often public or lower-risk on its own — a name, job title or business email — and becomes risky mainly when combined with other fields.
The GDPR has a stricter, named tier called special-category data whose processing is prohibited by default (Art. 9): data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, trade-union membership, genetic data, biometric data used to uniquely identify someone, health data, and data about a person’s sex life or sexual orientation. “Special category” is the GDPR term; “sensitive PII” is the looser US analogue.
The 18 HIPAA identifiers (Safe Harbor)
US health data has the most concrete inventory of what counts as an identifier. To de-identify a record under HIPAA’s Safe Harbor method, a covered entity must remove all 18 identifiers and have no actual knowledge the remainder could still identify the person: (Source: 45 CFR 164.514(b)(2).)
- Names
- Geographic data smaller than a state (only the first 3 ZIP digits, and only if the area has over 20,000 people)
- All dates except year (birth, admission, discharge, death); ages over 89 grouped as “90 or older”
- Telephone numbers
- Fax numbers
- Email addresses
- Social Security numbers
- Medical record numbers
- Health-plan beneficiary numbers
- Account numbers
- Certificate/licence numbers
- Vehicle identifiers and licence plates
- Device identifiers and serial numbers
- Web URLs
- IP addresses
- Biometric identifiers, including finger and voice prints
- Full-face photographs and comparable images
- Any other unique identifying number, characteristic or code
HIPAA’s second route, Expert Determination, instead asks a qualified expert to confirm the risk of re-identification is “very small” — a residual-risk standard, not a claim of zero risk. (Source: HHS.)
Anonymization vs pseudonymization vs redaction
These three get used interchangeably, but the law treats them very differently:
| Technique | Reversible? | Still personal data? | Goal |
|---|---|---|---|
| Redaction | The data is removed or blacked out | Removed from that document | Take identifiers out of a file before sharing it |
| Pseudonymization | Reversible — with the separately-kept mapping | Yes — still personal data (GDPR Art. 4(5)) | Reduce risk while keeping the data usable and restorable |
| Anonymization | Irreversible — no one can be re-identified | No — outside the GDPR if truly anonymous (Recital 26) | Make re-identification impossible |
The distinction is legally load-bearing. Under GDPR Art. 4(5), pseudonymised data — reversible with information kept separately — is still personal data; only truly anonymous data (Recital 26) falls outside the GDPR. The EDPB’s 2025 pseudonymisation guidelines confirm it: reversible tokenization reduces risk but keeps the data in scope.
Why removing names isn’t enough
Even a careful de-identification leaves some residual risk. A test of HIPAA Safe Harbor found that fewer than 1% of records could be re-identified when only year of birth, sex and 3-digit ZIP remained — small, but not zero. (Source: NISTIR 8053.) Combined with Sweeney’s 87% figure, the lesson is consistent: stripping names is the easy part; the real work is catching the indirect identifiers that quietly single people out — which is exactly what good redaction tooling has to do, not just match names.
How to remove PII from your files — locally
Once you know what to look for, the safe way to remove it is to keep the whole job on your own machine. Occlira is a Windows and macOS desktop app that detects and removes 40+ types of PII — names, national IDs, emails, financial identifiers, health and more — across documents, spreadsheets, email, audio and images, 100% on-device, with no cloud and no account. Because it also flags indirect identifiers, not just names, it addresses the re-identification gap above.
Its anonymization is reversible: it replaces each identifier with a neutral placeholder and keeps the mapping on your device, so you can restore the originals. In GDPR terms that is pseudonymization — a strong risk-reduction step, but, as above, the output is still personal data, so treat it as data minimization rather than a route out of the rules. See anonymize before ChatGPT or how to redact Word, PDF and Excel for the step-by-step.
Frequently asked questions
PII (personally identifiable information) is any information that can identify a specific person, on its own or combined with other data. NIST defines it in two parts: information that distinguishes or traces an identity (a name, SSN, date and place of birth, biometrics), plus any other information linked or linkable to that person (medical, financial, employment or education data).
“PII” is the common US term; “personal data” is the GDPR term and is broader — it covers anything relating to an identifiable person, directly or indirectly, and explicitly includes online identifiers like IP addresses and cookie IDs. Most things that are PII are also personal data; the GDPR simply casts a wider net.
PHI (protected health information) is the HIPAA subset: health information tied to one of 18 specific identifiers and held by a covered entity such as a hospital or health plan. All PHI is PII, but plenty of PII (a work email, say) is not PHI.
Yes. An email address points to a specific person — especially a name-based work address — so it counts as PII in the US and as personal data under the GDPR.
Under the GDPR, yes: Recital 30 names IP addresses among the online identifiers that make a person identifiable, so they are personal data. US definitions historically treated it as PII only in combination, but increasingly count it too.
A direct identifier singles someone out on its own (an SSN, a passport number, an email). Indirect or quasi-identifiers don’t identify anyone alone but do in combination — date of birth, ZIP code, gender, job title. NIST notes that a list of credit scores identifies no one, but add age, address and gender and the individuals become identifiable.
Pseudonymization replaces identifiers with reversible placeholders while the mapping is kept separately — under GDPR Art. 4(5) the result is still personal data. Anonymization aims to make re-identification impossible; only then (Recital 26) does the data fall outside the GDPR. The EDPB’s 2025 pseudonymisation guidelines confirm pseudonymised data stays in scope.
Yes. Because it can be re-attributed to a person using the additional information kept separately, pseudonymized data remains personal data under the GDPR. It lowers risk and enables reversibility, but it is not the same as anonymization.
Often, yes. A landmark study by Latanya Sweeney estimated that about 87% of the US population is uniquely identifiable from ZIP code, full date of birth and sex; a later replication put it near 63%. That’s why removing names alone doesn’t anonymize a dataset — the quasi-identifiers can still single people out.
Find and remove PII locally
Detect and remove personal data across your files, on your own computer. Free for 14 days on Windows and macOS.
More: use AI without breaking the GDPR · remove personal data from audio · how your data is handled