How to remove personal data from audio recordings
A recording is often the single most sensitive file you handle — a real voice plus names, numbers and health or legal detail. This guide shows how to redact the personal data out of audio on your own computer: transcribe locally, bleep the spoken identifiers, and share safely — without uploading the recording anywhere.
Why audio is the hardest kind of PII to redact
Personal data hides in ordinary files far more often than people expect: after scanning about 6.5 million Google Drive files, security firm Metomic found 40.2% contained sensitive information, including PII, and a third were shared externally. (Source: Metomic.) Audio is the worst case. Spoken identifiers scatter unpredictably across a long recording, they’re tedious to find by ear, and the recording carries the person’s actual voice — which, once processed to identify the speaker, can itself become biometric data under the GDPR (more on that below).
What audio redaction actually means
Audio redaction is removing identifying information from a recording by muting or bleeping the spoken segments — names, numbers, addresses — and masking the same values in the transcript. The key point that trips people up: masking a name in the transcript does nothing to the recording. If you share the audio, the name is still there to hear. Real anonymization acts on both the sound and the text, so no identifier survives in either.
On-device vs cloud transcription: the privacy fork
Most transcription happens in the cloud, and the scale is enormous — Otter.ai alone reports processing over 1 billion meetings for more than 35 million users. (Source: Otter.ai.) That convenience has a cost. In August 2025 Otter was hit with a federal class action alleging it “deceptively and surreptitiously” records conversations across Zoom, Meet and Teams and uses them to train its AI. (Source: NPR.)
It isn’t just one vendor. The Freedom of the Press Foundation warns that Otter, Rev, Trint and Descript all have the technical ability to access uploaded audio, and that several train on it or route it to third-party processors and human reviewers — advising journalists to prefer on-device tools like Whisper instead. (Source: Freedom of the Press Foundation.) Those human reviewers are a real exposure: a OneZero investigation found Rev’s freelance transcribers were exposed to deeply sensitive material, including audio describing child sexual abuse. (Source: OneZero.) On-device transcription — Whisper-quality models running on your own machine — sidesteps all of it: the raw recording is never uploaded, so there’s no retention policy, subprocessor or reviewer to trust.
Is a voice recording personal or biometric data under the GDPR?
A voice recording is always personal data if someone can be identified from it. It crosses into special-category biometric data once it’s technically processed to identify or verify the speaker — for example turning speech into a voiceprint. The ICO treats that as high-risk processing that in most cases requires a Data Protection Impact Assessment and, absent a tailored condition, explicit consent under Article 9(2)(a) alongside an Article 6 lawful basis. (Source: ICO.) Get it wrong and the exposure is the top GDPR fine tier — up to €20 million or 4% of global annual turnover.
Consent, and the law of recording voices
There are two strands to stay on the right side of. First, data protection: the EDPB considers it “very likely” that voice/audio services need a DPIA, and holds that an accidentally-triggered recording is “highly unlikely” to count as valid consent. (Source: EDPB Guidelines 02/2021.)
Second, wiretap law. US federal law allows recording with one party’s consent, but as of 2026 fifteen states require all-party consent: California, Connecticut, Delaware, Florida, Illinois, Maryland, Massachusetts, Michigan, Montana, Nevada, New Hampshire, Oregon, Pennsylvania, Vermont and Washington. (Source: Sembly AI.) The stakes are concrete: under California’s Invasion of Privacy Act, someone recorded without all-party consent can sue for statutory damages of $5,000 per violation or three times actual damages, whichever is greater. (Source: Shouse Law Group.) The practical takeaway is the same either way: get consent, keep only what you need, and redact identifiers before the recording goes anywhere.
When you must redact a recording before sharing it (DSARs)
A common trigger is a subject access request. When a DSAR covers a call recording that also contains other people’s data, the ICO expects you to remove or bleep the third-party identifiers before disclosure, respond within one month, and apply the redaction so the removed data can’t be recovered from the copy you release. (Source: ICO.) This is where on-device bleeping fits the rule neatly: because the bleeped audio can’t be un-bleeped, the copy you disclose stays redacted by design.
Who redacts audio, and why
- Journalists protecting a source’s name, employer and phone number before an interview file goes to an editor.
- Qualitative researchers anonymizing consented interview recordings for an ethics board (IRB) before sharing transcripts.
- HR teams removing a witness’s name and home address spoken aloud in a recorded investigation meeting.
- Legal teams bleeping a minor’s name or a Social Security number read into a deposition.
- Healthcare and therapists de-identifying a session recording before using it for supervision or notes.
How to remove PII from an audio recording, step by step
- Transcribe on your own computer. Open the recording in Occlira. It is transcribed locally with an on-device speech model — the raw audio, the most sensitive version, never leaves your machine or touches a cloud service.
- Separate the speakers. Speakers are told apart on-device, so you can see who said what and redact per speaker while keeping the conversation readable.
- Detect the spoken PII. Occlira flags personal data spoken aloud — names, phone numbers, addresses, and account or card numbers — in the synced transcript, each with its exact position in the audio.
- Review and confirm. Keep or redact each match in a quick review step, and select anything the detector missed to add it. Nothing is changed without your say-so.
- Bleep the audio and mask the transcript. For every confirmed item, Occlira mutes/bleeps that segment of the waveform and masks the same value in the transcript — so the identifier is gone from what people hear, not just what they read.
- Export a shareable copy. Save a redacted copy to share, archive or disclose. The bleeped audio is permanent — exactly what you want for a copy you hand out — while the transcript stays reversible within your working session, so you can double-check the originals before you finish.
Doing it by hand vs automatically
You can redact audio manually in an editor like Audacity — find each identifier by ear, select the region, and silence it. For a short clip that works; for a 90-minute interview or a batch of call recordings it’s slow and easy to miss a name on the second pass. Transcript-driven detection scales: it locates every occurrence by its timestamp, so you review a list instead of scrubbing a waveform, and nothing identifying slips through because it was said quietly at minute 72. That’s exactly what Occlira does: it transcribes on-device, lists every detected name and number by timestamp, and bleeps each one on your confirmation — and unlike Otter, Rev or Descript, the recording is never uploaded.
Audio-redaction checklist
Before you share, archive or disclose a recording, run through this:
- Confirm you have consent or another lawful basis to record and process the audio
- Run a DPIA if the recording involves systematic monitoring or turning voices into voiceprints
- Transcribe on your own device — don’t upload the raw recording to a cloud service
- Detect and review every spoken name, phone number, address and account/card number
- Bleep the audio AND mask the transcript — a masked transcript alone leaves the name audible
- Play back the redacted copy and confirm no identifier survives
- Strip file metadata before sharing
- For a disclosure/DSAR, export a fresh irreversible copy and keep the reversible mapping private
How Occlira does it — entirely on your device
Occlira is built for exactly this. It transcribes your recording locally with a Whisper-quality on-device model, separates the speakers, auto-detects the names and numbers spoken aloud, and — once you confirm — bleeps the audio while masking the transcript at each match. It handles MP3, WAV, M4A, FLAC and OGG, and works offline after activation on Windows and macOS. The audio you export is permanently bleeped, while the transcript stays reversible within your working session. No upload, no account, no cloud. See exactly what stays on your device.
Frequently asked questions
Audio redaction is the removal of identifying information from a recording by muting or bleeping the spoken segments — names, numbers, addresses — and masking the same values in any transcript. Unlike redacting a text document, it acts on the waveform itself, so the identifier can no longer be heard on playback.
A masked transcript only hides the text; if you share the underlying recording, the name is still audible. True anonymization requires acting on both — bleeping the audio at each match and masking the transcript — so nothing identifying survives in either the words or the sound.
Transcribe the audio, locate each spoken name in the transcript (with its timestamp in the recording), and mute or overlay a tone on that segment of the waveform. Occlira does this automatically: it detects the names on-device and bleeps every occurrence, which is far more reliable than hunting for them by hand in an editor.
A voice recording is always personal data if a person can be identified from it. It becomes special-category biometric data once it is technically processed to identify or verify the speaker — for example turning speech into a voiceprint — which the ICO treats as high-risk processing that usually needs a DPIA and explicit consent under Article 9.
Usually, yes. Under the GDPR you need a lawful basis, and the EDPB considers a DPIA “very likely” to be required for voice/audio services. In the US, federal law allows one-party consent, but many states require all-party consent (see below). The safe pattern is: get consent, minimize what you keep, and redact identifiers before sharing.
As of 2026, fifteen states require all-party consent: California, Connecticut, Delaware, Florida, Illinois, Maryland, Massachusetts, Michigan, Montana, Nevada, New Hampshire, Oregon, Pennsylvania, Vermont and Washington. On interstate calls the stricter law usually applies, so if any participant is in an all-party state such as California, you need everyone’s consent regardless of where you are.
Several cloud transcription services can, by default, use uploaded recordings to improve their models or route them to third-party processors and human reviewers. Because the raw audio is uploaded, you are trusting their retention and training policies. On-device transcription avoids the question entirely — the recording is never uploaded.
If a data subject access request covers a recording that also contains other people’s personal data, the ICO expects you to remove or bleep the third-party identifiers before disclosure (applying the balancing test), respond within one month, and apply the redaction so the removed data can’t be recovered from the copy you release.
Bleeping the audio is permanent: once you export the redacted recording, the muted segments can’t be restored — which is exactly what you want for a copy you share or disclose. The transcript stays reversible within your working session (via Occlira’s Deanonymize) so you can check the original values before you finish; the exported audio itself can’t be un-bleeped.
Redact audio without uploading it anywhere
Transcribe locally, bleep the personal data, share safely. Free for 14 days on Windows and macOS.
More: for law firms · what is PII? · use AI without breaking the GDPR · how your data is handled