'De-identified' data and what it can still be used for
When an AI scribe vendor tells you “we only use de-identified data to improve our models,” it sounds like a privacy promise. It is closer to a legal disclosure. Understanding de-identified data training starts with an uncomfortable fact: under HIPAA, once health information is properly de-identified, it stops being protected health information at all. It can be kept indefinitely, sold, combined with other datasets, and used to train AI models — lawfully, and often without telling you. The reassurance you heard is really a description of how broadly the data can travel.
This matters because your session recordings and the notes drawn from them are among the most sensitive material a person ever creates. So it is worth knowing exactly what de-identification does, what it does not do, and why running everything on your own machine sidesteps the whole question.
What HIPAA de-identification actually means
HIPAA recognizes two ways to de-identify protected health information. The first, Safe Harbor, requires removing the categories of identifiers the rule specifies — names, geographic detail smaller than a state, dates more specific than a year, contact information, record numbers, and so on. The second, Expert Determination, lets a qualified statistician certify that the risk of re-identification is “very small.” Both are legitimate. Neither is anonymization in the everyday sense, and that distinction does most of the work.
The key consequence is structural, not technical. HHS guidance is explicit that information meeting either standard is no longer protected health information. The Privacy Rule simply stops applying. There is no required retention limit, no required deletion, no continuing duty to safeguard it under HIPAA. A vendor that de-identifies your session data has, in legal terms, converted regulated clinical material into an unregulated asset.
Why “we only use de-identified data” is not reassuring
Read on its own, that sentence tells you almost nothing about your clients’ protection. It tells you about the vendor’s freedom. Once the de-identification step is complete, several things become permissible that you might assume are off the table:
- Indefinite retention. The data can be kept long after you stop using the product, or after the company is acquired and the dataset becomes part of the deal.
- Reuse and combination. De-identified clinical text is valuable. It can be merged with other corpora, licensed, or analyzed for purposes you never had in mind.
- Model training. This is the headline use. Your phrasing, your clients’ described experiences, the texture of real sessions — all of it can become training signal for systems sold back to the market.
There are real-world limits to how clean “de-identified” actually is. Therapy transcripts are dense with specifics: an unusual occupation, a town’s only hospital, the sequence of a family’s losses. Free text resists a tidy list of identifiers, and re-identification research has repeatedly suggested that records presented as anonymous can sometimes be matched back to individuals when enough quasi-identifiers remain. A statistician’s “very small” risk is a probability, not a guarantee — and it is assessed against the data as it exists today, not against tomorrow’s matching techniques.
Once data is de-identified, the question is no longer “is this protected?” It is “do you trust where this will go, with whom, and for how long?” — because HIPAA has stopped answering on your behalf.
This is the same gap I wrote about in why HIPAA compliance is not enough: a vendor can be fully compliant and still build a business on the data flowing through it. Compliance is the floor, not the ceiling.
Where to look in a privacy policy
If a tool processes sessions in the cloud, the de-identification clause is worth finding and reading slowly. A few questions cut through the marketing:
| Look for | Reassuring | Worth questioning |
|---|---|---|
| De-identification method | Named (Safe Harbor / Expert Determination) | Vague “anonymized” with no standard cited |
| Training use | Opt-in, or never | ”We may use de-identified data to improve our services” |
| Retention | Stated limit, deletion on cancellation | Silent, or “as long as necessary” |
| After de-identification | Treated as still off-limits | Explicitly outside the agreement’s protections |
The phrase “to improve our services” is the one to slow down on; it is broad enough to cover training, and it usually appears exactly where the data stops being yours. For a fuller walkthrough, see how to read an AI scribe’s privacy policy. Rules around de-identification and secondary use vary by state and by payer contract, and none of this is legal advice, so confirm specifics with your board or attorney rather than relying on a vendor’s summary.
How on-device processing settles the de-identified data training question
The cleanest answer to “what can they do with de-identified data” is to ensure there is no “they.” If transcription and drafting happen entirely on your own Mac, no session audio or transcript is uploaded, so there is nothing for a vendor to de-identify, retain, or train on. The de-identified data training debate becomes moot — not because the data is better protected after it leaves, but because it never leaves.
This is the model CouchNotes is built on: recording or dictation, on-device transcription, and a SOAP, DAP, or BIRP draft you review, edit, and sign, with audio that auto-deletes per your setting. No cloud, no account, no telemetry. The draft is a starting point; you remain the author of record, and the underlying material stays where the work happened.
De-identification is a real and useful legal tool, but it is a tool for the people holding your data, not a shield for the people in your sessions. When a vendor says it only trains on de-identified data, hear it for what it is: an honest statement that, somewhere in their systems, the version of your clients’ words has been stripped of HIPAA’s protection and put to work. The surest way to control that outcome is to keep the most revealing material from being uploaded in the first place — and the only way to be certain is for it never to leave the room where the session happened.