'De-identified' data and what it can still be used for

When an AI scribe vendor tells you “we only use de-identified data to improve our models,” it sounds like a privacy promise. It is closer to a legal disclosure. Understanding de-identified data training starts with an uncomfortable fact: under HIPAA, once health information is properly de-identified, it stops being protected health information at all. It can be kept indefinitely, sold, combined with other datasets, and used to train AI models — lawfully, and often without telling you. The reassurance you heard is really a description of how broadly the data can travel.

This matters because your session recordings and the notes drawn from them are among the most sensitive material a person ever creates. So it is worth knowing exactly what de-identification does, what it does not do, and why running everything on your own machine sidesteps the whole question.

What happens to data after de-identification A flow from protected session data, through a de-identification step, to three outcomes. The session audio and transcript are protected PHI on the left. After de-identification, the data is no longer PHI and can be retained indefinitely, reused or shared, or used to train models. Session audio + transcript PHI — protected De-identification strip identifiers No longer PHI — outside HIPAA's reach Retained indefinitely Reused or shared Used to train models
De-identification does not lock data down — it removes the HIPAA protection, after which the data can be retained, reused, or used to train models.

What HIPAA de-identification actually means

HIPAA recognizes two ways to de-identify protected health information. The first, Safe Harbor, requires removing the categories of identifiers the rule specifies — names, geographic detail smaller than a state, dates more specific than a year, contact information, record numbers, and so on. The second, Expert Determination, lets a qualified statistician certify that the risk of re-identification is “very small.” Both are legitimate. Neither is anonymization in the everyday sense, and that distinction does most of the work.

The key consequence is structural, not technical. HHS guidance is explicit that information meeting either standard is no longer protected health information. The Privacy Rule simply stops applying. There is no required retention limit, no required deletion, no continuing duty to safeguard it under HIPAA. A vendor that de-identifies your session data has, in legal terms, converted regulated clinical material into an unregulated asset.

Why “we only use de-identified data” is not reassuring

Read on its own, that sentence tells you almost nothing about your clients’ protection. It tells you about the vendor’s freedom. Once the de-identification step is complete, several things become permissible that you might assume are off the table:

  • Indefinite retention. The data can be kept long after you stop using the product, or after the company is acquired and the dataset becomes part of the deal.
  • Reuse and combination. De-identified clinical text is valuable. It can be merged with other corpora, licensed, or analyzed for purposes you never had in mind.
  • Model training. This is the headline use. Your phrasing, your clients’ described experiences, the texture of real sessions — all of it can become training signal for systems sold back to the market.

There are real-world limits to how clean “de-identified” actually is. Therapy transcripts are dense with specifics: an unusual occupation, a town’s only hospital, the sequence of a family’s losses. Free text resists a tidy list of identifiers, and re-identification research has repeatedly suggested that records presented as anonymous can sometimes be matched back to individuals when enough quasi-identifiers remain. A statistician’s “very small” risk is a probability, not a guarantee — and it is assessed against the data as it exists today, not against tomorrow’s matching techniques.

Once data is de-identified, the question is no longer “is this protected?” It is “do you trust where this will go, with whom, and for how long?” — because HIPAA has stopped answering on your behalf.

This is the same gap I wrote about in why HIPAA compliance is not enough: a vendor can be fully compliant and still build a business on the data flowing through it. Compliance is the floor, not the ceiling.

Where to look in a privacy policy

If a tool processes sessions in the cloud, the de-identification clause is worth finding and reading slowly. A few questions cut through the marketing:

Look forReassuringWorth questioning
De-identification methodNamed (Safe Harbor / Expert Determination)Vague “anonymized” with no standard cited
Training useOpt-in, or never”We may use de-identified data to improve our services”
RetentionStated limit, deletion on cancellationSilent, or “as long as necessary”
After de-identificationTreated as still off-limitsExplicitly outside the agreement’s protections

The phrase “to improve our services” is the one to slow down on; it is broad enough to cover training, and it usually appears exactly where the data stops being yours. For a fuller walkthrough, see how to read an AI scribe’s privacy policy. Rules around de-identification and secondary use vary by state and by payer contract, and none of this is legal advice, so confirm specifics with your board or attorney rather than relying on a vendor’s summary.

How on-device processing settles the de-identified data training question

The cleanest answer to “what can they do with de-identified data” is to ensure there is no “they.” If transcription and drafting happen entirely on your own Mac, no session audio or transcript is uploaded, so there is nothing for a vendor to de-identify, retain, or train on. The de-identified data training debate becomes moot — not because the data is better protected after it leaves, but because it never leaves.

This is the model CouchNotes is built on: recording or dictation, on-device transcription, and a SOAP, DAP, or BIRP draft you review, edit, and sign, with audio that auto-deletes per your setting. No cloud, no account, no telemetry. The draft is a starting point; you remain the author of record, and the underlying material stays where the work happened.

De-identification is a real and useful legal tool, but it is a tool for the people holding your data, not a shield for the people in your sessions. When a vendor says it only trains on de-identified data, hear it for what it is: an honest statement that, somewhere in their systems, the version of your clients’ words has been stripped of HIPAA’s protection and put to work. The surest way to control that outcome is to keep the most revealing material from being uploaded in the first place — and the only way to be certain is for it never to leave the room where the session happened.

Dario Valles

Building CouchNotes — on-device AI session notes for therapists on macOS and Windows. Sessions never leave your computer; that's the whole point.

Get the free beta