Skip to main navigation Skip to search Skip to main content

K-anonymity decay in multi-turn clinical large language model conversations

  • James Weatherhead
  • , Azra Hasan
  • , Jake Weatherhead
  • , George Golovko
  • , Bradley Grant
  • , Juan David Garcia
  • , Hugo Sebastian Certuche
  • , Reuben Peter Powell
  • , Jose Marri Abril
  • , Peter McCaffrey

Research output: Contribution to journalArticlepeer-review

Abstract

Per-prompt de-identification, a commonly adopted practice for protecting patient privacy in clinical artificial intelligence (AI) conversations, may be insufficient to bound re-identification risk when the same patient is discussed across multiple conversational turns. Under the Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method, eighteen specified identifier categories must be removed from each disclosure independently, but this approach does not assess cumulative re-identification risk: successive turns reveal additional quasi-identifiers that progressively narrow the set of matching records, eroding the privacy protection that any single turn appeared to provide. We quantified this vulnerability by simulating progressive quasi-identifier disclosure, modeled on clinical case presentations, against a synthetic electronic health record cohort. Under progressive disclosure, 79.9% of simulated patients fell below the small-cell threshold ((Formula presented)) by the end of their disclosure sequence, with a median of seven disclosure steps to reach this threshold; when rare attributes were disclosed first, the median decreased to four steps. Even when no individual step disclosed a direct identifier, the cumulative quasi-identifier profile degraded k-anonymity below accepted safety thresholds. These findings suggest an operational limitation in per-prompt de-identification as applied to multi-turn clinical AI and large language model conversations: although HIPAA Safe Harbor includes a provision requiring no actual knowledge that remaining information could identify an individual, clinicians lack tools to assess cumulative re-identification risk in real time.

Original languageEnglish (US)
Article number1832168
JournalFrontiers in Digital Health
Volume8
DOIs
StatePublished - 2026

Keywords

  • clinical decision support
  • de-identification
  • HIPAA
  • k-anonymity
  • large language models
  • re-identification
  • shadow AI
  • synthetic data

ASJC Scopus subject areas

  • Medicine (miscellaneous)
  • Biomedical Engineering
  • Health Informatics
  • Computer Science Applications

Fingerprint

Dive into the research topics of 'K-anonymity decay in multi-turn clinical large language model conversations'. Together they form a unique fingerprint.

Cite this