
De-identification, pseudonymisation and anonymisation are three different levels of privacy protection for health data, and none of them is a switch that flips cleanly from identifiable to anonymous. They sit on a spectrum of re-identification risk. De-identification is the broad goal of reducing that risk so individuals cannot reasonably be identified. Pseudonymisation lowers the risk while keeping a key that can reverse it, so the data stays legally personal. Anonymisation aims to reduce the risk so far that no one is reasonably likely to re-identify anyone, at which point the data is no longer personal. The hard part, and the part that matters most, is deciding how much residual risk is acceptable, and in what context.
In conversations about sharing or reusing health data, I often hear these three words used as if they mean the same thing, and I hear re-identification treated as though data is either identifiable or it is not. Neither is accurate, and the honest picture is not black and white. What actually matters is the level of risk that remains, how that risk is measured, and whether it is acceptable given who will hold the data and what else they could combine it with. Getting the language right is the first step. Getting the risk judgment right is the real work.
De-identification is the umbrella term for reducing the information in a dataset so that individuals can no longer reasonably be identified from it. Under HIPAA, it can be achieved in two recognised ways, set out in the guidance from the US Department of Health and Human Services. The first, often called Safe Harbor, requires removing eighteen specified categories of identifiers, from names and geographic detail smaller than a state to dates, contact details and record numbers, and requires that the organisation has no actual knowledge that the remaining information could still identify someone. The second, expert determination, is a documented risk assessment rather than a checklist, and I return to it below. Both are routes to the same goal, which is data that can be used responsibly without exposing the people it came from.
Pseudonymisation is a specific technique, and it is not the same as anonymisation. It replaces identifying details with tokens or keys, so that a record can still be linked and followed over time, for example to build a longitudinal view of a patient's care, without directly revealing who the person is. The defining feature is that it is reversible by design. Because a key exists that could re-link the data to an individual, the EU's GDPR treats pseudonymised data as personal data, with all the protections that implies. Pseudonymisation reduces risk and enables a great deal of legitimate work, but it does not make data anonymous, and it should never be described as if it does.
A common question follows from this: if you destroy the key, is the data then anonymised? Not automatically, and this is where the nuance really bites. While an organisation still holds the key, or any other information that allows re-linking, the data remains personal, and it is pseudonymised rather than anonymised. Destroying the key removes the most direct route back to a person, and it is an important step that moves the data much closer to anonymisation. But it does not, by itself, remove the risk that lives in the data. If rich indirect identifiers remain, someone may still single out or re-identify an individual by combining the details that are left, with no key at all.
There is also a contextual point worth understanding. The same dataset can be effectively anonymous in the hands of a recipient who holds no key and has no reasonable means to re-identify anyone, while remaining pseudonymised for the party that once held the key. Whether destroying the key gets you to anonymisation depends on the data that is left and the environment it sits in, not on the deletion of the key alone.
Anonymisation describes data whose re-identification risk has been reduced so far that no one is reasonably likely to identify an individual, by any means reasonably likely to be used. The important word is reasonably. Regulators do not require the elimination of every conceivable, hypothetical risk. The UK's data protection regulator, the Information Commissioner's Office (ICO), is explicit that anonymisation does not turn on a purely theoretical chance of identification, but on whether re-identification is reasonably likely in the circumstances, and it describes the aim as making that likelihood sufficiently remote. Its guidance is widely referenced internationally, which is why it is useful here. This is why anonymisation is best understood as a high bar on a risk spectrum rather than an absolute state. It is also genuinely difficult to reach, because removing the obvious identifiers is rarely enough on its own.
Re-identification rarely happens by reading a name that was left in by mistake. The EU's Article 29 Working Party set out three ways it tends to happen, and they remain a useful lens. The first is singling out, isolating one person's record from all the others. The second is linkability, connecting records about the same person across different datasets, sometimes called the mosaic effect. The third is inference, deducing something about a person from other values with high probability. This is why the risk is rarely about direct identifiers alone.
It helps to separate two kinds of information. Direct identifiers name a person outright, such as a name, a national identity number or a medical record number. Indirect identifiers, also called quasi-identifiers, do not identify anyone on their own but can single someone out in combination, such as a date of birth, a postal code, a sex, an admission date or a rare diagnosis. Research has shown that a small set of quasi-identifiers, such as date of birth, sex and postal code, can uniquely identify the majority of people in a population. Sound de-identification manages indirect identifiers with as much care as direct ones, using techniques such as generalising a value into a band, grouping rare categories together, or suppressing a detail that is simply too revealing.
This is the question that matters most, and it does not have a single numeric answer. Both major frameworks accept that some residual risk is unavoidable. The US Department of Health and Human Services states plainly that both HIPAA methods, even when properly applied, leave data that retains some risk of identification, and that although the risk is very small, it is not zero. The GDPR asks whether re-identification is reasonably likely, taking account of all the means reasonably likely to be used. So the real test is not whether risk exists, but whether it is low enough, in context, to be acceptable. HIPAA frames that as a very small risk that an anticipated recipient could identify someone using other reasonably available information. The ICO offers a practical way to probe it, the motivated intruder test, which asks whether a reasonably competent person, motivated to try and equipped with public resources and ordinary investigative techniques, could succeed. Acceptable risk, in other words, is a judgment made against a defined threshold, documented, and revisited as circumstances change, not a box that is either ticked or not.
Re-identification risk is never a property of a dataset in isolation. It depends on the environment the data sits in and, crucially, on what other data exists that could be linked to it. The same records can be personal data in one organisation's hands, where a key or linking information is held, and anonymous in another's, where it is not. Regulators expect the assessment to reflect this. A public release, where anyone can obtain the data and combine it with anything, demands a far more robust standard than a release to a named recipient under a data-sharing agreement with security controls, access restrictions and audit. Assessing risk therefore means asking not only what is in the data, but who will hold it, what else they could reasonably obtain, and what controls surround it. This is why the other data sources a recipient could bring to bear belong at the centre of any serious risk assessment, not at its edges.
Expert determination is a recognised HIPAA method, and it is worth being precise about what it is and is not. It requires a person with appropriate knowledge and experience in generally accepted statistical and scientific methods to assess the data and conclude that the risk is very small that an anticipated recipient could identify an individual, alone or in combination with other reasonably available information. HHS is clear that there is no specific certification for such an expert and no fixed numerical threshold that defines very small. The judgment is contextual, and the expert must document the methods and results behind it. Because conditions change as technology and available data change, determinations are often time-limited and reassessed before any re-release. Expert determination is therefore a rigorous and defensible way to reach a risk-based conclusion, but it is a judgment about probability in a context, not a guarantee that re-identification is impossible. Applied well, with the data environment and the anticipated recipient properly accounted for, it is one of the strongest tools available. Applied as a rubber stamp, it is worth very little.
Sound de-identification is a process, not a checkbox. It manages direct and indirect identifiers together. It measures residual risk rather than assuming it, whether through a motivated intruder test, a statistical risk assessment, or expert determination. It takes the data environment into account, matching the strength of the treatment to the release model and the controls around it. It includes independent quality assurance, so that protection is verified rather than asserted. It is documented, so the reasoning can be shown rather than claimed. And it is designed from the outset to comply with the frameworks that apply, which for health data commonly means HIPAA, GDPR and PDPA. The goal throughout is data that is safe to use and safe to share, with the individuals behind it reliably protected, and with the residual risk understood and defensible.
For any organisation thinking about the wider value of its clinical data, this is foundational. Data cannot become a usable, shareable asset until its re-identification risk has been reduced to a level that is genuinely acceptable in context, and that cannot be done while the basic terms, and the risk judgments behind them, are treated as interchangeable. Clarity of language and rigour of process go together, and both are what make health data something you can responsibly build on. I look at that wider value, and what readiness requires, in a companion piece on when health data becomes an asset.
At Jonda Health, de-identification and pseudonymisation sit alongside independent quality assurance in a process designed to comply with HIPAA, GDPR and PDPA, with re-identification risk assessed in context rather than assumed away. Jonda Health is ISO 27001 certified.
Is pseudonymised data still personal data?
Yes. Because pseudonymisation is reversible through a key that can re-link the data to an individual, GDPR continues to treat pseudonymised data as personal data, with the protections that implies. It reduces risk, but it does not remove personal-data status.
What is the difference between HIPAA Safe Harbor and expert determination?
Safe Harbor removes eighteen specified categories of identifiers and requires that the organisation has no actual knowledge that the remaining data could identify someone. Expert determination is a documented assessment in which a qualified expert concludes that the risk is very small that an anticipated recipient could re-identify an individual using other reasonably available information. Safe Harbor is a fixed rule; expert determination is a contextual risk judgment.
How much re-identification risk is acceptable?
There is no single number. No method reduces risk to zero, so both frameworks work to a threshold. HIPAA accepts a risk that is very small that an anticipated recipient could identify someone using reasonably available information. GDPR asks whether re-identification is reasonably likely given all the means reasonably likely to be used. The acceptable level is judged in context, documented, and revisited as circumstances change.
If you destroy the pseudonymisation key, is the data anonymised?
Not automatically. Destroying the key removes the most direct way to re-link records, and it is an important step, but data is only anonymous if no one is reasonably likely to re-identify a person by any means reasonably likely to be used. If rich indirect identifiers remain, re-identification can still be possible through singling out, linkability or inference. Whether you reach anonymisation depends on the remaining data and its context, not on deleting the key alone.
What are quasi-identifiers?
Quasi-identifiers are details that do not name a person directly but can identify them in combination, such as a date of birth, a postal code and a sex, or a rare diagnosis together with a date and a location. Because these combinations can single someone out, strong de-identification manages quasi-identifiers as carefully as direct identifiers, by generalising, grouping or suppressing the values that carry the most risk.
What is the motivated intruder test?
It is a practical way to assess re-identification risk, described by the UK's Information Commissioner's Office. It asks whether a reasonably competent person, motivated to try and using public resources and ordinary investigative techniques, could identify individuals in a dataset. It is a way of testing whether re-identification is reasonably likely rather than merely conceivable.
This article is general information about privacy and data-protection concepts. It is not legal advice, and specific obligations depend on your jurisdiction and circumstances.
Sources and further reading: HHS OCR, Guidance on De-identification of Protected Health Information · ICO, How do we ensure anonymisation is effective? · Article 29 Working Party, Opinion 05/2014 on Anonymisation Techniques
Related reading: Reading a Lab Report with AI Is Not the Same as Making It Usable · The Hidden Engineering Tax of Using Document AI in Healthcare
Suhina Singh is the founder and CEO of Jonda Health, a Singapore-based health data infrastructure company. A physician by training, she works with health systems across Asia-Pacific to harmonise, de-identify and standardise clinical data so it can be trusted and used. Jonda Health is ISO 27001 certified.
Real-time health data transformation engine. Any format in, clean structured output out.
Explore the engine →From 3-person startups to Fortune 500 insurers.
Talk to sales →Field notes, product updates, and customer stories from the health data frontier.
Browse all resources →

