The GDPR Definition of Pseudonymisation
The General Data Protection Regulation draws a sharp line between two concepts that are routinely conflated in practice: pseudonymisation and anonymisation. Getting this distinction wrong does not merely invite regulatory scrutiny — it can invalidate the legal basis for an entire data processing programme.
Article 4(5) of the GDPR defines pseudonymisation as:
"The processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person."
— GDPR, Article 4(5)
The critical phrase is "without the use of additional information." Pseudonymisation does not eliminate the link between data and individual — it separates the data from the key that enables re-identification. The key still exists. The link can be restored. And because the link can be restored, pseudonymised data remains personal data under the GDPR.
Anonymisation, by contrast, is the irreversible transformation of data such that the data subject is no longer identifiable by any means reasonably likely to be used. Truly anonymous data falls outside the scope of the GDPR entirely. The regulation does not apply to it. No lawful basis is required. No data subject rights attach. No data protection impact assessment is needed.
The gap between these two concepts is where most organisations encounter difficulty.
Why Pseudonymised Data Can Remain Personal Data
The test for whether data is personal under the GDPR is not whether a given organisation can re-identify an individual, but whether any party using means reasonably likely to be employed could do so. This is the "motivated intruder" test articulated by the Article 29 Working Party (now the European Data Protection Board) and reinforced by the Court of Justice of the European Union.
For genomic data, this test is particularly consequential. Genomic sequences are inherently unique to individuals. Even when direct identifiers (name, date of birth, medical record number) are removed and replaced with pseudonyms, the genomic data itself retains re-identification potential through several vectors:
- Familial matching — A genomic sequence can be linked to family members whose data appears in publicly available genealogy databases. Research published since 2013 has demonstrated that a significant proportion of individuals of European descent can be identified through long-range familial matching using consumer genomics databases.
- Phenotypic inference — Genomic data encodes physical characteristics (eye colour, hair colour, facial structure) that, combined with demographic information, can narrow identification to a small population.
- Linkage attacks — When pseudonymised genomic data is combined with other datasets — health records, geographic data, temporal metadata — the combination may be sufficient for re-identification, even when each dataset individually is not.
- Longitudinal uniqueness — Unlike a credit card number or email address, a genomic sequence cannot be changed. Any exposure creates permanent re-identification risk that cannot be mitigated by issuing new identifiers.
These characteristics mean that genomic data is, for all practical purposes, extremely difficult to anonymise. Most processing of genomic data in biotech research operates under pseudonymisation rather than anonymisation, and must be governed accordingly.
Practical Controls for Pseudonymisation
Effective pseudonymisation is not merely the act of replacing a name with a code. It is a set of technical and organisational measures that, taken together, reduce the risk of re-identification to an acceptable level while preserving the data's utility for its intended purpose.
Key Separation
The pseudonymisation key — the mapping between the pseudonym and the data subject's identity — must be stored separately from the pseudonymised data. "Separately" means not merely in a different database table, but in a distinct system with independent access controls, ideally managed by a different organisational unit. If the same administrator has access to both the pseudonymised dataset and the re-identification key, the pseudonymisation offers limited practical protection.
Access Controls
Access to pseudonymised data and access to the re-identification key must be governed by separate, independently managed access control policies. The principle of least privilege applies: researchers who need the pseudonymised data for analysis should not have access to the re-identification key. Personnel who manage the key (for example, for regulatory submissions requiring identified data) should not have routine access to the pseudonymised analytical datasets.
Role Separation
Organisational role separation reinforces technical access controls. The data protection officer, the custodian of the re-identification key, and the research team working with pseudonymised data should be distinct roles with distinct reporting lines. This separation reduces the risk of a single point of compromise enabling re-identification.
Encryption
Encryption at rest and in transit is a baseline expectation for pseudonymised data, but it is not a substitute for pseudonymisation. Encryption protects against unauthorised access to the data as a whole; pseudonymisation protects against attribution of the data to an individual even when the data is accessed by authorised users. Both controls are necessary; neither alone is sufficient.
Data-Tier Architecture Implications
The pseudonymisation-anonymisation distinction has direct implications for how data tiers are designed within a research platform. Two architectural patterns illustrate the trade-offs:
Federated Processing (D-03 Pattern)
In a federated model, pseudonymised data remains at the data source. Analytical workloads are sent to the data rather than bringing the data to a centralised platform. The re-identification key never leaves the originating institution. This pattern minimises data movement and reduces the attack surface for re-identification, because the pseudonymised data and the key are never co-located in a single environment.
The trade-off is operational complexity. Federated processing requires standardised interfaces, consistent data quality across sources, and governance frameworks that span institutional boundaries. It is well suited to multi-site clinical genomics research where data sovereignty is a primary concern.
Centralised Processing (D-04 Pattern)
In a centralised model, pseudonymised data is ingested into a single platform where analytical workloads are executed. This pattern offers simpler operations, better performance for cross-dataset analysis, and easier reproducibility. However, it concentrates risk: the pseudonymised data from multiple sources is co-located, and the platform operator must implement robust controls to prevent re-identification through data combination.
The centralised model requires additional safeguards: data use agreements with each contributing institution, technical controls preventing the joining of pseudonymised datasets in ways that could enable re-identification, and audit trails demonstrating that access to the data was limited to authorised purposes. Under GDPR, the platform operator in a D-04 model is likely a data processor (or joint controller, depending on the arrangement), with corresponding contractual and regulatory obligations.
| Consideration | Federated (D-03) | Centralised (D-04) |
|---|---|---|
| Data movement | Minimal — data stays at source | Data ingested to central platform |
| Re-identification risk | Lower — key never co-located | Higher — requires compensating controls |
| Cross-dataset analysis | Complex — requires standardised APIs | Simpler — co-located data |
| Operational complexity | Higher — multi-site coordination | Lower — single platform |
| GDPR controller role | Each institution retains control | Platform operator as processor or joint controller |
| Audit requirements | Per-site audit trails | Centralised audit trail with access justification |
Special-Category Data: Genetic and Health Data Under Article 9
The GDPR classifies genetic data and data concerning health as special categories of personal data under Article 9. Processing of special-category data is prohibited by default, with a limited set of exceptions. The most relevant exceptions for biotech research are:
- Explicit consent (Article 9(2)(a)) — The data subject has given explicit consent to the processing for one or more specified purposes. In research contexts, the specificity requirement means that broad, open-ended consent forms may not satisfy the standard.
- Scientific research (Article 9(2)(j)) — Processing is necessary for scientific research purposes in accordance with Article 89(1), which requires appropriate safeguards including, where feasible, pseudonymisation. This is the most commonly relied-upon basis for genomic research, but it is not a blanket exemption — the safeguards must be demonstrable.
- Substantial public interest (Article 9(2)(g)) — Processing is necessary for reasons of substantial public interest, on the basis of Union or Member State law. This basis is used in public health genomics and epidemiological research, but requires a specific legal foundation in national law.
Critically, pseudonymisation is explicitly named in Article 89(1) as one of the safeguards that should be implemented when processing personal data for scientific research. The GDPR does not merely permit pseudonymisation in this context — it expects it. Organisations that process genomic data for research without pseudonymisation must justify why it was not feasible, and must demonstrate alternative safeguards of equivalent or greater effectiveness.
"Processing for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes, shall be subject to appropriate safeguards [...] Those safeguards shall ensure that technical and organisational measures are in place in particular in order to ensure respect for the principle of data minimisation. Those measures may include pseudonymisation."
— GDPR, Article 89(1)
National Variations
Member States have enacted varying national implementations of the Article 9 research exemption. Germany's BDSG imposes additional requirements including the appointment of a data protection officer with specific qualifications. France's framework through the CNIL includes sector-specific reference methodologies for health data research. Switzerland, while not an EU Member State, has adopted provisions through the revised Federal Act on Data Protection (revFADP) that align closely with GDPR principles for cross-border data transfers. Organisations processing genomic data across multiple jurisdictions must map these national variations and design their pseudonymisation controls to satisfy the strictest applicable requirements.
Pseudonymisation reduces risk but does not remove GDPR obligations. Pseudonymised genomic data remains personal data, remains subject to data subject rights, and — as special-category data under Article 9 — requires a specific legal basis for processing. Design your data architecture with this reality in mind: the question is not whether to apply GDPR controls, but which controls to apply and how to implement them at the infrastructure level.
Practical Implications for Platform Design
For organisations building or procuring platforms to process genomic data, the pseudonymisation-anonymisation distinction drives several architectural decisions. First, the platform must support key separation as a native capability, not an afterthought. Second, access control models must accommodate the dual-layer structure of pseudonymised data and re-identification keys, with independent policy enforcement for each. Third, audit trail requirements are heightened for special-category data — every access to pseudonymised genomic data should be logged with a purpose justification, and these logs must be reviewable by data protection and compliance functions.
Finally, the platform's data retention and deletion capabilities must support the exercise of data subject rights. Even for pseudonymised data, data subjects may exercise their right to erasure (Article 17), and the platform must be able to identify and delete all data associated with a given pseudonym — including derived data, cached copies, and backup instances. This capability must be designed into the data model from the outset; it cannot be reliably added to a system that was not built to accommodate it.
References & Further Reading
- EU GMP Annex 11: Computerised Systems Primary source for GMP computerised system expectations including validation, audit trails, security, backup/restore, and periodic evaluation.
- FDA 21 CFR Part 11: Electronic Records; Electronic Signatures US regulatory requirements for electronic records/signatures including audit trails, validation, and authorised access controls.
- FDA Part 11 Scope and Application Guidance Official FDA interpretive guidance on the scope and practical application of Part 11 requirements.
- PIC/S PI 041-1: Good Practices for Data Management and Integrity Pharmaceutical inspection guidance on data integrity expectations, ALCOA+ principles, and governance frameworks.
- Swissmedic: EU GMP and PIC/S GMP in Switzerland Confirms Switzerland recognises both EU GMP and PIC/S GMP standards — relevant for Swiss-hosted platform compliance.
- Swiss Federal Act on Data Protection (FADP) Primary Swiss legal text governing personal data processing and protection obligations.
- GDPR Full Text (WIPO Lex) Complete GDPR legal text including pseudonymisation definition (Article 4), genetic data classification, and special-category protections (Article 9).

