Practical guide · GDPR
What is document anonymisation | Guide 2026
What anonymising a document means, how it differs from blacking data out, when it is mandatory under GDPR and the National Security Framework, and which techniques actually work.
Anonymising a document means removing or transforming the personal data it contains so that no person, using reasonably available means, can re-identify the individual appearing in it. It is not the same as blacking out a name or deleting a couple of fields: it is a technical process with a legal criterion behind it, defined by European data protection law.
Legal definition of anonymisation
Under the General Data Protection Regulation (GDPR), data is considered anonymised when it falls outside the scope of the regulation because it is no longer "personal data": it does not identify an individual, directly or indirectly. The Spanish Data Protection Agency (AEPD) summarises it in its own guide on anonymisation and pseudonymisation: data is considered anonymised when there is no reasonable probability that someone can re-identify the person, taking into account the cost, time and technological means required to do so, both current and foreseeable in the future (AEPD blog, "Anonimización y seudonimización").
That last part is often overlooked: a document may look "anonymised" today and cease to be so tomorrow if new ways of cross-referencing data or reversing obfuscation appear.
Anonymising is not the same as redacting
Covering a name with a black rectangle in a PDF, or replacing it with asterisks in Word, is not anonymisation: it is visual concealment. The original text usually remains underneath the graphic layer, in the file metadata, in the revision history or in the selectable text layer of the PDF. The AEPD has fined bodies and companies precisely for publishing documents where the "redacted" data could still be recovered by copying and pasting the text or opening the file with another program (see real cases in our article on AEPD fines for failing to anonymise documents).
Real anonymisation requires:
- Removing the data from the text layer, not just covering it visually.
- Cleaning the file metadata (author, revisions, comments, GPS coordinates in images, etc.).
- Applying the same treatment consistently throughout the document and in all derived formats (if you generate a PDF from Word, you must anonymise before exporting, not after).
Which data must be anonymised?
It is not only about names and surnames. The types of data that usually require anonymisation in administrative, healthcare or legal documents include:
- Name, surnames and signature.
- ID card/passport number, social security number.
- Postal and email addresses.
- Telephone numbers.
- IBAN and banking details.
- Vehicle licence plates.
- Health data, ethnic origin, trade union membership or religious beliefs (special categories with a higher level of protection under Article 9 GDPR).
- Data relating to minors.
When is anonymising a document mandatory?
In practice, the obligation arises whenever a document containing personal data is going to:
- Be published (transparency portals, gazettes, resolutions, minutes).
- Be shared with third parties that have no legal basis for processing that specific data (suppliers, other administrations, collaborating firms).
- Be used to train or feed AI systems, where the risk that the data "leaks" into future model responses is an additional reason for caution.
- Be kept beyond the period for which the data was collected, when the original purpose no longer applies but the document has historical or statistical value.
For public bodies, this connects directly with the Spanish Transparency Act (Ley 19/2013) and the National Security Framework (ENS): the higher the security category of the system managing the information, the more the processing —including anonymisation before publishing or sharing— must be documented and auditable. We develop this in detail in document anonymisation in the Public Administration.
Anonymisation vs. pseudonymisation: the most common confusion
They are distinct concepts with distinct legal consequences: pseudonymisation replaces the data with a reference (a code, a hash) but allows identity to be recovered if the key or additional information is available, so it remains within the GDPR. Anonymisation, if properly done, takes the document outside the scope of the regulation because it is no longer reversible. We explain this in detail in anonymisation vs. pseudonymisation under the AEPD.
Frequently asked questions
Is it enough to delete the name and leave the rest of the text unchanged?
Not always. If the rest of the document contains enough context (position, dates, location, a very specific case), it may be possible to re-identify the person by combining those data, even if the name does not appear. This is known as inference re-identification risk.
Is anonymisation reversible?
It should not be. If the process allows the original data to be recovered with a key or additional record, it is not anonymisation: it is pseudonymisation, and it remains subject to the GDPR.
Is manually anonymising in Word or Acrobat sufficient?
It may be for occasional documents if done correctly (removing the actual text, not just covering it, and cleaning metadata), but it becomes unworkable and error-prone for volume: case files, notifications or recurring resolutions. That is where automation makes sense. For a practical guide, see how to anonymise a PDF step by step.
Sources and references
Sources and references
Everything on this page is based on published legislation and on the guidance of the Spanish Data Protection Agency. These are the original texts:
Do you anonymise documents daily?
Stop redacting by hand. Automate it with anonimizia.
Upload your PDFs and get GDPR-compliant anonymised documents in seconds.
Try it free