Practical guide · GDPR

AI tests privacy: anonymise before publishing, sharing or asking

9 min read

Every day we use more information to work, and we share it more too. We publish documents on transparency portals, hand them over in response to access requests, send them to suppliers, add them to case files and, more recently, copy entire documents into artificial intelligence tools to summarise, analyse or draft a reply. The problem is that personal data still travels inside those documents. The arrival of AI makes it more urgent than ever to anonymise before publishing, sharing or asking.

AI did not create the problem, but it has changed its scale

David Bollero recently explained this in Público in the article “La IA pone patas arriba la privacidad”, published on 15 August 2026. The article focuses on a particularly relevant issue: the inference capability of AI systems allows apparently innocuous information to be linked and, in certain scenarios, reconstruct personal information or facilitate re-identification.

Until now, when we talked about protecting personal information, we tended to think of obvious identifiers: first and last name, ID number, phone, address, email, account number. Removing these is still essential, but no longer always enough. A combination of dates, locations, job titles, personal circumstances, economic information or references contained in different documents can identify a person even if their name has been removed.

The UK Information Commissioner's Office and other European regulators remind us that removing direct identifiers is only a first stage, and that a de-identified dataset can be re-linked to a person when combined with other available information. So anonymising should not be understood simply as crossing out a name.

Before worrying about re-identification, we have a simpler problem

On many occasions we are giving data directly without anonymising it. Think of everyday situations:

  • “Summarise this case file.”
  • “Analyse these submissions.”
  • “Show me the differences between these contracts.”
  • “Draft a reply based on this documentation.”
  • “Extract the important information from this PDF.”

Artificial intelligence can save hours of work. The problem appears when that PDF contains names, ID numbers, signatures, phone numbers, addresses, bank accounts, identification codes, economic information, social circumstances or any other personal data that was not necessary to provide for that task.

The debate should not be reduced to whether a particular AI tool is safe or unsafe. There are very different solutions, configurations and corporate environments. The previous question should be much simpler: does the tool really need to receive that personal data to do what we are asking? If the answer is no, that data should be removed or anonymised first. This is a practical application of the GDPR data minimisation principle: work only with the information necessary for a specific purpose.

AI is only one destination for the document

This issue does not only appear when someone uses ChatGPT, Copilot, Gemini or another tool. The same documentation can end up:

  • published on a transparency portal;
  • available on an electronic office;
  • added to a public board or platform;
  • handed over following a freedom of information request;
  • sent to a third party;
  • reused internally for new purposes.

Transparency and data protection must coexist. Freedom of information law recognises the right of access, but when that information contains personal data, the limits derived from its protection must also be analysed. Regulators expressly remind us of this need for balance and distinguish between access to information cases and active publication cases.

Therefore, anonymising correctly does not mean hiding everything. It means identifying what information can be shown, what must be protected, and doing so before the document leaves the environment in which it was created.

The challenge is doing it at scale

Manually reviewing ten pages is feasible. Reviewing hundreds or thousands of documents before publishing, handing over or processing them is another matter. Moreover, personal data does not always appear in the same way or place: it can be in the body of the document, tables, annexes, spreadsheets or metadata.

That is precisely where a technological solution is needed. That is why we developed anonimizia: a tool created to help organisations and public administrations detect and anonymise sensitive information in documents before they are published, shared, delivered to third parties or used in other digital processes.

Anonymise first, not afterwards

For years we have talked about privacy mainly in terms of databases, cyberattacks or leaks. Now we need to add another question to the normal workflow: what data does this document contain before sending it to the next system? Before publishing it. Before handing it over. Before sharing it. And yes, also before uploading it to artificial intelligence.

Because the greater our ability to process and relate information, the more important it becomes to apply privacy from the source. Artificial intelligence is changing the rules. Data protection must evolve too.

Reference

Source that inspired this reflection: David Bollero, “La IA pone patas arriba la privacidad”, Público, 15 August 2026. Additional references: UK Information Commissioner's Office and European data protection authority guidance on anonymisation and transparency.

Do you anonymise documents daily?

Stop redacting by hand. Automate it with anonimizia.

Upload your PDFs and get GDPR-compliant anonymised documents in seconds.

Try it free