Practical guide · GDPR

How to remove PDF metadata before publishing a document

7 min read

You can redact every name in a PDF and still publish personal data: metadata stores the author, the machine user, the software, creation and modification dates, and even GPS coordinates in scanned documents.

What metadata a PDF carries

  • Author and creator: usually the employee's real name or network user.
  • Title and subject: sometimes the data subject's name or a case number.
  • Dates of creation and modification.
  • Source application and file path.
  • EXIF data in embedded images, including camera model and GPS.
  • Layers, comments and bookmarks holding text you thought was gone.

How to check it

  1. Windows: right-click the file → Properties → Details.
  2. Any viewer: File → Document properties.
  3. Review comments, bookmarks and layers too.

How to remove it

  1. Clear the fields: author, title, subject, keywords.
  2. Flatten the PDF (print to PDF or export) to drop layers, form fields and comments.
  3. Inspect the document in Word before exporting.
  4. Re-check the final file's properties — the step most people skip.

Metadata and the GDPR

Metadata containing a person's name is personal data. Publishing a PDF whose author is identifiable is an unintended processing operation and a common source of complaints.

Do it in one step

anonimizia removes personal data from the content and strips the metadata of the resulting file, with an auditable report.

Do you anonymise documents daily?

Stop redacting by hand. Automate it with anonimizia.

Upload your PDFs and get GDPR-compliant anonymised documents in seconds.

Try it free