In 2026, a company's biggest GDPR risk is not some exotic cyberattack: it is an employee pasting a contract, a CV or an accounting export into ChatGPT to save time. No filter, no trace. Banning AI does not work, people just keep doing it quietly. This guide explains how to frame that usage: what the GDPR says, the method, and what to mask depending on the document.
The reality: shadow AI is already here
In most organizations, staff already use generative AI on internal documents. It is productive, but hard to audit. The right reflex is not punishment, it is to provide a gateway: anonymize before the AI, keep control, re-identify on the way back if needed.
What the GDPR says
Sending a document to a third-party model is a processing of personal data. Pseudonymization (article 4.5) strongly reduces exposure without taking the data out of the GDPR; anonymization (irréversible) takes it out of scope if it is real and lasting. Health data (article 9) is sensitive and demands extra care. In all cases: minimization, a legal basis, and no uncontrolled transfer.
The universal method, in three steps
Whatever the document: 1) pseudonymize before the AI (identifying data becomes tokens), 2) use the AI of your choice on the masked version, 3) re-identify the result locally, the key staying with you.
What to mask depending on the document
A quick reminder by common document type:
| Document type | Typical sensitive data | Before sending to the AI |
|---|---|---|
| Client contract | Names, amounts, IBAN, références | Pseudonymize (réversible) |
| CV / HR file | Name, age, photo, address | Anonymize identity, screen on skills |
| Médical report | Patient, national ID, dates, conditions | Pseudonymize, médical confidentiality |
| Accounting export | SIREN, balances, salariés | Pseudonymize, generalize amounts |
| Client email | Name, contacts, attachment | Mask before pasting into the AI |
By sector
The stakes vary by sector. See the dedicated pages: law firms, accounting, HR, healthcare.
Mistakes to avoid
Believing detection is perfect (human review needed), forgetting scans and screenshots (OCR then mask them), confusing blur with masking (the text stays extractable), and storing the key anywhere (it must stay encrypted, under your control).
Safe-Doc pseudonymizes or anonymizes before the AI, keeps the layout, processes in the European Union then purges, and lets you re-identify locally. A GDPR gateway between your teams and the LLMs.
Test on a real document, réversible pseudonymization, layout preserved.