Blog · Data & AI · · 1 min read

How much personal data is in a typical document?

People underestimate it. Almost every time. A single work document hides far more personal data than anyone expects, and an ordinary contract or report will easily line up dozens of items: names, dates, addresses, identifiers, amounts, all of it. That is precisely what slips out the moment you paste the file into an AI, with no filter in between. Here is an order of magnitude, by document type.

Personal data in a typical document
Indicative order of magnitude of identifying éléments per document.
Excel export (batch)
30 PII
Médical report
24 PII
Client contract
18 PII
CV / résumé
12 PII
Work email
6 PII

Method: indicative figures, based on typical documents. The real number depends on the document. The point is to show an order of magnitude, not an exact statistic.

Why it is more than you think

The name is obvious. The rest, far less so. Yet a document teems with identifying signals that quietly add up: a date of birth, an address, an IBAN, a file number, sometimes a medical condition or a salary buried in a single line. Taken together, they can pin down a person even when the name appears nowhere at all. Which is the whole point: masking "just the name" is not enough.

What it changes before the AI

More identifying data means more exposure the instant you send a raw file to a third-party model. The logic is relentless. The fix, though, never changes, whatever the volume involved: pseudonymize before the AI, keep the mapping key, then re-identify locally once the work is done. The AI reasons over structure and meaning. Not over identity.

Safe-Doc handles it. It automatically detects and masks these dozens of items, preserves the original layout, processes everything inside the European Union, then purges without keeping a trace. To go further, see the ChatGPT and GDPR at work guide.

See what is detected on your own document. Yours, not a sample. In seconds.

Part of the guide : Fundamentals ↗