One last read. You drop the contract into ChatGPT for a summary. And you assume you only shared the text. In fact, that file carries far more than that: the name of the colleague who left a margin comment, the clause a partner struck out the day before, and your firm's name quietly stored in document properties nobody ever bothers to open. The visible text? Just the top layer.
A .docx is a zipped archive
- The visible text names, addresses, amounts in the body of the document.
- Comments often forgotten, they hold names and internal remarks.
- Tracked changes revisions keep the history and the authors of changes.
- Metadata author, company, file path: invisible on screen, very much there.
Comments and tracked changes: the document's memory
Comments stay between colleagues. That is the whole point: "ask for an extension", "check with the client before signing". They carry their author's name and, more often than not, details you would never put in the body of the text. Tracked changes go further still: until a revision is accepted, the file keeps the old version, the new one, the author of every change and the exact moment it was made. A corrected amount. A replaced name. A deleted paragraph. All of it stays recoverable.
Metadata: the author, the company and the file path
Open the document properties. You will find the author, the company registered in the Office license, the creation date and, sometimes, the full save path that gives away your very login name (C:\Users\first.last\...). Add to that the old versions Word quietly keeps and the usernames tied to every revision. None of it shows on screen. But all of it leaves with the file the moment you send it as is.
The reflex: before the AI, accept or remove revisions, delete comments and clean the document properties (flatten the file), on top of masking the text. Otherwise data slips out through a channel nobody watches.
Prepare the Word, not just the text
You pseudonymize the content. You remove the metadata. You flatten the revisions. The AI only ever sees the masked version, and once its answer comes back, the re-identification happens locally, on your own machine. Nothing identifying gets out. Through any channel.
Safe-Doc handles all of this. It pseudonymizes the Word content and cleans its metadata while keeping your layout, styles and tables exactly as they were. For the wider picture, see the ChatGPT and GDPR at work guide.
The visible text is just one layer. Pseudonymize the Word, clean its metadata, and only then send it to an AI.