A colleague forwarded me a forty-file set to run past an AI. Articles of association, shareholder agreements, emails, account spreadsheets. The instinct? Open each document, mask the names by hand, move to the next. Except by the tenth file, without even noticing, you are no longer masking quite the way you did on the first, and the AI has no way of knowing that "the company in file 3" and "the company in file 28" are one and the same party. Batch processing exists for exactly this. You drop the whole set, and every person, every company gets the same pseudonym everywhere.
- Each file opened and masked separately
- Inconsistent tokens from one file to the next
- Risk of missing a document in the batch
- Slow, so often rushed
- You drop several files together
- The whole batch anonymized at once
- Consistent tokens across files (same party = same token)
- Key kept, ré-identification at the end
The same name becomes the same pseudonym, across every file
This is the heart of it. When you pseudonymize a batch in a single pass, Smith becomes PERSON_1 in the contract, in the email and in the spreadsheet, the very same label, from the first file to the last. The AI can then follow that party from document to document, compare what they sign here against what they claim there, without ever reading a real name. Masking each file on its own breaks that thread. The consistency is lost. And cross-analysis with it.
On a large set, volume is the real danger
Forty files by hand means forty chances to miss something: a signature at the foot of a page, an address in a header, a name tucked into a footnote. The bigger the set, the more review tires. And the more it lets slip. Dropping the whole set together changes that, because the same detection applies at the same standard to every page, with no drop in attention by the thirtieth document.
Watch the cross-file re-identification risk
A harmless-looking detail, on its own, says nothing. Cross-referenced with another file in the set, it can re-identify someone: a date in one, a job title in another, an amount in a third. Treating the set as one coherent whole, with pseudonyms that stay stable from piece to piece, closes that door. What identifies in one file identifies in the others. So it is masked everywhere, the same way.
The method, on a complete set
The move is simple. You drop all the files together, the whole batch is pseudonymized in one pass, then you work with the AI on the masked version: summary, red flags, comparison across documents. After that, you re-identify the deliverable locally. The mapping key never moves: it stays on your side, encrypted, and nothing goes to a third-party AI.
Safe-Doc does exactly that. It handles several files together (PDF, Word, spreadsheets) with consistent pseudonyms across the whole set, keeps the original layout and processes your documents in the European Union before purging them. Handy on large sets. In due diligence, for example.
Drop a whole set at once. Consistent pseudonyms across every document. The key, kept on your side.