An amendment lands. And the question never changes: what moved since the last version? A non-compete clause stretched out, a liability cap raised, a deadline pushed back three months: AI is brilliant at hunting down that kind of gap, line by line, faster and more thoroughly than any manual read-through ever could. But wait. Before you paste both files into a chat window, ask yourself one simple thing: to compare clauses, does the AI really need to know who is signing? No. It needs the legal skeleton, not the identity of the parties.
The same party carries the same token in both versions. The AI spots clause differences without ever seeing a real name.
Compare clauses, not identities
It all lives in the text. A gap between two contracts surfaces in the body of the clauses: an obligation added here, an amount changed there, a deadline suddenly cut short. The supplier's name? Its company number? Neither one matters. So you can pseudonymize the parties, the contact details, the amounts and the sensitive dates, then let the AI work calmly on the legal skeleton, while the identity of the signatories never once leaves your side.
Token consistency does all the work
Here is the trap. When you mask two files on their own, the same company can turn into ORGANISATION_3 in version 1 and ORGANISATION_7 in version 2, and then everything breaks: the AI thinks it is staring at two different parties and the comparison falls apart. The rule fits in one line. The same actor keeps an identical token from one document to the next. It is that consistency, and that alone, never the real name, that lets you trace a clause from one version to the following one.
What stays in clear text: the matter to compare
Nothing essential gets touched. Not the clause text, not the article headings, not the logic of the cross-references: everything you actually want to set side by side stays intact, word for word, right where it belongs. Only the parties' names, the addresses, the internal references and the amounts you deem confidential disappear. The outcome? A perfectly comparable contract, stripped of whatever identifies it.
In practice, across two versions
The flow is short. You pseudonymize both versions with consistent tokens, the AI draws up the list of differences, then you re-identify the report locally, on your own machine, out of reach. The real names reappear only on your side. The tokens stay reversible for you; the AI itself never saw a single identifying detail.
Safe-Doc does exactly that. It masks several documents with consistent tokens before the AI, preserves the original layout and processes everything in the European Union before purging. To dig deeper, see the ChatGPT and GDPR at work guide.
Compare without exposing. Consistent tokens from one version to the next, and the parties' identity that never leaves your machine.