Skip to content
Dustin's AI Lab
Go back

Anthropic, Same Week: Safety Report Backfires, De-identification Goes Full Auto

Anthropic pulled off two opposite moves in the same week: a safety report that backfired on customer conversation privacy, and a feature that let me push my de-identification tool to full auto-block.

Anthropic had a genuinely funny week. They did two completely opposite, contradictory things at the same time.

First, they patted themselves on the back and published a safety report, and it backfired the moment people realized: wait, you’ve been casually looking at customer conversations this whole time? Palantir and NVIDIA both went pale.

Second, they shipped a powerful Claude Mod that let me finally push my de-identification tool, pii-guard-tw to its full auto-block form. You don’t need to touch anything by hand anymore, just install the updated plugin. From then on, on Anthropic’s servers, the model only ever sees <PERSON_1>, <TW_MOBILE_1>, <TW_ID_1>. What comes back to you is always the restored, real content.

Here’s the mechanism:

A three-step diagram showing pii-guard replacing real PII with placeholder codes before AI sees it, then restoring the real content locally

The flow runs in three steps. Step one happens on your own machine: your files (client: Chen Dawen, phone: 0912-345-678, ID: A123456789) and whatever text you type both pass through pii-guard first, which detects them locally and swaps them for safe placeholder codes before anything gets sent to the AI. Step two is what Claude actually sees: just placeholders, client <PERSON_1>, phone <TW_MOBILE_1>, ID <TW_ID_1>. No real PII, ever. It does the work and sends the result back. Step three is back on your machine, writing the file restores the real content automatically. Your real data never leaves your own computer.


15-Minute AI Adoption Diagnosis

Fill in a short form to book your free 1-on-1 diagnosis, and find out how you can work with AI without the pain and get a real productivity lift.

Book my free diagnosis
Share this post on:
Previous Post
Zero Findings: Nothing Wrong, or the Checker Isn't Checking
Next Post
How I Split My AI Subscriptions: Fable Buys Judgment, Luna Max Buys Value