Two weeks after Anthropic began watermarking Claude’s output under the EU AI Act, tools claiming to strip those marks are spreading faster than anyone can check whether they work. The most popular, an open-source project called watermarks-remover, stood at 13,000 GitHub stars on Monday.

Anthropic disclosed on August 2 that new Claude models embed invisible watermarks in generated text and attach signed provenance metadata to generated files. The date was not the company’s choice. Article 50 of the EU AI Act became enforceable that day, requiring machine-readable marking of AI-generated content, with penalties up to 15 million euros or 3 percent of global turnover. Anthropic has said the mark indicates content was processed by Claude, not necessarily authored by it.

What Anthropic has not done is publish how the text watermark works or release a detector. That absence defines the current moment: removal tools can claim success, and nobody, including their authors, can verify it.

The authors largely admit this. The watermarks-remover README describes its statistical-watermark layer as best-effort rewording, warns that it degrades the copy, and states plainly that no tool can honestly certify its output fails the official check. Guillaume Meyer, its developer, has acknowledged the tool reliably removes only what is checkable: invisible Unicode characters and file metadata.

The checkable layers are also the weakest. Zero-width characters strip with a script. The C2PA metadata attached to files disappears with a screenshot, a format conversion or a social-media upload. The statistical mark woven into word choice is the only layer with teeth, and it is precisely the one nobody outside Anthropic can test. Paraphrase is believed to defeat it, which every rewording tool relies on but none has demonstrated.

Around the flagship project, a market has formed: claude-watermark-cleaner and remove-ai-watermarks on GitHub, web services with names like claudewatermark.com, and established paraphrasers such as StealthGPT adding Claude-specific removal features, all selling an outcome none can measure.

There is a quieter risk in the packaging. Several of these tools install as agent skills, which means granting code execution and document access to software written to defeat provenance controls. Agent-facing supply chains have been productive attack surface this year, from split instructions that doubled compliance rates in coding agents to a single email planting persistent false memories in assistants.

The marking requirement sits inside a wider relationship. The same August 2 date brought the Act’s high-risk provisions into force, and Brussels has separately been seeking access to Anthropic’s Mythos for its cybersecurity agency.

The verification gap closes only when Anthropic ships its detector. Until then, compliance marking and its countermeasures are both scaling on faith.

Sources: BleepingComputer, GitHub

–
By the Control Plane Editorial Team