AI Watermarking Defeated in Hours as Anthropic Claude's Text Labels Face Instant Workarounds

Reviewed byNidhi Govil

9 Sources

Share

Anthropic introduced invisible watermarks on Claude-generated text to meet EU AI Act requirements by December 2024. Developer Guillaume Meyer published a bypass tool within 4 hours that's been bookmarked over 20,000 times. Experts warn watermarking generated text won't work due to easy removal through light edits, translation, or open-source models.

Anthropic Claude Deploys Invisible Watermarks to Meet EU AI Act Deadline

Anthropic announced that its Claude models will embed invisible watermarks into AI-generated text and files, marking a significant shift in AI transparency efforts

1

2

. The move comes in response to the EU AI Act, which came into force on 2 August and requires all AI models to provide watermarking by 2 December this year

1

. Claude models launched on or after 2 August now include watermarking across all products, including Claude API, Claude Code, Claude Cowork, and Claude Tag

4

. The EU's Code of Practice on Transparency of AI-generated Content has been signed by 190 organizations, including OpenAI, Microsoft, and Meta

2

.

Source: CNET

Source: CNET

The watermarking system uses Google's SynthID technology, which creates a detectable signal by leaving patterns in Claude's choice of words and phrases that remain indiscernible to human readers but can be identified by machines

2

4

. Large language models typically create sentences by selecting statistically likely words, and by varying this selection—perhaps alternating between the most likely and second most likely word—models can embed patterns without changing the text's meaning

1

. For images, Claude attaches cryptographically signed notes in metadata of files like .png, .jpg, or .svg, following the C2PA standard used in photo-editing software

4

.

Source: New Yorker

Source: New Yorker

Developer Bypasses Watermarks Within 4 Hours of Announcement

Within four hours of Anthropic confirming the watermarking implementation, developer Guillaume Meyer published circumvention tools on GitHub

2

. Meyer's code has been bookmarked more than 20,000 times on X and has drawn more than 100 contributors, with many incorporating the technology into their own projects

2

. The removal method uses a non-watermarking large language model to generate multiple rewrites, swapping in synonyms and slightly reorganizing content

2

.

Source: New Scientist

Source: New Scientist

Other developers quickly followed suit. Software engineer Erik Hughes created a tool in 15 minutes that removes invisible and look-alike characters, reorders sentences within paragraphs, and swaps several words for synonyms

2

. Leon Chlon, a Visiting Fellow at the University of Oxford, demonstrated that watermarks can be removed by condensing Claude's response, translating it into Arabic—which has very different semantics compared to English—and then translating it back

2

. Wayne Pan, CTO and cofounder at Silicon Valley-based sovereign AI startup Haimaker, incorporated Meyer's open-source tool into his platform

2

.

Experts Warn Watermarking Generated Text Faces Fundamental Limitations

James Padolsey at NOPE, an AI safety firm, explains that while image and video watermarks are more reliable because they're composed of dense and multi-layered data, watermarks in text will always be more difficult to implement

1

. Padolsey created an online tool called declaude that takes AI-generated text and lightly changes it so existing AI detectors no longer work, demonstrating that watermarking patterns are susceptible to deletion by even light edits

1

.

The availability of open-source models presents another challenge to AI watermarking effectiveness. "Any bad actors will just use those. And they'll probably be cheaper and easier to run," Padolsey notes

1

. These open-source AI models don't abide by the EU's rules because they aren't controlled by any one tech firm

1

. Anthropic itself acknowledged that heavily edited, paraphrased, or translated content might not carry a watermark

2

.

False Positives Raise Concerns About AI Content Attribution Accuracy

Peter Scarfe at the University of Reading, UK, highlights concerns about false positives—when text is incorrectly identified as AI-generated or flagged because a model had even a minor role in producing work

1

. "If it says there's a 67.8 per cent chance that this text was generated in some way by AI, as an educator, how would I act upon that? I'm not really sure," Scarfe says

1

.

Meyer expresses similar concerns about the risk of false positives and the inability of watermarking to distinguish between light or heavy AI use

2

. As a native French speaker who often uses Claude and other AI tools like Grammarly to edit his writing, Meyer worries that using the watermark as evidence could lead to employers unfairly rejecting candidates or overblown accusations of researchers using artificial intelligence

2

. Watermark-detection tools could flag text that is simply processed by AI, not just text entirely generated by it—meaning a student could face consequences even if they only asked a model to scan their manually written work for grammatical mistakes

1

.

AI Transparency Debate Intensifies as Users React to Labeling AI-Generated Content

The announcement sparked swift backlash from AI users, with some reportedly canceling their Claude subscriptions in protest

3

. Freelance content writers and social media creators contacted Meyer asking for assistance using the bypass code

2

. In 2024, Slack found that 48% of desk workers said they would be uncomfortable disclosing they used AI

3

.

Only 4% of US adults are using AI chatbots "constantly," according to a June Pew Research Center report

3

. Yet many (51%) of US adults want better labels on social media for AI-generated content, CNET found in February

3

. Across the internet, criticism of the allegedly broken contract between machines and their human prompters ballooned, with users questioning whether L.L.M.s were specifically advertised as being able to present a machine's work as one's own

5

.

Combating Misinformation Remains Central Goal Despite Implementation Challenges

A European Commission spokesperson stated: "These rules will help people recognise when they are interacting with AI or when content has been generated or altered by AI. The adversarial robustness of marking and detection solutions must be assessed in terms of resilience to malicious behaviour, such as copying, removal, regeneration and modification attacks on the markings"

1

.

Scarfe acknowledges that AI can be used for various purposes and doesn't inherently spread misinformation, suggesting a blanket policy on watermarking might be a blunt instrument

1

. However, he notes AI can make mistakes and hallucinate facts, and can be used to disseminate lies quickly—such as by powering disinformation bots on social media

1

. "I don't think [watermarking] is a magic bullet. I don't think it's going to suddenly solve what the EU maybe are hoping it's going to solve," Scarfe concludes

1

. The development highlights ongoing tension between regulation compliance and the cat-and-mouse game between AI providers and users seeking to circumvent provenance tools.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved