The mark flags anything Claude processed, even human writing it only edited.
"Look at that subtle off-white coloring. The tasteful thickness of it. Oh my God, it even has a watermark." Credit: Aurich Lawson | American Psycho (Lions Gate Films)
Anthropic has revealed that it will soon watermark content that is processed (not just generated!) by any of its models. In a support article, Anthropic explained that it was rolling out machine-readable watermarks to comply with the European Union’s AI Act, which requires all AI system providers to watermark AI-generated or manipulated audio, image, text, and video outputs. The law applies to any AI model released after August 2 and provides a grace period until December 2026 for providers to update previously released models.
Anthropic confirmed that moving forward, all new models offered globally—not just in the EU—will mark AI-generated content “from day one.” Text outputs will “carry embedded watermarks,” invisible to the user, and other “generated files will include digitally signed provenance metadata where supported,” Anthropic said.
Notably, Anthropic is deploying a “nuke it from orbit” approach, applying the watermarks to all processed content where supported, even though the EU does not require it for cases where an AI system performs “an assistive function for standard editing” (the guidance’s own example is grammar correction), or where it doesn’t “substantially alter” the user’s text or its meaning.
A watermark applied at the model level can’t tell wholesale generation from a comma fix, so Claude may end up stamping exactly the content the law was written to leave alone. How thoroughly it truly watermarks will not be known until Anthropic releases a detection tool that can be tested. The company said that it plans to eventually share details about how to detect marks in order to offer technical support that the EU’s law requires.
Anthropic also noted that the watermarks won’t work on “some platforms or features” that don’t support them. For non-text content, Anthropic will use the C2PA metadata approach to record provenance.
Easy to get around, easy to misunderstand
The approach described by the EU and implemented by Anthropic is unfortunately trivially easy for bad actors to bypass, while potentially punishing users who trust the system to accurately label their outputs. Text watermarks work by biasing the model’s word choices in a pattern spread across the entire document, only detectable in aggregate by the right tool. The catch is that “invisible” can also mean the model occasionally trades the best word for a slightly worse one, just to keep the signal intact.
Anthropic noted that those marks “will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.” But if watermarked text is pasted into another chatbot system that edits the text, the watermark could be destroyed. With image and video content, screenshotting/recording or using any decent metadata editing tool will suffice to remove this information, too. And once Anthropic tells the world how to identify these watermarks, building a system to remove them would be trivial.
Further, the potential for misinterpretation seems high; the watermark is not particularly informative. Anthropic explained that a “detected mark provides a signal that content was processed by Claude, but is not fully conclusive.” The only real message the mark sends is that “the content may have been processed by Claude,” Anthropic said, and the mark may even appear on content that was not generated by Claude.
On top of this, you have the general public, who may not grasp the difference between processed text and wholly generated text. If the system watermarks human-authored text simply because it was edited in a workflow that touches Claude, suddenly it carries the same denotation as wholly generated text does. And all of this in a system where the “lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed,” Anthropic said.
Claude may label more AI content than required
To its credit, Anthropic acknowledges that it may be marking some content that the AI Act does not require to be labeled: “People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source.” So, despite being explicitly exempted by the law, Claude will mark any editing work done on original writing.
If the watermarks become a catch-all for any Claude use, from spellchecks to complete rewrites, Anthropic’s solution will likely frustrate users by going further than necessary to label content in ways that could inadvertently muddle provenance. A teacher or professor, for instance, should not interpret this watermark signal as anything beyond “AI touched this,” but it is easy to imagine they will. After all, what is the point of a watermark that cannot differentiate between light editing and wholly generated text?
The matter becomes even murkier when we turn to another section of the EU AI Act, 50(4), which addresses how those publishing such text must explicitly label content. Under their labeling regime, wholly generated AI text does not require a label in most instances. AI-written novel? No label. Marketing copy generated by AI? Nope. But if the text is meant to “inform the public on matters of public interest,” you have to label it unless a human reviews it editorially (meaning, the editor is known and accountable). So, we have a strange situation where, on the model level, it’s “watermark all the things,” but on the public disclosure front, it’s “you don’t need to label this AI text if Joe looked at it.”
Ars reached out to Anthropic to see if there’s a timeline for details on detection to be released or results from any testing the company can share assessing the likelihood for false negatives or positives. We also asked if Anthropic could address how its watermarks may conflict with standard editing and other exemptions from the AI Act, but Anthropic sent a statement that did not address Ars’ questions.
“We’re adding marking to Claude’s output to comply with the EU AI Act, and other labs are taking similar steps,” Anthropic said. “It’s hard to identify AI-generated text, and this gives people better tools for identification. Text from supported Claude models, including output from Claude Code, will carry an invisible watermark, and it doesn’t change the meaning, quality, or readability of Claude’s responses. We also plan to ship a text detection API so users can do more of this themselves.”
The fun is just starting
In the EU, transparency requirements are meant to ensure AI tools like Claude don’t upset “the integrity and trust in the information ecosystem, raising new risks of misinformation and manipulation at scale, fraud, impersonation, and consumer deception.” One EU support article forecasted that the obligations would be the “primary compliance challenge” for many AI firms.
“People should know when they are interacting with AI or exposed to AI-generated content,” the European Commission’s guidelines said. “This will help them make informed decisions, calibrate their trust and reliance on AI, and avoid misinformation or deception.” Still, it’s hard to square this with the fact that a wholly generated article on a matter of public interest gets a watermark, but not a reader-facing label, if an editor properly reviews it.
AI firms like Anthropic are best positioned to develop watermarking solutions, the EU expects, since AI moves fast and there will be an ongoing “need for new methods and techniques to trace origin of information.”
But that largely leaves the societal value of such marks up to tech firms to decide, with the EU only stipulating that “techniques and methods should be sufficiently reliable, interoperable, effective and robust as far as this is technically feasible.”
In its post, Anthropic said it plans to continue working on its watermarks and detection methods that meet the EU’s demands. If Claude’s labels fail, the AI Act carries steep penalties for violations, including fines up to 15 million euros, or 3 percent of a company’s worldwide annual revenue.
Ars Editor-in-Chief Ken Fisher contributed to this report. This story was updated on August 13 to add a statement from Anthropic.










English (US) ·