1 hr ago
Anthropic's AI Watermark Survives Copying but Fails Against Rewriting
Anthropic has created a hidden mark for some content made by its Claude AI.
For text, the mark is hidden in the pattern of words the AI chooses.
Copying the text usually keeps the mark.
But another AI can rewrite or translate the text, making the mark disappear.
Images can also contain digital information showing where they came from.
Saving an image with software that removes extra information can erase that record.
These tools may catch people who use AI content without trying to hide it.
However, finding no mark does not prove that a person wrote the content.
Anthropic says Claude-generated text has carried an invisible statistical watermark since 2 August.
The text signal survives copying and some editing because it is embedded in word-selection patterns.
Rewriting, translating, or thoroughly paraphrasing the text removes the watermark signal.
Claude-generated images carry C2PA credentials, but software that discards metadata can remove them.
The watermark may identify unchanged AI output, but a negative result does not prove human authorship.
- Who
- Anthropic and its Claude AI models.
- What
- Anthropic added invisible text watermarks and C2PA Content Credentials to certain Claude-generated content, while documenting how the signals can be defeated.
- Where
- In Claude-generated text and image files, including when the content is copied between applications.
- When
- Since 2 August; Claude Fable 5.1 and Mythos 5.1 included the marks at launch, as did models released afterward.
- Why
- To support AI-content detection and provenance, while providing infrastructure that could be used for future disclosure requirements.
Reasons to use the system
Limitations and risks
Practical detection
Reasons to use the system
The watermark can identify straightforward cases where Claude text is pasted unchanged, including content used by students, employees, or contractors who are not trying to conceal its origin.
Limitations and risks
Anyone can defeat the text signal by asking another model to rewrite, translate, or paraphrase the passage.
Provenance infrastructure
Reasons to use the system
C2PA provides a provenance standard that could support future rules requiring disclosure of AI-generated content.
Limitations and risks
Image credentials can disappear when files are re-saved through software that does not preserve metadata, and metadata stripping is routine.
Meaning of a negative result
Reasons to use the system
A positive detection can provide useful evidence that content came from Claude.
Limitations and risks
No detection cannot reliably show that content was written by a person, and treating it that way could make institutional decisions less reliable.
Key facts
- Text watermark
- A statistical signal embedded in the sequence of words selected by the model.
- Text durability
- The signal survives unchanged copying and may survive some editing.
- Main limitation
- Rewriting, translation, or thorough paraphrasing by another model removes the signal.
- Image provenance
- Generated image files carry cryptographically signed C2PA Content Credentials.
- Image limitation
- Software that does not preserve metadata can remove the credentials.
- Deployment date
- The marking began on 2 August.
- Interpretation risk
- A positive detection is informative, but an absent signal does not establish human authorship.








