Anthropic Watermarks Claude-Generated Text and Files to Identify AI Content

Anthropic is adding machine-readable watermarks to text and files written or processed by Claude to improve transparency of AI content. The invisible text watermark is robust to copying and some alteration and files may include signed provenance metadata based on the C2PA standard.

Anthropic watermarks Claude-generated text and files to identify AI content
Anthropic watermarks Claude-generated text and files to identify AI content

Worries about the proliferation of low-quality AI-generated content, or "AI slop", have persisted ever since ChatGPT and Claude entered the mainstream. Also, it's not always easy to tell if someone used AI to write something.

Now Anthropic is watermarking all of Claude's output, claiming that users won't be able to detect it. To better assist users in identifying AI-generated material, Anthropic has now revised its support site. In an effort to meet the openness standards set out by Article 50 of the EU's AI Act, the business claims it is marking information produced by Claude models with machine-readable indicators.

How Anthropic’s New Watermarking will Work?

There are two ways this will operate, says Anthropic. A supported Claude model can incorporate a subtle watermark into text without the user even realising it's happening. Users will not be able to see this watermark, according to the business. However, users can copy and paste the text to another document without losing it, and it may even "persist through some editing."

However, neither the content nor the quality of Claude's reply are affected by the watermark. When it comes to other file kinds, such as .svg, .png and .Claude will include signed provenance metadata with the attached jpg. In order to document information regarding the origin of content, the metadata adheres to the C2PA open standard. According to Anthropic, a signed metadata label can indicate that Claude processed a file and reveal whether the file has been altered. There were a variety of responses to this update on social media. Some considered terminating their memberships in protest after seeing Claude watermark everything.

Anthropic Working on Watermark Detection Mechanism

According to the firm, they are actively trying to make it possible for consumers and third parties to discover the provenance metadata and embedded watermarks in Claude, as mandated by EU law. The detection process would look for the presence of a supported Claude mark in a given file or text. Details on detecting mechanisms will soon be shared, according to Anthropic. The new system, however, is not without its limitations. A discovered mark indicates that content may have been handled by Claude, according to Anthropic, but it does not prove entire provenance.

People use the program to proofread, translate, summarise, or convert files. Hence marked content may also be altered, excerpted, or mixed with other material after Claude processes it; and even if a mark is identified, Claude might not be the original creator. A lack of a recognised mark, according to Anthropic, does not rule out the possibility that the material was processed or created by AI. If the content is from a previous model that did not enable marking, or if the text has been significantly modified, paraphrased, or blended with other writing, then a mark might not be there.