IT Brief Asia - Technology news for CIOs & IT decision-makers
Asia
Anthropic to add text watermarks to future Claude models

Anthropic to add text watermarks to future Claude models

Wed, 26th Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Anthropic will add text watermarks to outputs from future Claude models as part of its compliance with the EU AI Act.

It is applying the watermark globally at launch because it does not yet have a reliable way to limit the measure by region. Several other major AI model developers have signed the same EU Code of Practise on Transparency of AI-Generated Content and are introducing their own marking systems.

The move reflects a broader push by regulators and AI companies to identify whether generative AI tools were involved in producing content. Anthropic's system is designed to estimate the likelihood that Claude was used to write a piece of text, rather than make a definitive claim about authorship.

The watermark works by influencing some of the low-stakes word choices a model makes as it generates text. The process does not insert hidden characters, add visible markings or require extra tokens.

Instead, the pattern is embedded in the sequence of words the model selects. Anyone with the relevant key can then test whether that sequence is consistent with text produced under the watermarking system.

Anthropic said its approach is based on SynthID-Text, a method published by Google DeepMind. It described the technique as changing the source of randomness used when a model chooses between words that are similarly suitable in context.

Internal testing found no effect on output quality, creativity or readability. Anthropic also cited Google DeepMind research that found no statistically significant difference in user ratings between watermarked and unwatermarked model outputs.

Limits remain

Anthropic also outlined several limits to the system. A watermark can only indicate how likely it is that Claude was involved in writing text; it cannot establish whether material was written by a human or by another AI model.

Detection is also less reliable on short samples because there are fewer word choices to analyse. Factual writing, proofreading and code present further constraints because these tasks often require exact wording, leaving less room for the model to make interchangeable choices.

That means a lightly edited document may carry little or no detectable signal if most of the wording came from a person rather than the model. By contrast, a translation generated by Claude would carry a watermark because the model selects every word.

Code would generally contain less watermarking than ordinary prose for the same reason. Where there is little flexibility, such as in mathematical completions or precise syntax, the system does not push the model toward an alternative token.

The watermark does not contain identifying information about users, organisations or chats. Anthropic also said the process has a negligible effect on model speed and does not alter the cost of serving or using Claude.

Detection plans

Anthropic said it will offer a watermark detection API, though it has not yet provided implementation details. The tool is expected to help users and third parties test whether a piece of text is likely to have been produced or processed by Claude.

Anthropic drew a distinction between watermarking and AI detection software sold by third parties. External detectors do not have access to Anthropic's key and instead look for patterns in phrasing that may suggest AI use.

For images and other supported file types, Claude will attach content credentials through C2PA metadata rather than use the same watermarking method as text. Those credentials indicate that a file was made or processed with Claude and can be read by tools that support the standard.

The policy will also apply to older Claude models launched before the EU deadline after a transition period, with rollout expected over the coming months. As Anthropic put it: "A watermark can only determine that Claude was likely involved with the content at some point. It cannot distinguish 'Claude wrote this' from 'Claude heavily edited this.'"