“It’s not cheating, it’s quiet efficiency”: What is AI watermarking, and how should we respond to it?

By Julian Koplin, Monash University

Image: Nahrizul Kadri - Unsplash

On 2 August 2026, the European Union’s new AI transparency rules came into force. Under Article 50 of the EU AI Act, text generated by AI will need to be marked so that it can be more easily detected.

Anthropic is the first major AI lab to implement this requirement. Claude, its flagship chatbot, has begun embedding invisible ‘watermarks’ in the text it writes. Many other AI companies have agreed to do the same, so Anthropic probably won’t be alone for long.

This move towards transparency is a good thing. But it also carries a risk.

We need a shared social understanding of which uses of AI are a moral problem, and which are not. Until we have this, there is a danger that we will treat AI watermarks as a sign that the human ‘author’ has done something wrong – and dissuade people from using AI for those purposes for which it is genuinely useful.

How does AI watermarking work?

According to Anthropic, the company embeds ‘watermarks’ in AI-generated text by subtly steering the word choices made by its AI models.

Large language models work by generating text one word at a time. Sometimes, there are several candidate words that the model thinks would work equally well. For example, it might be unclear whether a wombat is best described as ‘stocky’ or ‘sturdy’. Previously, this tie would be settled using the digital equivalent of a coin-flip. Now, however, Anthropic uses a pseudo-random algorithm to select one word rather than the other. 

To readers, the outputs appear normal. However, those with access to the algorithm used to guide these choices can check whether a passage of text matches the word choices that Claude would make.

No single word choice is definitive proof of AI use; even if describing the wombat as ‘stocky’ is exactly what Claude would have done, it is entirely possible that a human writer simply happened to choose the same word. However, a long passage of text might contain hundreds of these apparent coin-flips. If all of them come up heads, the most likely explanation is that the coin was loaded.

The reaction

Following the news, many Claude users cancelled their subscriptions and took to social media to complain. Some are worried that if a watermark is detected in their own writing, others will wrongly assume that the AI has done all the work – rather than, say, making a minor contribution to their writing process. As one Reddit comment put it: “I gave the instructions, context, decisions, and countless refinements, Claude was [just] the tool.”

Not everybody is sympathetic to these complaints. As another commenter put it: “The only reason you wouldn’t want [watermarks] is to lie to people.”
This is an understandable reaction. There is something that looks self-serving about wanting to enjoy the benefits of AI while concealing that you have done so – especially when the internet is awash with AI-generated slop. If AI watermarks make it harder to get away with producing low-effort online content, this is a good thing.

Yet the frustrated Claude users get something right. The presence of an AI watermark cannot tell us how the human ‘author’ has used AI, nor whether their use of it ought to be condemned.

What does an AI watermark mean?

As Anthropic’s documentation states, AI watermarks tell us just one thing: whether the content was “processed by Claude”. They do not tell us how much influence Claude had over the ideas in the writing. This is important to understand, since people use AI in many different ways, and they are not all equally deserving of condemnation.

Consider, on the one hand, typical examples of ‘AI slop’: vacuous AI-generated LinkedIn posts listing productivity hacks that are already well-known (“Try the pomodoro technique!”), and university essays produced by ChatGPT (and submitted without changes) after students upload their assignment instructions to the system. These AI outputs are distinctly low effort; their human authors deserve little credit for them.

On the other hand, imagine a scientific paper produced by a researcher who is not a native English speaker, and who has asked Claude to help render their hard-won ideas into fluent English. Or imagine a philosopher who has turned to Claude for copy-editing advice and decided its suggestions would make their argument clearer or more accessible. I write philosophy myself; I can confirm we are bad at this.

In this second set of cases, the human author has done a large share of the work. They also deliberately set out to communicate something that they, the human author, thought mattered.

These are not typical cases of ‘slop’. Arguably, they represent some of the best applications of generative AI. However, because Claude contributed to the phrasing of these texts, they will carry the same watermarks as a list of generic productivity hacks wholly produced by AI. 

What does the absence of an AI watermark mean?

The presence of a watermark cannot tell us whether the ‘author’ has behaved badly. Equally, the absence of a watermark does not mean the ‘author’ produced the work themselves.

As Anthropic acknowledges, editing Claude’s outputs can replace the distinctive word choices on which its AI detection efforts depend.

This means two things. First: a human ‘author’ can have Claude generate a piece of writing for them, then make some simple word substitutions to wash away the watermarks. They might contribute nothing to the text but a few synonyms, yet pass the watermark test with flying colours.

Second: a committed cheat can spare even the slight effort associated with thinking of synonyms by using a second AI to do this for them. AI ‘humanisation’ apps – already a cottage industry – tweak AI outputs so that they will not be flagged by current AI detection tools. No doubt these apps will soon be able to launder watermarked text.

Alternatively, a dedicated AI cheat could turn to a service that does not watermark its text. Such services will presumably continue to be offered by companies operating outside the EU’s jurisdiction, as well as via open-weight models that individuals can run on their own hardware.

What next?

AI watermarking will bring greater transparency around AI use. Whether this transparency is an unmitigated good depends on how thoughtfully we respond to it.

At a time of growing AI backlash, I worry that watermarks will be treated as a clear sign of wrongdoing, even when the human author had good reason to use AI. I also worry that if watermark detection comes to be seen as the authoritative test for AI use, technically savvy AI users will be even better able to escape suspicion than they are now.

Many people want nothing to do with AI. This is a respectable position; some people have legitimate objections to how large language models were trained, and others are understandably concerned that relying on AI might erode their ability to think for themselves. And yet, many people do wish to use AI, and it is important to recognise that not all the things they might use it for are pernicious.

We need to build norms that distinguish these uses from the low-effort ones, and to make sure that watermarks are not seen as incontrovertible evidence of wrongdoing. One starting point might be to look at whose ideas, perspectives, and values are driving the writing. When AI helps people express something that matters to them, it has achieved something worthwhile; when it is used to replace human thinking, it has definitely not.

We also need to avoid outsourcing our judgements about human authorship to AI detection tools. Already, many of us are developing a kind of ‘sixth sense’ for the prose produced by popular AI models; a well-tuned ear will hear a phrase like “It’s not cheating, it’s quiet efficiency” and sound the alarm. Reading for traces of human thought (or their absence) will matter even more in an era when watermarks offer a test for AI use that looks objective but can be gamed.



A watermark can only ever tell us that a tool was involved. Precisely when we should welcome the use of this tool remains a question for us humans to decide.

Fact-Checked by Irfan Ahmad.

Read next: AI agents can now remember and hackers can ‘poison’ their memories — a new cybersecurity threat
Previous Post Next Post