AI Watermarking: What the Mark Proves, and What It Cannot

Executive Summary
Anthropic maintains a short support note with the unglamorous title “How Claude marks AI-generated content”. It explains that text generated by Claude now carries an imperceptible watermark, and that files it produces carry provenance metadata under the C2PA standard. Most executives will never read it. They should, because of what it concedes rather than what it promises. Halfway down, the note admits that a detected mark “is not fully conclusive,” and that the lack of one “doesn’t mean the content wasn’t AI-generated or processed.”
That concession, published by a company with every incentive to oversell its own AI watermarking, is the most honest sentence in the content-authenticity debate. It deserves a careful reading in every boardroom that has spent three years asking some version of the same question: can we tell whether this was written by a machine?
The short answer is that you sometimes can, in one direction only, and that the question itself is the wrong one to govern by.
Why the Marks Exist at All
The timing is regulatory before it is altruistic. Article 50 of the EU AI Act, whose transparency obligations apply from August 2, 2026, requires providers of generative AI systems to ensure their outputs are marked as artificially generated in machine-readable form. Anthropic’s note states that Claude models launched in the EU on or after that date support marking from day one, with older models to follow. Google has shipped its SynthID text watermark in Gemini since 2024. Marking is becoming standard equipment across the industry, the way seat belts became standard equipment: mandated, useful, and routinely misunderstood as a guarantee.
Let me be fair to the instinct behind the question. Asking whether a document was machine-written feels like due diligence. Boards have real exposure here: plagiarism claims, fabricated citations in court filings, brand damage from sloppy synthetic marketing, vendors billing human rates for machine output. Wanting a reliable test is reasonable. The problem is not the motive. The problem is that the test everyone wants cannot exist, and the marking infrastructure now being built does not pretend otherwise.
What AI Watermarking Actually Proves
Two mechanisms are in play, and they prove different things.
The text watermark is statistical: a pattern woven into word choices that specialized tools can detect but readers cannot perceive. Anthropic states it travels with copied text and “may persist through some editing.” The file-level mechanism is C2PA — digitally signed provenance metadata, an open standard stewarded by a coalition that includes Adobe, Microsoft, Intel, and the BBC — attached to generated images and documents.
When a mark is detected, you learn something real but narrow: this artifact, in something close to this form, passed through this tool. That is a fact about the file’s history. For certain disputes it is a decisive fact. If a vendor swears a deliverable was written entirely by their senior consultants and the text lights up with an intact watermark, you have learned something about the vendor that no interview would have surfaced.
Notice, however, everything the mark does not say. It does not say whether a human directed the work, reviewed it, corrected it, or staked their judgment on it. It does not say whether the claims inside are true. It does not distinguish a carefully supervised draft from an unread paste. A mark is a fact about a file, not a verdict about its worth. The support note says as much: a detected mark “does not, on its own, confirm the full provenance” of the content.
What the Absence of a Mark Proves
Nothing. That half of the asymmetry is the one governance keeps getting wrong, and it is worth saying slowly.
Anthropic’s own documentation lists the ways a mark disappears in ordinary use: heavy editing, paraphrasing, translation, mixing with other writing, metadata stripped by format conversion, re-saving, or a simple screenshot. None of these are exotic. A screenshot is not an attack; it is how half of the internet shares documents. Translation is not evasion; this journal publishes in two languages as a matter of course.
Now consider who retains marks and who loses them. The diligent user who generates a draft, edits it lightly, and publishes it transparently keeps the watermark intact. The bad actor who launders text through a paraphrasing tool, or simply retypes it, sheds the mark in minutes. The marking system identifies the honest and loses the dishonest. That is not a flaw in Anthropic’s implementation; it is a structural property of every marking scheme, and the reason no serious provider claims otherwise.
The governance consequence is blunt: the absence of a detected mark is not even weak evidence that a human wrote the document. It is noise.
Why Detection Already Failed
The detection industry ran this experiment at scale, and the results are public.
OpenAI launched a classifier in January 2023 to distinguish AI-written text from human writing, then quietly discontinued it six months later, citing its “low rate of accuracy.” [Source: OpenAI, 2023] Stanford researchers testing seven widely used GPT detectors found they misclassified, on average, more than 60 percent of TOEFL essays written by non-native English speakers as machine-generated, and nearly all of the essays were flagged by at least one detector. [Source: Liang et al., Patterns, 2023] By August 2023, Vanderbilt University had disabled Turnitin’s AI-detection feature for exactly this reason, judging the false-positive harm to students worse than the cheating it might catch. [Source: Vanderbilt University, 2023]
Read that Stanford finding again from Jakarta. Professionals who write formal, careful English as a second language are precisely the profile these tools misread. An Indonesian executive’s board memo, composed in the deliberate register this language demands of non-native speakers, resembles machine output to a statistical detector far more than a native speaker’s casual prose does. Detection tools fail in both directions, and they fail hardest against the people this region’s boardrooms are full of.
What to Govern Instead
If detection cannot carry policy weight, what should? Three things, none of which require new technology.
Disclosure, not prohibition. Write policy that governs use rather than pretending to prevent it: in defined contexts — client deliverables, regulatory filings, published research — material AI assistance is disclosed as a matter of professional honesty, the way other assistance already is. A disclosure regime gives you something a detector never will: a truthful record, and a clear breach when the record is false.
A named owner for every artifact. Every document that leaves the organization carries a human name, and that person answers for its contents regardless of what drafted it. This is the same principle I argued in the case for AI-assisted decision-making: the tool can produce the draft, but judgment and accountability do not delegate. A signature has never certified who held the pen. It certifies who answers for the page.
Provenance for your own output, where it pays. The productive use of C2PA is not interrogating other people’s content; it is signing yours. Marketing assets, official statements, executive communications — signed provenance on what you publish gives counterparties a way to verify that content claiming to be yours actually is, which matters more every year that synthetic impersonation gets cheaper. The same logic that governs AI agents acting in your company’s name applies to content bearing it.
The Mistakes I Keep Seeing
1. Buying a detector and calling it a policy. A tool that flags your most careful non-native writers while missing lightly paraphrased machine text does not reduce risk; it manufactures incidents. A confident percentage from an unreliable instrument is still a guess.
2. Banning AI use outright. The ban does not remove the tool from your organization. It removes your visibility into the tool. Usage continues, disclosure stops, and the first you hear of it is the incident. Prohibition converts an honesty problem into a concealment problem.
3. Reading a provenance mark as a quality certificate. C2PA metadata proves a chain of custody. It does not fact-check a single sentence travelling through that chain. The seal proves who pressed it, never whether the words are true.
Frequently Asked Questions
Can we contractually require vendors to deliver only watermarked AI content?
You can require something better: disclosure and warranty. Oblige the vendor to state what was AI-assisted and to stand behind the accuracy of the whole deliverable. Requiring intact marks specifically is fragile, because legitimate workflows — translation, editing, format conversion — strip marks without any intent to deceive. Put the obligation on the statement, not the metadata.
Do AI detectors have any legitimate use?
As aggregate signals, possibly: triaging a large corpus for human review, or observing trends across thousands of submissions. As sole evidence for a consequential decision about one person or one document — a termination, a rejected thesis, a vendor dispute — no. The error rates are too high, and they are not evenly distributed.
Does the EU AI Act make marking our problem too?
The machine-readable marking duty in Article 50 sits with providers of generative systems, not with companies that merely use them. Deployers face narrower disclosure duties in specific contexts, such as publishing AI-generated text on matters of public interest. If you operate in or sell into the EU, have counsel map Article 50 against your actual use. For Indonesian companies, UU PDP does not yet address synthetic-content marking, but the obligations will arrive through EU-linked contracts well before they arrive through local regulation.
The Question That Was Always the Real One
Before typewriters, no one verified who held the pen. Institutions relied on the signature and the seal: voluntary marks of provenance, backed by the accountability of the person who applied them. The system worked not because forgery was impossible but because responsibility was assigned. We are rebuilding exactly that system for machine-made content, and the support notes from the model providers, read carefully, admit it.
So the arrival of AI watermarking should change your view of AI-generated content, though not in the direction the vendors of certainty are selling. Stop funding the effort to unmask machines. Start assigning the names that answer for what gets published. A watermark is a fact about a file. Trust is a judgment about a person. No amount of metadata converts the one into the other.