The House Doesn't Publish Its Tells
2026-08-21 — customer acquisition strategy India
I watched an AI training data broker pitch corporate procurement teams last month. The pitch: physical books. Paper. Binding. Things that existed before the slop machines.
The bet underneath was clean. If you want to train on real data, buy it from the dead tree era. Don't ingest what the models already spit out.
Nobody's doing this by accident.
Here's what I'm watching instead: the content vendors who sell you machine-generated articles are the same people saying detection is broken. The companies building detection tools are selling to the AI labs that want to stay invisible. And somewhere in the middle, marketing teams are publishing AI text without marking it, betting that watermarking either doesn't exist or doesn't matter.
One of those bets is about to be wrong.
Watermarking Went Mainstream in August
Anthropic embedded watermarks into text generated by Claude models launched on or after August 2, 2026, worldwide. Not in Europe alone. The EU AI Act triggered it, but the watermark ships inside the model itself, so Brussels became everyone's default.
This is infrastructure, not theater.
The mark covers output across Claude, the API, Claude Code, and all cloud deployments. It's not metadata you can strip. The watermark changes the source of randomness used when the model generates each token. Edit the text lightly and it survives. Rewrite everything and it's gone—but at that point, most consultants would argue the content's yours anyway.
Not all, though.
The mark signals that Claude processed the text, not that it authored it. Claude proofreads your memo. Claude reformats your slide deck. Claude corrects your grammar. All three get marked. That matters for anyone who needs to know whether a document touched AI.
Academics will hate it.
Publishers will use it. Lawyers will cite it. Compliance teams will start asking for it.
Images and Audio Got There First
This isn't new for other media. Google's SynthID watermarked over 100 billion images and videos, plus 60,000 years of audio, by May 2026. That figure was 10 billion a year earlier. The acceleration tells you what adoption looks like at scale.
OpenAI, Kakao, ElevenLabs, and Nvidia all use SynthID now. Not because they love watermarking. Because maintaining separate model behavior by region is expensive and hard at scale. You ship it everywhere or nowhere.
The gap in text is closing fast.
Actually, that's not quite right. Text is different. Text under 200 tokens (roughly 150 words) is exempt from the watermarking requirement. Which means short posts, snippets, and punchy product descriptions live in a gap where nothing's marked. Which is fine until it isn't.
What This Means for You
Your AI content strategy is quietly betting that one of three things is true. Either watermarks don't work. Or nobody can detect them. Or they don't matter to your audience.
I'd test all three assumptions.
Anthropic says detection APIs are coming. Google's SynthID detector works on images, audio, video, and text. These aren't experimental. They're shipping now.
The second assumption—nobody will check—might hold longer. But journalists will. Publishers will. Fact-checkers will start running content through detectors the moment they have easy access to them.
The third assumption is the real one: that your audience doesn't care. Maybe they don't. Maybe they should, and maybe they will once they know the difference.
Most consultants get this wrong, and some of my colleagues at Tanjore are definitely in that group. They're treating watermarking like it's a compliance checkbox, something to route around. It's not. It's the moment the industry stops lying about where the stuff came from.
If you're publishing AI content without marking it, you're publishing on borrowed time.