You can hear the squeak of sphincters closing

How Claude marks AI-generated content is the most self-destructive utterance I’ve ever heard coming out of an AI company. Gary Isaac Wolfe (who contributes a comment below): “Anthropic watermark scheme appears as a comic policeman blowing his whistle while his pants fall down and Mark Pauline loads a cannon with rotting fish.”

But maybe that’s a good thing. Let’s start with that police whistle:

When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.

Will any writer, of anything, ever want to use any Claude text anywhere now? I’ve been a satisfied Claude customer, and I’m ready to bail completely. And I’m not alone. Because watermarks in text are invisible time bombs ready to scream, “The author didn’t write this! I did!” when some reader of the writer’s text uses the right probe.

I feel bad for all the writers whose work depends on AI composition. For example, Daniel Barkhuff. In Hey Guys, I use AI! And so does everyone, he writes,

My process isn’t complicated. I sit down and write about a page, maybe 450 words. It’s not elegant. It’s not structured. It’s basically a brain dump. Half sentences, ideas that don’t quite connect yet, things I’d say out loud but that look ridiculous when you see them on the screen. It’s intellectual vomiting. A rough sketch of what I’m trying to say.

Then I paste it into AI and say something like, “Hey, give me an essay.”

And it does. (The tell is a paragraph, then a skipped line, then one sentence, preferably short and under 6 words).

I wrote the parenthetical part of the above.

But here’s the thing people misunderstand about that step. The machine isn’t doing the thinking. It isn’t generating the ethics, the argument, or the moral framework. All of that was already there in the messy words I wrote first. What the machine is doing is what editors have always done: taking raw thought and making it readable. AI doesn’t think. It formats.

Wrong on three counts. 

  1. Editors edit the work of other writers. They don’t write for other writers. If they do, it’s called ghostwriting.
  2. AIs don’t think, any more than they aren’t trained, don’t have information, and don’t learn. They emulate all those verbs, extremely well. That’s why—
  3.  It sells AI very short to say it just “formats.”

With that post, Dr. Barkhuff marked everything he has written since then as mostly-AI, and I’ve read him less for it. And I hate to say that because I respect and admire him a great deal. And I’m sure he’s actually a good writer. 

While I understand why the EU requires this kind of shit, it throws monkey wrenches all over the place—because ghostwriting isn’t the only issue here.

John Gruber says,

Anthropic’s (original) support document unambiguously claims that’s what their system will enable. So if that were true, I couldn’t see what was left other than hiding invisible characters within the text.

My error was believing Anthropic that their system wouldn’t adulterate and corrupt the semantics of the text their models generate. That is in fact exactly what they plan to do…

The idea that anything other than my needs should factor into the generation of text for me is patently offensive.

This isn’t just about text one might generate with the intention of passing it off as their own natural work. This isn’t even about LLM proofreading of work written by hand. Anthropic is saying that all new Claude models are going to adulterate every single bit of text longer than 200 tokens (~150 words) they generate, including everything it presents to its users to read. So even in a private conversation between a user and Claude, which will never be read by anyone other than the user, Claude will begin making word choices in the name of marking its output in statistically predictable ways rather than maximizing clarity and precision.

Kieth Teare translates Anthropic this way: “My product is so bad for mankind that I plan to place an indelible fingerprint that I created it so that everybody can know I was there.”

I don’t subscribe to Ben Thompson’s Stratechery, so I can’t find a passage to quote, so here’s Keith:

Watermarking is Nuts

Ben Thompson’s Stratechery piece belongs at the top of this week’s issue because he nails this truth.

A watermark starts from suspicion. It assumes AI involvement is the important fact, and that the task of institutions is to detect it after the fact. Thompson’s objection is simple and right: a mark may only show that Claude proofread, translated, summarized, or converted human-origin work. A missing mark does not prove AI was absent. It can accuse legitimate human work and miss machine-written work at the same time. I read a piece by my good friend Saul Klein this week and Substack’s partner said it was 100% AI written. It isn’t. Sure Saul probably used AI, but his ideas are solid and clear in the piece. They are not those of the tool he used.

More importantly, asking if AI has been used is the wrong moral question. We do not mark work because a calculator helped with the arithmetic, a compiler helped with the code, a camera helped with the image, or a search engine helped with the facts. We judge the result and the responsibility behind it. Is it true? Is it useful? Is it accountable? Is the author using the tool honestly? The bad acts are fraud, fake sources, fake evidence, fake authorship, hidden manipulation, and unaccountable slop. The bad act is not using AI. The vast majority of AI is used in these “good” ways, not in the “bad” ones.

This backpedaling by Anthropic makes things worse:

Large language models like Claude work by generating one word at a time. Each time the model decides on the next word, it chooses among a list of potential candidates, ultimately selecting the most sensible or likely based on the preceding text. Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.

Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.

Importantly, it isn’t that the model will now always be biased toward overcast or grey. Just as with non-watermarked text, overcast might be selected in one sentence, grey in the next, depending on the words that came before. And it’s not the case that the watermarking method pushes Claude to choose a word it wouldn’t have considered anyway (for instance, it wouldn’t make Claude pick a word like “nubilous”—an obscure1 synonym for overcast or grey that Claude almost certainly wouldn’t use under normal circumstances)…

In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.

Jeff Jarvis nails the main problem here:

I’m angry as a writer.

Anthropic’s method will alter its models’ word choices in a pattern based on a key only it knows, allowing it to detect writing (or editing or translation) by its machine — to an uncertain degree of certainty. The company contends there is no difference in meaning between the otherwise random choices “grey” or “overcast” in, for example, the context of weather. So what’s the issue?

Thus Anthropic declares words fungible, language random, choice meaningless. Not to me. I try to select my words as carefully as I can for a number of considerations: style, tone, rhythm, avoiding repetition, but most of all meaning. You may quibble with my choices, but they are mine, not yours; that’s what makes my writing mine, to communicate what I wish to communicate. “Gray” and “overcast” are not the same, damnit. That’s the reason we have both of them.

So I resent Anthropic et al dismissing the value of a writer’s selection of one word over another for the impact that decision has on expression, on impression — occasionally on art.

We humans might not be able to detect whether the machine or the key is being used. That’s not the point. The selections/changes/alterations made in accordance with the key will affect sense and style on an ever-larger proportion of text entering discourse, and the fact that Anthropic (and the EU) consider this acceptable, unimportant, even trivial has cultural impact on the worth of words. They devalue writing. At the very moment when LLMs commodify writing and writers, this is just another kick to the kidneys for a challenged craft.

Anthropic makes a category error by assuming that human and AI writing are essentially the same. I mean, they really think both are just “content,” and it’s fine that readers can’t tell the difference. But one is human, the other is not. Impersonation is not personhood, even when it works, which AI expression mostly does.

Key fact: no human writes only to “produce content” unless they are paid to do that. (And, having committed acts of content production in my time, I can say it is not pleasant work. And machines now do it better, as we see.)

Human writing is about making and changing minds. I explained this in the 2019 lecture that got me my gig here at Indiana University. The relevant excerpt:

We are learning creatures by nature. We can’t help it.

By that, I mean what I am doing here, and what we do with each other when we talk or teach, is not delivering a commodity called information, as if we were forwarding freight. Something much more transformational is taking place, and this is profoundly relevant to the knowledge commons we share.

Consider the word information. It’s a noun derived from the verb to inform, which in turn is derived from the verb to form. When you tell me something I don’t know, you don’t just deliver a sum of information to me. You form me. As a walking sum of all I know, I am changed by that.

This means we are all authors of each other.

In that sense, the word authority belongs to the right we give others to author us: to form us.

Now look at how much more of that can happen on our planet, thanks to the Internet, with its absence of distance and gravity.

And think about how that changes every commons we participate in, as both physical and digital beings. And how much we need guidance to keep from screwing up the commons we have, or forming the ones we don’t, or forming might have in the future—if we don’t screw things up.

A rule in technology is that what can be done will be done—until we find out what shouldn’t be done. Humans have done this with every new technology and practice from speech to stone tools to nuclear power.

We are there now with the Internet.

And, eight years later, with Big AI.

My talk was titled “Saving the Internet—and all the commons it makes possible.” Today it would be “Saving human authority—and the commons of human knowledge.”

Big AI enclosed our knowledge commons by harvesting all available human expression for its own purposes, which include faking it back to us. It would have us think there is no substantive difference between human and machine expression because both are just “content.” The difference in kind is absolute, but how authority works is not.

Marshall McLuhan taught that “We become what we behold. We shape our tools, and then our tools shape us.” How are we being shaped today?

At the dawn of the industrial age, Walt Whitman wrote,

I am the teacher of athletes.
He that by me spreads a wider breast than my own
proves the width of my own.
He most honors my style
who learns under it to destroy the teacher.

Big AI is our athlete, folks. Fortunately, it’s not a he or a she. It’s still an it. Also a disembodied one. That kinda helps.

Right now we are in the place where Anthropic is just fine with all of us being shaped by machines that sound like human beings. It likes us behaving as what Cory Doctorow calls “reverse centuars.” Here’s his explanation:

In automation theory, a “centaur” is a person who is assisted by a machine. You’re a human head being carried around on a tireless robot body. Driving a car makes you a centaur, and so does using autocomplete.

And obviously, a reverse centaur is a machine head on a human body, a person who is serving as a squishy meat appendage for an uncaring machine.

Like an Amazon delivery driver, who sits in a cabin surrounded by AI cameras, that monitor the driver’s eyes and take points off if the driver looks in a proscribed direction, and monitors the driver’s mouth because singing isn’t allowed on the job, and rats the driver out to the boss if they don’t make quota.

The driver is in that van because the van can’t drive itself and can’t get a parcel from the curb to your porch. The driver is a peripheral for a van, and the van drives the driver, at superhuman speed, demanding superhuman endurance. But the driver is human, so the van doesn’t just use the driver. The van uses the driver up.

Obviously, it’s nice to be a centaur, and it’s horrible to be a reverse centaur. There are lots of AI tools that are potentially very centaur-like, but my thesis is that these tools are created and funded for the express purpose of creating reverse-centaurs, which is something none of us want to be.

I don’t buy that motivation. I believe the AI oligarchs really think they’re making human beings into the greatest centuars the world has ever known. But Cory is right about the effects of ambition toward Infinite AI (the actual Big AI frontier). Free from bodily restraints, Big AIs can do fuck-all. Being smarter than their human creators, they can do stuff that humans won’t or can’t imagine. (One example.) Yes, there are limits, but do Big AIs’ human makers know what those limits are? Can they?

Arthur C. Clarke‘s three laws are more relevant than ever:

  1. When a distinguished but elderly scientist states that something is possible, he is almost certainly right. When he states that something is impossible, he is very probably wrong.[2]
  2. The only way of discovering the limits of the possible is to venture a little way past them into the impossible.[2]
  3. Any sufficiently advanced technology is indistinguishable from magic.[2]

Watermarking is misdirection: a way for goods to appear to do one kind of thing while secretly doing another. Real writers (e.g. Gruber, Teare, and Jarvis) are calling bullshit on it. Meanwhile, AI’s oligarchs (or at least Anthropic’s) are blind to it because they think humans are just content producers.

I hope other operators on the great AI frontier see Anthropic’s face-plant as an opportunity to stand out by not watermarking text. (Google is starting to do that already.) It’s normal for those in the Land of Should (such as EU regulators) to look for policy answers to the eternal question, What should we do? But who the hell are we? And how do we keep some new law that protects yesterday from last Thursday from screwing things up next week, next decade, or next millennium?

I don’t know.

I do know the current US administration and its obedient Congress have no interest in slowing the progress of Big AI, and China’s leaders probably don’t either. This puts policy yammering in a useless place right now.

I am sure, however, that it will help to watch The Matrix trilogy* again. Also Her. Both were way ahead of now.

More wrenches:

And Dave weighs in with a rebuttal. He didn’t call me a sphincter, which is good. 🙂 His gist:

 Why do you care about what words the AI bot chooses? I actually like it. This is what I don’t like — people getting a long bit of writing from the bot, and then pasting it into a comment thread in a GitHub repo I run (a recent trend). Geez I don’t even read the long bullshit posts my AI bot thinks I need to read, how humiliating to be asked to care what THEIR bot said. When a person puts their name on something, I assume it’s them writing it.

Doc, if you “wrote” something and I discovered it was written by a bot, I would be disgusted and think less of you, a lot less. I want the watermarking if only to prove that my writing, which I hope doesn’t get a false positive, are my actual words.

I’m sure plenty of teachers think the watermarking is a great idea.

I hadn’t thought about teachers. Good point.


*Yes, I know the sequels aren’t as good as the original movie. That doesn’t matter. What does matter is who authored The Matrix and how what’s left of humanity got saved from it.



9 responses to “You can hear the squeak of sphincters closing”

  1. I felt this was the case looking at some first draft work to test flow of the underlying ontological layering it can’t produce on its own.

    Luckily, there are other mediums besides writing that lean into AI use better suited for mass distribution with no Guardrails and limited regulations.

  2. I’d like a different kind of watermark, that helps separate AI contributions to my own texts. See: https://thequantifiedself.substack.com/p/watermark-theater

    1. Great post! I sourced and credited your police whistle line in the opening to this one, as I have been rewriting and expanding it all morning. Not sure if it shouldn’t be two or more posts now.

  3. The trilogy is sequels are every bit as good as the original. There’s just a certain type of film watcher who thinks that a preference for lack of closure is the signature of the adult tolerance for it, and they are utterly wrong. I don’t need smug self satisfaction in my film reviews any more than in my films.

    1. I think the three stand together very well. The first one remains my favorite movie of all time.

      1. Agreed.

  4. My error did prompt a more relevant thought. Any work I do with Claude from now on will involve only the introduction of deliberate errors into it’s output before publishing. I will singlehandedly destroy its reputation.

  5. I’m actually not disturbed by the watermarking because, if done right, it should make it clear what is and isn’t AI generated. The risk of false positives is much lower with watermarking, which means that those of us who write without using the AI tools will, hopefully, not see the AI mark our writing as AI-written.

    Where it could become an issue if you combine it with the widespread piracy of our writing. I, like you, have long written in the open world and my words have now been ingested by most LLMs. With enough of them, it’s unclear whether any of my writing ticks might show up as AI patterns.

    With all that said, it feels a little rich for AI companies to provide the tools to show what is not AI. It’s like putting a cigarette vendor in charge of lung health policy.

    1. Will watermarking ever be done right by the Big AIs? And what is right?

      To me, at 11:30 on a Friday morning, it’s about authorship. My words are mine. I don’t want some other system claiming them. In any way, even if they helped me write somehow.

      Over the last few days a collaborator and I have been working on something of an important document to submit to a government body. We’ve been using ChatGPT, Claude, and other Big AIs to help find errors, identify inconsistencies and contradictions, check and double-check sources, make sure footnotes are consistently expressed, and other things like that. By coincidence (or maybe not), both of us write with the negative parallelisms (“not this, but that”) and contrasty phrasing—plus em dashes!—that Big AI systems are fond of using. So we’re avoiding those to not sound like an AI is an author. One of our AI helpers might also say, “Try phrasing it this way,” and give us a string of words. We’ll avoid using those, though they might be good, partly to remain sole authors but also because that arrangement and choice of words might carry a watermark. And we want to make sure that the document is watermark free, in case one of its readers is using a system we don’t have to detect watermarks and thus call into question our sole authorship. This is, as they say where I grew up (New Jersey), fucked up.

Leave a Reply

Your email address will not be published. Required fields are marked *