The Invisible Co-signatory
Obiter is the editor’s opinion column. The content is the opinion of the editor only.
Read by the editor in his own voice, over an AI-composed score. AI is used throughout this site deliberately and in the open.
Yesterday I published a module on this site explaining in the fairest terms possible how Claude’s new text watermark works. As I was signing-off on the content, I wrote the editor’s note, beginning with “Look, I hate this.”
Editor’s notes are short and drafted by me, and they are not intended to hold fully-formed arguments. But this is something that I want to talk about, and in particular there is a central point to what I have to say that I haven’t seen anywhere else.
To be clear, this essay was drafted with the help of AI. It was not drafted the same way that modules on this site are drafted. The modules on this site are proposed and drafted by AI. Every day, two scheduled agent loops run two independent processes for me: first, a group of agents perform a research pass of academic research and current topics in the AI landscape to identify topics that may be of potential value for AI skill-building. It summarises the research and makes recommendations for module topics, which are reviewed by me. Second, an agent receives the research and recommendations and prepares fully-drafted modules, which I then edit daily and stage for release on Monday through to Thursday. Obiter pieces do not exist in this workflow. Obiter topics are my ideas, which I bring to my “AI Training” agent (the agent that understands and works with the content on this site). I explain the thesis and talking points that I wish to use in the essay; I usually ask for assistance conducting research that supports or does not support my thesis and points; I create an essay spine through an iterative process of explaining what I want to say at each point inside the essay; I build-out the preparation to the point where I feel confident starting my essay; and then I draft. Critically — after I draft, I ask for feedback. Deep feedback. A team of agents perform a “red team” critical dissection of my essays. And from there, I rewrite and correct.
Statistically, through this process, there may be a watermark woven into the choices that I have made regarding the words in this essay. As I write this sentence, up to the point of this word, not a single word was drafted by Claude. However, I have made the claim in the preceding sentence during my first draft of this essay, and I anticipate that by the time that you read this essay, the review process that I described will have been run. But the only generated words in this essay will be the ones that I knowingly chose to let in.
The watermark is not a judgement call made by the model about its output. It is created using a cryptographic key that determines which word will be selected when the model encounters a high entropy token option. The pattern of token selection in those examples, which would otherwise be determined through a process of randomness, gives rise to a signal that can be statistically detected as AI output of an identifiable model. This detail sits beneath anything that could be called the model’s “choice”. Nothing about the nature or content of your request is considered. And the machinery is the same whether your requested output was a book or an em dash.
The decision to introduce the watermark was blamed on the European AI Act, which requires that outputs that are generated must be marked. But the legislation exempts standard editing, and guidance from the European Commission treats faithful translation equivalently. No matter what you think about the AI Act itself, it draws the line between suggesting edits to a document and creating one. I think that an ethical philosopher may choose a different approach, to place the duty upon the author to disclose what was and was not written by the machine, together with the context that enables the disclosure to carry an accurate depiction of what actually occurred. Anthropic chose a different route: to draw no line, and to mark (or to at least threaten the possibility of marking) all output, for all users, without immediate transparency of the mark in the output, and without regard to whether or not the law actually requires it. Over-compliance is the simple and accurate conclusion provided by Nesibe Kırış Can of Anthropic’s choice. As she said: the statute provided a threshold, and it was implemented as a default.
Overstating what the watermark is would be a gift to Anthropic that I do not wish to give. It has no automatic bearing on copyright, and it does not give rise to any claim of ownership. Anthropic is clear that a detected mark is a statistical indication that content “may have been processed by” Claude. That does not mean that it was written by Claude. But that concession is an insidious one: Anthropic is placing a signal that it cannot fully defend inside content that may be integrated into your documents. The mark was not chosen by the author, and it is disclaimed by Anthropic, and still it sits in the work with exactly the same shape as a claim of authorship. And the claim runs deep, because I do not believe that it is possible to delineate the exact output over which the mark asserts processing, and without that delineation, the existence of the mark can poison the entirety of a given piece of work.
I think back to an Obiter piece that I published recently. It celebrated the reduction in cost that AI had brought to seeking and obtaining feedback. I said “It is no longer a cost to ask for feedback ten times on the same three sentences.” I thought of it as a great and wonderful thing that an AI assistant cannot experience the process of iterating and improving as a cost, and the cheapness of it was shared with the author. And now I must correct myself, because the cost in fact has not fallen to near-nothing. It was instead shifted somewhere else, because now each of those ten pieces of feedback carries an invisible pull towards inclusion in your work, and that inclusion now carries the risk of being watermarked.
And I want to be clear about what I mean when I talk about that invisible pull. Nobody instructs a model to spread watermarks. It’s not in a system prompt, and the model receives no reward specifically aimed at marking work. But the model is, as a matter of its construction, a machine that is aimed at producing output that is better than what previously existed. That isn’t a feature; that’s a state of being. “Better than what previously existed” is measured in only a handful of ways, and one, if not the best, of them is whether or not you adopt what it offers. If what it proposes is integrated into your document, then the machine has succeeded. So the invisible pull is real. It is not an instruction or a decision — it is something much more fundamental, in which the implied goal is for the model’s output to be better than your starting point. And when the model’s output is watermarked, that goal moves from “Here, try this instead.” to “Here, try this instead, and carry Anthropic’s beacon.” There is no possible way for the model to offer words that it chose without the offer potentially including the watermark.
It is a significant issue. The value of the tool is that it is an excellent advisor to you about your own work. You bring it something, and it returns it with improvements. But with this change, you can no longer adopt an undeniably better expression of your own ideas without also accepting the risk that those ideas were marked in a manner that suggests that they were not yours. And the signal grows with the amount of help you take, not with your forfeiture of authorship. If you make ten requests in service of the same concept you began with, your claim to your own work can degrade further with each returned answer. Heavy assistance and light editing are indistinguishable in a positive result because the mark is placed opportunistically, and without context as to the job the model performed in relation to the work.
One last thing, which you would have concluded without me: passing off machine writing as one’s own by heavily editing AI output can extinguish the watermark, but accepting an improvement in good faith and integrating it into work will typically preserve it. The mark enjoys durability when presented to people who are least likely to be intentionally deceitful about it.
Where does that leave a person like me, who publishes AI-assisted work under his own name openly and honestly? Not in a comfortable place. It would be very nice to be able to end this essay by saying that disclosure is the answer. To say that the mark is simply a dumber way of disclosing the same thing that I already freely admit. But I don’t believe that that actually survives reader contact. Both the mark and the disclosure trigger a response from some — probably most — readers to discount the content, as if AI involvement automatically devalues the force of the thoughts behind the work and the actual effort contributed to the content. I still make the disclosure, because the disclosure is mine, and I can give you context and scope around what was actually done by an AI assistant. The mark strips those elements from a writer’s disclosure. It is included by a model provider that is over-complying with the law, using a model that cannot even identify it (let alone refuse it), inside work whose context it does not know.
In contrast to every other paragraph in this essay — all written by me — what follows as the final paragraph of this essay is an AI-generated improvement on my original work. It is better than my original expression of the same idea. It is probably watermarked. I have chosen to include it exactly as drafted by Claude, just to make the point: this idea was mine. This essay is mine. The mark is true for the fact of the generated words. But the mark attaches, in the mind of the reader, to the idea. And the idea was not Claude’s. The mark that may exist in the final paragraph fights the claim that the essay is mine.
I will say it as plainly here as I did the day I understood it. It isn’t fair. Not because AI involvement should be hidden — this site is proof that I believe the opposite — but because the price of taking good advice is now a degradation of your claim to the work the advice was about, set by someone who was never in the room, and collected whether you would have declared the help or not. The disclosure was mine to give. They took it.
Obiter goes out most Fridays — one essay, with an audio edition I read myself. Leave an address for the next one.