DL 444

Verified, Not True

Published: August 19, 2026 β€’ πŸ“§ Newsletter

Hi all, welcome back to Digitally Literate.

In this week's issue, we focus on generative AI watermarking. We look at recent developments, how it works, the challenges it faces, and the impact on education and writing in the coming years.

As always, your support is valued. Reach out anytime at hello@wiobyrne.com, and subscribe if you haven't already.


Compliance, Not Conscience

There was plenty of buzz this week as Anthropic, the maker of Claude, announced it is rolling out AI watermarking for text. Generative AI watermarking embeds invisible, machine-readable signals into synthetic text, images, audio, and video. As AI-generated work becomes more ubiquitous and indistinguishable from human work, there is a growing need to verify the origin and authenticity of content.

Anthropic didn't do this out of a growing need. This is less about ethics and more about compliance. California's SB 942 took effect on January 1, 2026, and enforcement of the EU AI Act's Article 50 transparency rules begins this August. These rules are meant to establish mandatory transparency measures to prevent deception and confusion caused by AI. Anthropic decided that since its models are global products, they should enforce the rules for all.

Claude is not alone in this work. Anthropic, along with roughly 190 other signatories, signed the EU's Code of Practice on Transparency of AI-Generated Content back in July. The industry was already headed this way. As we'll see below, the industry already has some tools in place for identifying provenance in images and audio, but text is a much harder challenge.

That compliance aspect is very important. Article 50 splits into two separate duties on two separate parties, providers and deployers. Providers (Anthropic, OpenAI, Google) must mark their own AI-generated outputs. Deployers (schools, newsrooms, and content creators using AI-generated materials) have an obligation to disclose that the content was AI-generated unless it underwent human editorial review with a named person responsible for it.

For example, in this newsletter, I use Generative AI to help source and organize information. I draft this out and use AI to help me understand what is really happening in these stories and how to explain it to the layperson. Under direction from Article 50, this would mean that the tools I use (Claude, ChatGPT, and local models) need to mark and identify their outputs. I should declare in my work that I am the final reviewer and responsible for all content presented. Lastly, I use a popular writing tool and an AI model to clean up my writing before I publish. I use these the same way I use a copyeditor for other publications I send out. I want to remove the friction that unclear writing creates between my ideas and the reader.

Invisible by Design

When you create and share digital content, it carries metadata (information about the information) and a C2PA, a standard that acts like a digital passport. The C2PA carries labels about the content, but they can be removed. On the other hand, generative AI watermarking is often an invisible mark embedded in the content as it is built. For the most part, this watermarking is kept proprietary, as it would require knowledge of methods or keys that could be left behind as models create content.

One of the most well-known versions is Google's SynthID, first described in a 2024 Nature paper. SynthID works by embedding an invisible pattern into pixels or audio frequencies when content is generated. Keep in mind, this is not a universal AI detector. SynthID works with a dual-network approach. A neural network embeds or injects a watermark into the content, and then a detector identifies and reads the watermarked signal.

Text is a much harder problem, as text watermarks are fragile, so meaning can survive content rewording. Text watermarks are also often vulnerable to paraphrasing, copy/paste modifications, and back-translation.

Anthropic's attempt at watermarking text is very interesting, and it's a version of that same SynthID-Text approach. To explain this, think about how AI writes. When Claude or ChatGPT writes something, it stops at each word and considers the probabilities of the next possible word. So if it writes "the cat sat on the," the model might pick "hat," "mat," "couch," "window sill," etc. The generative AI text watermark lives in that variability, as the model imperceptibly nudges word choices multiple times throughout the document. In a way, it's like the model is enforcing tone or style into the text. Taken together, all of these form a slight statistical pattern that you'd never identify.

It's important to note that this is also without problems. Generative-AI text watermarking also barely applies to code or factual, low-variation text. For example, this view writing as a spectrum between high-variation (creative or descriptive) or low-variation (factual or code) content. With high-variation text, you have some freedom in word choice. With low-variation text, no variability is allowed in the content. This means that if the model is talking about the capital of France, it needs to say Paris. If it's code to be completed, it needs to be shared correctly, or it will not run.

What the Mark Can't Prove

This is still a developing story, and the responses online have been interesting. My main focus has been on what this means for education, literacy, and, most of all, the future of writing.

Here's what I'm currently thinking:

The detection fantasy is the wrong frame for schools. One of the key things I hear bubbling up is discussion about how watermarking may exist, and now we use it to catch kids using AI. Text watermarking can fail under paraphrasing and falsely flag human writing. This makes it unreliable for decisions about academic integrity. It also becomes problematic and potentially inequitable, given earlier detectors' problems with non-native English writing.

As an example, OpenAI's own AI-text classifier, released in 2023, caught just 26% of AI-written text with a 9% false-positive rate before it was quietly withdrawn. A Stanford study found detectors falsely flagged human writing at a 61% rate on TOEFL essays written by non-native English speakers. NeurIPS desk-rejected 178 papers in 2026 based on detector output that was later questioned as uncalibrated.

I think we're poised to enter a world where the detection of watermarks carries more weight than the system can actually, authentically bear. A detected watermark doesn't prove that Generative AI wrote something. No detected watermark doesn't prove a human did. Detection should be a signal, not a verdict.

Provenance is a new literacy. Provenance means the chronology or the ownership, custody, or location of a historical object. In plain English, provenance is like a diary or travel log for an old object. It's basically a way to prove that I this item (a painting or signed object) is real, wasn't stolen, and how did it get to me. People need to understand what provenance systems can and cannot prove.

In creative efforts like writing, we generally expect to see a straight line from an idea to evidence to the ultimate output. But this is not entirely true. Writing is not just a mode of output or expression. It's a mechanism for making sense of things. You figure out things as you write.

In addition, we teach students how to research, obtain evidence, and cite their sources in their work. To properly teach provenance, we need to understand and track the evolution of thought. That means we need to better understand how an idea germinated, how it advanced, how it was modified, co-opted, and published by others. Many of our systems are not built to, and we often don't spend the time digging in to track editing histories and iterations over time.

In a sense, we move beyond asking who said something first to asking whose thinking actually shaped the claim. Instead of valuing only the finished product or superficial mechanics, our focus shifts to assessing the authentic inquiry behind the work.

The Understory

In 1300, Edward I signed a statute that gold and silver could not be sold in England unless it had been assayed and stamped with an official mark at Goldsmiths' Hall in London. The mark was a leopard's head. It is why we still call this kind of official stamp a "hallmark," seven centuries later.

The crown didn't do this to protect buyers out of principle. Debasement was rampant as silversmiths were cutting precious metal with cheaper alloys and passing it off at full value. This ultimately impacted the currency that the crown itself depended on. The hallmark was an instrument of compliance, aimed at a specific, measurable harm,

The hallmark was meant to prove that the metal in an object was what it claimed to be, at the moment it was assayed. It said nothing about how the silversmith acquired the silver. It said nothing about whether the finished piece was honestly represented to the buyer, or who actually owned it, or what happened to it after it left the hall. A perfectly hallmarked candlestick could still be stolen goods sold under a false story. The mark verified the origin, not the whole story of how it got to you.

That gap didn't stop people from treating the mark as if it covered more ground than it did, and it didn't stop people from attacking it directly. Forging a hallmark became a capital felony by the Plate Offences Act of 1738. Despite this, the forgers kept trying because the mark had become valuable enough to fake.

In our current discussion, a watermark tells you a passage is statistically consistent with an output from generative AI. It doesn't tell you what inspired that work, whether the creator actually understood what they submitted, fabricated claims, or was honest or sincere in their work. In an attempt to identify what is true, we're once again looking for a mark to serve as a substitute for thinking about thinking and learning. We're having the same arguments, but this time with tokens instead of silver.

See you next Wednesday. As always, my email is hello@wiobyrne.com.

Digitally Literate is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.


Previous: DL 443 β€’ Archive: πŸ“§ Newsletter


πŸ•ΈοΈ Connected Concepts: