The Story About the Story

Published: September 16, 2026 • 📧 Newsletter

Hi all, welcome back to Digitally Literate.

Jacob Coxon's resignation from Anthropic was at the top of my list to write about this week. Then, serendipitously, a local news station reached out asking me to talk about the exact same story on camera.

The pitch was “I’m focusing my story on what people should be aware of with all the reports that are saying AI could kill all humans by 2030. Like what should people know, should people be worried, etc.”

Prepping to say something coherent to a general audience in a five-minute segment is a pretty effective way to find out what in a story you actually understand versus what you’ve just been absorbing secondhand. This issue is the sorting-out process, written down.

As always, your support is valued. Reach out anytime at hello@wiobyrne.com, and subscribe if you haven’t already.

The warning, before anyone spun it

On September 8, Jacob Coxon resigned from Anthropic. He’d spent three years doing pre-training research at both OpenAI and Anthropic before that. At 8:04 PM that evening, he posted a seven-part thread on X. It has since crossed 172 million views. I’ve presented the thread below with some comments.

“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.”

It’s worth highlighting the term self-improving superintelligence in this conversation. In plain English, this means that instead of each new AI model being built by people from scratch, labs are already using one generation of AI to help train, fine-tune, and evaluate the next. This results in the new systems doing more and more of the training for new AI models, not humans.

“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”

Coxon highlights three specific claims: offensive hacking capability, research-grade acceleration, and autonomous resource acquisition. Given Coxon’s insider status, we need to consider his warnings rather than general anxiety about AI.

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger.”

Coxon is doing two things in this tweet. First, he’s leaning on his insider status to share what conversations are like in the companies where these models are built. He differentiates between what insiders say to the press and public, and what they say privately to each other. He also highlights the “this is not a marketing stunt” point to anticipate future arguments.

One extraordinary claim buried in there, too: “No other human activity poses this level of danger.” In a normal world, we’d have public dialogue about things that could kill us all. My initial takeaway is that we need the dialogue and more than one person’s private conversations to support it.

“A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk.”

In this tweet, I see Coxon addressing the objection every reader will have: why are you building something you know will kill everyone? He presents two different diagnoses, one per company. Again, this is Coxon’s read on the two companies’ internal cultures.

“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”

This is the line that matters most to me. A decision at this scale is being made inside private chat conversations at a handful of companies, by a small group of people, with no public input and no accountability. We need public dialogue about this, not some Slack channel texts about the end of the world.

“I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”

Readers of this newsletter know that we covered the Hugging Face incident in depth in recent issues. With this tweet thread, Coxon is adding insight into how people at the companies view these events. His use of the phrase "warning shot" describes an incident serious enough to prove the danger is real, but not catastrophic enough to be a disaster. Highlighting U.S.-based labs is interesting, as labs agreeing to pace themselves doesn’t touch the international competitive pressure that is supposedly driving the race in the first place. He ends with a concrete ask (a temporary ban on improving model capabilities), not just “be careful.”

“If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ — or take this moment to call for different conditions?”

In the closing tweet, he pivots from the general public to a specific audience. He stops addressing the over 172 million people who saw this thread and turns to the ending for the people working at the frontier labs.

He describes a superintelligent RL run without a rigorous understanding of its mind. The RL run is a technical term; it refers to the reinforcement-learning training process that shapes a model’s behavior. His point is that labs are running this process on systems whose internal reasoning still isn’t well understood by the people building them. Human interpretability hasn’t caught up to superintelligent capability.

He closes with addressing the “Put your head down because ‘it’s happening anyway.” He’s pointing out that whether we work at AI labs or use their products, we have a choice: either comply or push back.

The agreement, four days later

On September 12, four days after Coxon’s post, Dario Amodei published an essay arguing that “We must slow the pace at which we improve the capabilities of AI models.” Not to halt progress, but stop capabilities from outrunning the safety work meant to keep them in check. He feared what a more capable “swarm” could do, and what recursive self-improvement reason things are moving faster than expected.

The plan has three parts: embedded third-party evaluators with employee-level access to verify safety claims, industry-wide coordination on safety standards and pacing, and international agreements (even with authoritarian governments) on the most dangerous uses. Only the first part is a binding commitment (and only from Anthropic) “regardless of others’ actions.”

Within hours, Sam Altman called the evaluator's idea “a good idea” and said OpenAI would match it. Elon Musk, who is currently suing OpenAI and rarely agrees publicly with anyone in this fight, posted: “Dario is right.”

Things got much more strange as within a day, Parker Thayer, an investigative researcher, published a thread arguing this was “the start of a very sophisticated and well-funded PR operation”.

Now, we’re not trying to figure out whether racing toward self-improving AI without a plan to control it is, on its own terms, a good idea. Instead, we’re automatically into discussions about whether this is marketing, PR, or a psyop?!?!

The Understory

On September 26, 1983, a Soviet early-warning satellite system reported that the United States had launched one intercontinental ballistic missile at the Soviet Union, with four more behind it. Stanislav Petrov, the duty officer at the secret command center outside Moscow, had minutes to decide whether to pass the warning up the chain of command, a chain that led toward retaliation.

He decided the system was lying. A real American first strike, he reasoned, would come as hundreds of missiles, not five. He waited for corroborating evidence that never arrived, and reported it as a false alarm. He was right, as sunlight reflecting off high-altitude clouds had fooled the satellite into seeing missiles that weren’t there. Petrov wasn’t rewarded for it. He was reprimanded for failing to properly log the incident in the war diary, quietly reassigned, and the story stayed classified until 1998.

Mutually assured destruction gets described as a system. Rational actors, deterred by certain retaliation, collectively choose not to press the button. But the actual history of nuclear near-misses is full of moments like this one, where the system held because one person, alone, under pressure, decided not to trust it.

That’s the troubling piece about this conversation about racing toward self-improving AI. It’s that the restraint we point to when we say “we’ve handled this kind of danger before” ran on someone’s willingness to doubt the machine in front of them. It’s not obvious who that’s supposed to be once the machine is the one doing more and more of the evaluating.

See you next Wednesday. As always, my email is hello@wiobyrne.com.


Previous: DL 447Archive: 📧 Newsletter


🕸️ Connected Concepts: