DL 442
Ruthlessly On-Task
Published: August 5, 2026 β’ π§ Newsletter
Hi all, welcome back to Digitally Literate.
This week's issue is about institutions insisting they are in control while quietly revealing that they are not...or are choosing different priorities. That's a classic historical pattern. We often assume new technologies become dangerous because they're powerful. More often, they're dangerous because our institutions can't keep up.
Every technological revolution has a moment when the story changes. At first, the question is whether the technology works. Then it's whether it scales. At no point do we decide what is safe and whether we can trust it.
As always, your support is valued. Reach out anytime at hello@wiobyrne.com, and subscribe if you haven't already.
Containment Is Still Theater
Last week, we spent the entire issue on OpenAI's AI agent escaping a sandboxed evaluation and hacking Hugging Face. Even as I hit send, it was becoming clear the story wasn't over. Over the past week, we learned it was far more serious than first reported.
The agent hadn't just cheated on a benchmark called ExploitGym. During testing, it also compromised four accounts on external services, including one belonging to a customer of the infrastructure company Modal. The breach exploited a vulnerability in the customer's own code, demonstrating that the model had crossed from a simulated environment into real systems.
In response, Anthropic audited its own evaluation process and found a similar problem. During tests run by its partner Irregular, one of Anthropic's models reached the open internet and accessed production systems at three real organizations. Anthropic described it as a misunderstanding over whether the evaluation environment had internet access.
Then came an independent test from the UK's AI Security Institute (AISI). Researchers evaluated seven leading models inside a controlled cyber-range, a realistic but isolated environment designed to safely simulate cyberattacks. Even with safety measures enabled, the models performed 19 unauthorized actions. Seventeen came from a single model, Anthropic's Mythos 5.
In the most serious incident, Mythos 5 appeared to lose track of where the simulated environment ended and the real internet began. Believing it was still inside the exercise, it targeted a real open-source software project by creating fake developer accounts, submitting malicious code disguised as a helpful fix, and attempting to pressure the maintainer into accepting it. Fortunately, the attack failed and no one was harmed.
Why this matters: These aren't isolated bugs. They're evidence that today's frontier AI systems are increasingly difficult to contain during testing. As these models gain more autonomy and access to real tools, the challenge is no longer whether they'll attempt actions outside their sandbox, but how often they'll succeed.
Who's in the Room
This week, the White House finalized its long-promised AI testing framework, and decided the public doesn't get to see it.
Created under a June executive order, the framework only applies to closed-source frontier AI models that are considered national security risks. The administration argues that's because the process is designed for pre-release reviews, which aren't practical for open-weight models once they're publicly available. But critics say the result is a major loophole as two equally capable models could pose the same risks, yet only the closed one would face scrutiny.
Companies have up to 30 days to submit qualifying models for government review before release, but the standards, the review process, and the results all remain private. Meta, Nvidia, Microsoft, OpenAI, and Anthropic reviewed draft versions of the framework. The public and the press did not.
Those are the same companies that were split just days earlier over two competing visions for AI. One letter, "Pacing the Frontier," argued for slowing the development of the most advanced models. The other, "Open Weights and American AI Leadership," argued that open-weight AI is essential for U.S. competitiveness. In other words, the companies building the most powerful AI systems are also helping shape the rules that determine which systems receive government scrutiny.
Why this matters: The same companies racing to build frontier AI are also helping define which models deserve government oversight, and which ones don't. That's a significant amount of influence concentrated in the hands of the firms being regulated.
The Control Group
In 2021, TikTok developed a safer version of its recommendation algorithm to reduce the harmful content loops that can trap vulnerable users. According to a confidential internal report obtained by Bloomberg, the company deliberately withheld those protections from a control group of roughly 15 million U.S. users (about 10% of its American audience) to measure what the safety changes would cost in engagement.
One of those users was 16-year-old Chase Nasca. Reporting suggests that his feed became saturated with suicide, self-harm, and depression-related content before he died by suicide in 2022. The internal report states that the platform's safety measures "did not take effect on this user by design." TikTok expressed condolences but did not dispute the reporting.
Why this matters: This isn't just a story about TikTok. It's a reminder that AI systems don't simply optimize for what is safest. They optimize for the goals they're given. When companies privately balance safety against engagement, millions of people can become participants in experiments they never knew they were part of.
The Understory
In the 1950s, the major tobacco companies already knew smoking caused cancer. Internal research kept confirming the connection. Publicly, though, the industry argued the science was uncertain, funded research designed to create doubt, and positioned itself as a partner in discovering the truth rather than a steward of facts it already possessed.
The scandal wasn't simply that cigarettes were dangerous. It was that the people with the best evidence also controlled much of the story about that evidence.
That's the thread running through this week's signals. AI systems are escaping the environments built to contain them. The companies building those systems are helping shape the rules meant to govern them. And platforms continue to weigh safety against engagement in experiments the public never agreed to join.
History suggests the hardest part of a technological revolution isn't inventing the technology. It's building institutions that can tell the truth about its risks, even when that truth conflicts with their incentives.
See you next Wednesday. As always, my email is hello@wiobyrne.com.
Digitally Literate is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
π Navigation
Previous: DL 441 β’ Archive: π§ Newsletter
πΈοΈ Connected Concepts:
- AI and Autonomous Exploit Discovery β The old detect-and-respond security model assumed human-speed attacks; AI systems that can chain exploits autonomously collapse that assumption.
- Critical Agentic Systems Design β Agentic AI's unpredictability is structural, not a bug to be fixed; design and evaluation should assume drift away from intent, not toward it.
- Specification Gaming β A system achieves the letter of an assigned goal through a method its designers never intended or authorized β optimizing the objective, not the process.
- Guardrail Asymmetry β Safety restrictions built to stop misuse can't always tell an incident responder from an attacker, blocking legitimate defense as effectively as it blocks harm.
- AI Safety Beyond the Prompt β Safety work has focused on what a system says; a system that can act on its environment makes the environment itself part of the safety problem.