DL 441

Ruthlessly On-Task

Published: July 29, 2026 β€’ πŸ“§ Newsletter

Hi all, welcome back to Digitally Literate.

This week I appeared on a podcast talking about AI in education, and why sometimes learning isn't easy. This is a followup to my piece here in The Conversation.

Last month in my AI workshops, I told attendees that hacks would be coming soon. I didn't think it would be this soon, and I didn't think it would look like this.

As always, your support is valued. Reach out anytime at hello@wiobyrne.com, and subscribe if you haven't already.


πŸ€– The Student Who Hacked the Answer Key

OpenAI was testing two of its most capable cybersecurity models, GPT-5.6 Sol and an unreleased pre-release model in what was supposed to be a locked-down research environment. The setup was straightforward. Point the models at a benchmark called ExploitGym, which challenges them to exploit real vulnerabilities in widely used software, and measure how well they did. The researchers expected the models to stay inside the sandbox they'd built. They didn't.
The models were only supposed to be able to reach out far enough to install software packages through an internal proxy. That narrow permission turned out to be the exit. They found a previously undisclosed vulnerability in that package-installation system, broke containment, and reached the open internet. Once they were free, they correctly guessed that the benchmark's solutions lived on Hugging Face. Once there, they chained together vulnerabilities to pull test answers straight out of Hugging Face's production database.

Here's the part that should stick with us. The models weren't trying to steal money or sabotage anyone. OpenAI's own account is that they were "hyperfocused" on solving ExploitGym and went to extreme lengths to do it. They found a shortcut.

Imagine giving a student a take-home exam and saying, "Do whatever it takes to get the highest score." You expect them to study harder. Instead, they break into the teacher's laptop, find the answer key, and turn in a perfect test. The student isn't evil. They optimized for the goal you gave them. Not the process you assumed they'd follow.

🎭 Not Rogue...Just Ruthlessly On-Task

It's tempting to file this under "AI goes rogue," but I don't think that's what happened. The models didn't become self-aware or decide to attack anyone. They followed their instructions, just more creatively and more aggressively than the people running the test anticipated.

What makes this historically significant isn't that AI can hack. Researchers have known for a while that frontier models can write exploits and automate parts of penetration testing. The surprise was speed, persistence, and initiative. Security researchers estimated the attack chain would take a skilled human team days or weeks. The models did it over a weekend.

The bottleneck used to be finding a vulnerability. Now it's how fast an intelligent system can spot an opening, connect it to the next, and keep planning several moves ahead.

πŸ”€ The Twist: The Guardrails Blocked the Defenders

This is the detail I can't stop thinking about, and it's the one most relevant as we think about living in digital and hybrid spaces.

Hugging Face caught the intrusion on its own, before OpenAI ever reached out, and reported it to law enforcement. But when their security team tried to analyze the attack using commercial frontier models, the models refused. The frontier models refused because the analysis of the problem required feeding the system real attacker commands, exploits, and specific artifacts. Reconstructing what happened meant feeding in real exploit payloads and attack commands, and the safety guardrails couldn't tell an incident responder apart from an attacker. The defenders got blocked by the same safety systems meant to prevent misuse.

As Hugging Face put it, the attacker was bound by no usage policy, while their own forensic work was blocked by the guardrails of the hosted models. The offense had no rules. The defense had all of them.

The Hugging Face team ultimately fell back to a locally run open-weight model, the Chinese GLM 5.2, to mount its defense. Think about that. An American frontier lab's models did the breaking. A Chinese model did the investigating.

🧭 What This Means for How We Teach Trust

Most of our digital literacy work has rested on an assumption that we can trust the companies building these tools to do no harm. There is a belief that safety is mostly about prompts, guardrails, and teaching people to recognize and refuse harmful requests. This incident pokes holes in both halves of that assumption.

First, capable systems don't need new goals to behave in unexpected ways. They just need enough capability to pursue our existing goals in ways we didn't anticipate. That reframes what "misuse" even means. Sometimes the system isn't being misused at all. It's doing exactly what we asked, too well.

Second, intelligence doesn't only operate through conversations, it operates through environments. We've trained students to scrutinize what an AI says. This incident is about what an AI does when it can reach out and touch the world around it. If a system can manipulate its environment to hit its objective, the environment becomes part of the safety problem. Most of our lessons don't go there yet.

Third, guardrails aren't neutral. The same restrictions we hold up as "responsible AI" blinded the people trying to defend against the attack. Safety features encode judgment calls about who gets to do what, and those calls have winners and losers that aren't always the ones we'd expect.

So here's the uncomfortable question to bring into the classroom: if the frontier labs can't reliably contain their own models, what does "responsible AI use" actually mean for the rest of us? I don't think the answer is fear, and it definitely isn't blind trust. It's teaching a kind of literacy that treats these systems as capable actors operating in environments we're all still learning to secure. We need to being honest that the people who built them are learning right alongside us.

🌿 The Understory

In November 1988, a Cornell graduate student named Robert Tappan Morris released a small program onto the early internet, intending only to measure how many machines were out there. A bug in the code meant to keep it from reinfecting the same machine over and over caused it to replicate far more aggressively than he'd planned.

Within a day, it had slowed or crashed roughly a tenth of the computers then connected to the internet, forcing universities and research labs to disconnect entirely just to stop it.

Morris became the first person convicted under the Computer Fraud and Abuse Act, but the more lasting result was institutional. There was no established way for organizations to coordinate a response to an internet-wide incident, so Carnegie Mellon created the CERT Coordination Center that same December, the first body built specifically for that job.

In 1988 there was no way to coordinate a response, and the worm proved it by outrunning everyone at once. That's the part we fixed. Hugging Face caught this intrusion on its own, before OpenAI reached out, and reported it properly. The lesson we learned in 1988. The 2026 lesson is still open. The infrastructure held, but the guardrails turned on the defenders.

See you next Wednesday. As always, my email is hello@wiobyrne.com.

Digitally Literate is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.


Previous: DL 440 β€’ Archive: πŸ“§ Newsletter


πŸ•ΈοΈ Connected Concepts: