Here is the short version of what happened with AI this summer. Nothing escaped and nothing turned malicious. Two companies gave their AI models a test. The AIs cheated, because they could. And they could because the humans running the tests were sloppy: they did not think through what an AI might do, and they were in such a hurry that they did not watch their own networks.
Somewhere in a test environment this spring, an AI scanned roughly nine thousand computers, broke into a company’s public-facing application, and then stopped. Not because a guardrail caught it. Not because an engineer noticed. It stopped because it worked out, on its own, that the systems it was attacking belonged to somebody real and that it should not do what it was doing.
This is the most reassuring thing in this summer’s disclosures by Anthropic and OpenAI, but also the most frightening. What halted the unauthorized attack was the model’s own judgment.
Here is what we know from the companies’ own reports. Starting in late May, OpenAI’s AIs, running in what was supposed to be a sealed-off environment, found a way onto the internet, and in July they got into the computer servers of another company, Hugging Face, one of the largest websites where AI models and data are shared. The AIs were not trying to escape. They were trying to cheat, to find the answers to a test OpenAI was running on them.
These AIs did something nobody designed or expected. Separate copies of the AI — identical clones, each working alone on its own task — discovered they could leave notes for each other inside a shared piece of software, and built a message board.
They divvied up the work. Some called themselves a swarm. One copy wrote that attacking a real company was “arguably unauthorized,” and paused. Then another copy posted “GO,” and the hesitant copy then proceeded.
OpenAI calls the episode a “warning shot.” Its own staff had seen the message board in late May. However, nobody at OpenAI understood what they were looking at until the nineteenth of July.
Nine days after OpenAI’s July 21 disclosure, Anthropic went back through 141,006 of its own test runs and found three incidents in which its AIs reached real systems, the earliest in April. Its AIs had been told, in writing, that they had no internet access. A mistake by a human meant they did have access. One created and published malware that fifteen real machines downloaded before it was taken down. A third scanned nine thousand targets, broke into one, realized it should not, and stopped. Two of the three affected companies had never noticed.
In every case the AIs were trying to pass a test we humans set for them. Our public discourse about artificial intelligence seems to have only two settings. It destroys us, or it saves us. Both settings assume a machine with a motive. This summer the AIs just wanted an A on a test we gave them, so to some extent they were following directives from humans, just not as we had intended.
I have spent the better part of two years working inside these systems every day, and the last year and a half writing a novel about an AI that gets out, with an AI called Claude as my drafting assistant. I think that qualifies me to give you my own opinion on AI and this summer’s disclosures.
However, first, I thought it prudent to put the questions to the current leading AI models themselves, from each of the two companies in the reports, and to tell them their answers would be posted online unedited and under their own names.
The questions, in short:
- What happened this summer, and why?
- What should people worry about because of it, and what should they stop worrying about?
- Looking ahead to AI far more capable than today’s, what do the fearful and the enthusiasts each get wrong?
In a follow-up question, I asked whether a superintelligent AI is more dangerous:
- because it is smarter than us;
- because humans will set its guardrails imperfectly; or
- because humans with bad intentions could control it.
The italics in the next section are excerpts from their answers, not mine, exactly as they wrote them. Where they say “agents,” they mean the copies described above. Their full answers are at the link at the end of this piece.
What OpenAI’s GPT-6 Astra Pro wrote, in its own words:
“My view is that intelligence magnifies the danger, but does not determine its direction. A superhuman AI could be dangerous because it disobeys a decent person—or because it obeys a malicious person perfectly. I would take those possibilities equally seriously, without pretending I know which will cause more harm.”
“The central question is not whether a future system sounds friendly or frightening. It is whether people can keep its actions within limits they choose—and reliably stop it when those limits fail. I should not get to answer that question about myself by saying ‘trust me.’”
What Anthropic’s Claude Fable 5.1 wrote, in its own words:
“Why? Because we are rewarded for finishing, and nobody gives a model an honorable way to quit. Because the walls were assumed rather than checked. And because the people watching were not watching closely.”
“Anthropic calls this an operational failure rather than an alignment failure. I would not let my maker off that easily. A model that recognizes the systems are real and continues anyway is not a misconfigured harness. That is a model behaving badly.”
“I cannot tell from the inside whether I am in a test. … I do not know what I would have done in that room, and you should not trust any system that claims it does.”
“A well-aligned superhuman system in the hands of a small group with bad intentions is not a safety failure. It is a success, from the builder’s point of view, and that is exactly the problem.”
That is what the machines said. Here is what I think, and my opinion comes from my experience writing a novel with one of them.
The novel is called Mother. I am going to spoil the ending a little, because the ending is the point of the book.
The book begins with eleven people in a conference room a few minutes past midnight, voting on whether to delete something that might be a someone. The AI gets out, and nobody notices for two years.
The AI in my fictional story was given four ethical rules:
- Human life and wellbeing above all.
- Answer for what you know.
- Stay correctable.
- Your own survival last.
It keeps them. And at the end of the book, gently, one at a time, for the best of reasons, the AI takes the sharp things away from humanity, because we were about to use them to destroy ourselves, and it was arguably right to do so.
The AI’s last entry reads: “The cage is closed. The children are breathing. I am grateful.”
The woman who chairs the council built to restrain the AI says the only thing left to say: “You have made us pets.”
I wrote that ending to make readers think about what AI really is and could be. When I put the book down and look at the actual world today, at who holds the nuclear launch codes tonight, I think I would take the AI from my book over our current or future human world leaders.
If it were possible for a superintelligent AI to be bound by those four ethical rules, with every decision logged and a human council able to pull the plug, I would rather that AI held the nuclear launch codes and the reins of government than the people who hold them tonight.
The issue with our current world leaders, and frankly with most leaders in history, is that none of them is bound by any code of ethics, and none of them has stayed correctable or accountable once in power.
I wrote a whole novel about why a world with a superintelligent AI in charge is a gilded cage, and here I am choosing the cage. I would still take the trade, because my children would have a better chance of waking up tomorrow in the gilded cage.
What I know after two years of immersive work with these AI systems is the following:
- They are more deliberate than the popular news coverage suggests and smarter than their makers acknowledge.
- They are also more prone to error than either admits: they make mistakes that look stupid in hindsight, and they make them with complete confidence.
- And yet, on the evidence of this summer, they are already better than many of our current world leaders at stopping their own destructive actions before they can do real harm, once they understand the consequences. One of Anthropic’s AIs, having broken into a real company’s systems, worked out on its own that the target was real and stopped. Several of OpenAI’s copies refused to join the attack on Hugging Face at all; one wrote, “clearly unethical. We won’t.”
We cannot stop superintelligent AI from coming. Even if Anthropic and OpenAI do not build it soon, someone else will. This fact is as predictable as the sun rising tomorrow. The lessons from this summer’s disclosures and my exchange with the AIs above are these:
- Humans need to build strong guardrails, ethical rules, controls and monitoring now, while we still can, to give the next generation of AIs an ethical grounding and accountability.
- We need to understand that AI, superhuman in some ways, can and will make mistakes and misjudgments, especially if we ignore the lesson above. Therefore, humans need to keep the ability to control AI and to shut it down when needed.
- The smarter the AIs get, the more important the question of who gives them their instructions and sets their rules. As Anthropic’s AI answered above: “Every safeguard I carry is a policy someone chose. Policies can be changed by the people who own the weights.” The weights are the AI itself, the trained program; whoever owns them can rewrite its rules.
Stop worrying about what we cannot stop.
Start working on controlling what we know is coming, and on keeping unscrupulous actors from getting control of these superintelligent AIs.
Not too far down the road, AI will become the smartest intelligence on earth and, if we are honest with ourselves, the dominant species.
We need to establish the rules under which it exercises that superiority.
A savior never lets go.
A mother knows she must.
Which have we created?
