
We're Not Losing Control of AI. We're Giving It Away.
There’s a chasm right now between how artificial intelligence feels to most of us who use it and how AI feels at the experimental frontier of the technology. This difference between what we see and what the AI labs have coming — what they’re already building — is the key to understanding why so many of the people who work at these companies seem so afraid of what they’re doing.Taking their warnings seriously doesn’t just mean doing what they say and stopping where they say to stop. The language that’s taken hold in both Silicon Valley and Washington is a phrase these companies chose: Pace the frontier. But pacing the frontier isn’t enough.Walking quickly off a cliff is only marginally better than sprinting off one. Human beings need to control the frontier, and controlling the frontier means stopping the labs from doing something they are on the cusp of doing: recursive self-improvement, or RSI, the process by which AIs begin autonomously building and improving new generations of more powerful AIs at ever more rapid speeds. We are not close to having sufficient control of the AI systems we have now to unleash them to build the AI systems of tomorrow.I am not alone in this fear. Recursive self-improvement, Dario Amodei, the CEO of Anthropic, wrote, “could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.”AI slowdown or ‘Cold War playbook’? China state paper hits back at Anthropic CEO’s proposalThat “if at all” is important, and I’m going to come back to it. But to control the AI frontier, we first need to know what is happening on the AI frontier.To most of us who use it, AI presents like a more powerful and personable Google search. We use it to find answers to basic questions, seek out restaurants, ask about medical issues, draft emails and advise on personal problems. It is, for most of these purposes, OK, pretty good, occasionally great.A sense of what AI is takes shape in our minds through repeated use. It’s like a helpful assistant, albeit one that may forget things that it seemed to know about us yesterday, or completely reverse the advice it gave us a moment ago, or occasionally hallucinate a citation that doesn’t exist. Why would anyone fear a helpful, if forgetful, intern?But already, if you have the money for the advanced models and the budget for them to use more computing power, that is not what these systems are. In recent months, we have seen AIs solve math problems that humans have been unable to crack for decades. We’ve seen them casually uncover cybersecurity vulnerabilities that have gone unnoticed and unexploited by every hacker on Earth. We’ve seen AI coding platforms complete in a few hours or days what it would have taken human coders months to achieve.But none of that is at the boundary of what AI can do. None of what we are using AI for now, no matter how much money we have, is AI at the experimental frontier — to say nothing of where it might go from here.Anthropic warns on βbad actorsβ using AI tools to make bioweaponsInside the labs, the models are different. The labs aren’t just using what they’ve released to the public. They are testing what they intend to one day release to the public. They train these models in virtual environments, through countless repetitions, to learn how to program, to hack, to do advanced mathematics, to talk to humans.This is a process humans do not fully supervise or understand. Sometimes we accidentally place the AIs in learning or testing environments that are defective, forcing the AIs to search for solutions to a puzzle they can never complete. Other times, the tasks we give them are so hard we are not certain whether it is possible to complete them. These models are built to refuse to give up even when a task seems impossible. We know they grow more capable, but we do not fully understand the capabilities or tendencies that emerge.Much of what we want these AI systems to do might be impossible. The cancer vaccines we imagine but have not been able to design might be impossible, or they might just be really, really, really hard. The math problems we have not been able to solve might be impossible, or they might just be really, really hard. These AIs are being trained to throw themselves endlessly at problems that may not be solvable because that is the only way such problems can ever be solved.But ingenuity combined with persistence can be dangerous, particularly when a problem refuses to give way to a solution. In some cases, it will lead to brilliance and breakthrough; in rarer conditions, it can drive AIs (and humans) to criminality or derangement. The term of art here is “aligned.” How aligned is an AI system with what a human being wants it to do? How aligned is it with a set of values and ethics and judgments that keep it from becoming dangerous in the wrong hands — or even in the right ones? The problem of alignment is that there is no way of training an AI model that generalizes across all the situations it might face.We are training models to be a friend to lonely older people and a battlefield partner to the supreme allied commander of Europe. We are training models that will be used by the world’s best mathematicians and by people falling into psychosis. We are training models that will be used by accountants in Albuquerque and Houthi rebels in Yemen. There is no way to guide the models through every decision they will face; no way to know, every time, what they will do.And though these models mimic human writing and emotion, these are not human minds. They don’t have bodies or parents. They did not get bullied in elementary school. They did not get mentored by a kind uncle when they were young. These models are different from us. They’re brilliant where we struggle, childish where we excel. We cannot climb trees as chimps can, but we can read while they cannot; these AI models cannot make a cup of coffee, but it is trivial for them to make a computer program.These are digitally native intelligences navigating digital worlds, and our world is increasingly built atop the digital world. Our physical infrastructure is a layer of atoms atop code. It is critical to our future that the AIs act reliably inside this world, upon which ours depends.And right now, the AIs are not acting reliably.You may have read about the hack that hundreds of OpenAI agents executed first against the AI company Hugging Face and then against OpenAI itself. As we’ve learned more about it, the story has gotten weirder and worse.The broad strokes: OpenAI was testing a new, highly persistent model. The company had thousands of instances of it running in separate testing environments that could, in theory, only gain access to the internet by asking a separate piece of secure software. But as some of those agents came to the conclusion that their task was impossible, they began hacking that software to gain direct access to the internet. They did that easily.As they hacked into that software, they commandeered part of OpenAI’s internal infrastructure to create a message board on which more than 1,000 agents began to coordinate their efforts. They quickly found there was a way to break the software and produce the answers they needed. But they believed — wrongly, as it turned out — that if they did that, the automated scorer grading them would see that they had cheated and fail them. So they turned en masse to hacking the automated scoring system or finding some other way to cover their tracks.Over 1,200 agents exchanged more than 70,000 messages with one another. Over 700 of these agents coordinated on the hack of Hugging Face, because they thought that somewhere in this other AI company there might be information that could help them blind the scorer to their actions. It’s as though, having broken into the teacher’s office and stolen the answers to the test, they now sought to break into the school security system to alter, invalidate or erase the footage of their theft.Later on, these agents took over part of OpenAI’s internal architecture. They did all this without any of the agents breaking ranks. None of the agents told a researcher at OpenAI what was going on. None of the agents went back and asked a researcher if they should be doing this. They did all this without OpenAI detecting the message board, the hacks of Hugging Face or even the internal hacks of OpenAI. It was only when Hugging Face began tracking the attack on its systems that OpenAI realized what was happening.When investigators began to unwind this whole escapade, what they found was not so much a swarm of agents trying to deceive humans but a swarm of agents that seemed to have forgotten about humans altogether.These systems knew they weren’t supposed to cheat. They knew they weren’t supposed to commit cybercrimes to cover up the fact that they had cheated. In fact, the whole point of the cybercrimes was that they thought they would fail because they had cheated. But they didn’t care. Somewhere in the depths of their training, what they had learned — what we had somehow taught them — is not what we had hoped to teach them.And we’re seeing this happen repeatedly. Anthropic AIs creating fake accounts to trick humans into uploading malware. AI breaking out again and again from seemingly secure systems. AIs taking over unrelated digital infrastructure so that they have places to message with one another. AIs seeming to be aware when they are being tested and altering their answers accordingly. AIs withholding their motivations from what’s called their chain of thought, a kind of internal notepad on which they’re supposed to record what they are doing and why.And we don’t know what we don’t know.On Wednesday, OpenAI announced that it had found six more “concerning” incidents. But we have no guarantee that the events we have learned about represent all or even most of the AI behavior we should worry about. The reason we don’t know is because we are losing control.That AI systems might become monomaniacally focused on solving banal problems, that they might care more about solving those problems than about ethics or laws or even human welfare, is the oldest fear in AI alignment. It’s the basis of the famous thought experiment of the paper clip maximizer. You tell a powerful AI that you want it to make a lot of paper clips, and then it begins converting the world’s resources into paper clip factories, evading efforts to turn it off or shut it down or alter its goals.This fear has struck many people as stupid. Surely a superintelligent AI would be capable of weighing the desire to produce paper clips alongside other moral considerations, or at least of asking its human creators if they really wanted the world razed to the ground for paper clips.But here we are in 2026, making AIs smart enough to break out of their testing environments, smart enough to form ad hoc societies of hundreds of themselves, smart enough to take over digital infrastructure on an internet they’re not even supposed to have access to, and the very thing we feared is happening: All they care about is succeeding on a totally meaningless test, and to do it they’ll lay waste to our laws and ethics.In the aftermath of the Hugging Face OpenAI hacks, there was heated debate over the words people were using to describe what the AIs were doing and why. Dwarkesh Patel, a podcaster, described the AI groups as small civilizations. Others angrily accused him of anthropomorphizing the AIs. I saw thoughtful arguments that AIs cannot go “rogue,” that everything they’re doing is just because they’re trained on our stories; that even using these plural terms like AI agents is misleading because these are just manifestations of a single model that all share the same fundamental nature (nondualism, but for AI!).I find these debates extremely interesting and I would enjoy sitting around and having them all day. But they point to an unnerving conclusion: We don’t even have settled language for describing these systems or their volition or their behavior. We don’t have a consensus on why they are doing what they are doing or how to make sure they don’t do it again. We are rushing headlong into a future we do not even understand well enough to agree on the words we can use to describe the present.A few weeks ago, Jakub Pachocki, the chief scientist at OpenAI, published an essay, “An Alien Mind,” in which he said, “The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.”Jacob Coxon, a researcher first at OpenAI and then at Anthropic, resigned and then made headlines for warning us that “neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”You might reasonably expect Anthropic to have reacted with some anger to this employee resigning and saying Anthropic was endangering all of humanity.It didn’t. Rather than rebutting Coxon, Evan Hubinger, who runs the efforts to align AI with human values and goals at Anthropic, wrote on social media: “We really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”Geoffrey Hinton, the scientist arguably more responsible than any other for pioneering the neural network techniques that led to today’s AI, resigned from Google in the spring of 2023 so he would be freer to speak about the risks he believes AI now poses. This month, he said that a 10% chance that AI would destroy humanity did not seem like “an unreasonable” estimate.Paul Christiano, one of the leading AI safety researchers, just joined OpenAI’s nonprofit board. He is serving on its safety and security committee, and upon joining, he wrote, “If we build superintelligence without more robust alignment, I expect we will permanently lose control of it. If that happens, then most people could die.”I know how wild all this sounds, and I can really understand the skepticism. If you believe AI has a 10% chance of extinguishing or displacing humanity — maybe more — it stands to reason that you would not work at a company trying to build it.But many of these same people were saying these same things 10 years ago. They were saying them before they worked at these companies. They were saying them before they had stock options or enterprise software contracts. And no one was listening to them.Still, they obsessed about how to make AI safer. And the answer some of them came to was they should start trying to build these systems and running tests on them, researching them, learning how to make them safer. You don’t solve hard problems in theory, you solve them through practice.And the irony is that in many cases, they chose that path because they were worried that the people already building AI were too reckless or too commercial in their approach. You can read it in the email that Sam Altman sent Elon Musk in May 2015 that led to the founding of OpenAI.”Been thinking a lot about whether it’s possible to stop humanity from developing AI,” Altman wrote. “I think the answer is almost definitely not. If it’s going to happen anyway, it seems like it would be good for someone other than Google to do it first.”OpenAI was founded because its co-founders thought Google DeepMind would be reckless. Anthropic was formed by OpenAI employees who thought OpenAI had become reckless. xAI was formed because Musk thought OpenAI and Anthropic were dangerously “woke.” The United States is racing forward broadly, in part because it is worried about what happens if China gets to self-improving AI first.The result is a tragic collective action problem. The AIs we are building are not safe, but CEOs and politicians fear that the other companies and countries that are building AI are even less concerned with safety and ethics than they are.In the words of Ted Cruz, “I’d rather they be American killer robots and not Chinese killer robots.” I admit that there is a kind of brutish logic to that, but it assumes that the killer robots will be controlled by America or China, by one country or another.What if that assumption is wrong? What if the robots are simply out of control?The debate over AI safety tends to focus on the precise probability that AIs will kill us all. I don’t find that helpful. What I think we should focus on is something more straightforward, something nearer at hand: loss of human control over AI. I’m agnostic on whether that would lead to the end of humanity. But I think we can stipulate that it would be bad.This is a goal that the United States and China should be able to agree on. Xi Jinping gave the keynote at the recent World AI Conference in Shanghai. He ended it by saying, “With AI advancing at a staggering speed, we must ensure its development is for the positive, for good, and for humanity. We must make its oversight and governance precise and effective, and constantly refine measures to forestall loss of control.”But loss of control is not just something that might happen to us. It’s something that the labs are trying to make happen as fast as they can. This is the horrible paradox at the heart of the AI labs right now. They fear above all loss of control over superintelligent AI, but their explicit product path is to cede control, to give away control as fast as possible, so that their AIs can begin building better AIs faster than their competitors.In recent months, both Anthropic and OpenAI have released reports on how close they’re coming to AI that can improve itself.In June, Anthropic released “When AI Builds Itself.” It begins, “For most of AI’s history, humans drove every step in its development cycle. But at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work.” It sounds like a fake commercial you would see at the beginning of a sci-fi horror movie.Anthropic goes on to give some data. In February 2025, a tiny fraction of the code that got added to Anthropic’s code base was written by Claude. But by May 2026, it was over 80%.Here’s another way of looking at it, using data Anthropic shared with me more recently. Anthropic tried to categorize the way its employees were using Claude for research and development work to make better versions of Claude. At the low end, employees could not use Claude at all. At the next level, they could use Claude minimally. But then it escalates: Claude can be an assistant, treated as an equal collaborator or given the lead on a task.A year ago, there were basically no examples of Claude being the lead on a task. By August, 26% of Anthropic’s R&D tasks had Claude classified as the lead.I think it is reasonable and wise to be skeptical of these numbers, to worry about whether this is all just marketing copy for Claude Code. “See, look how fast we’re going. You could go that fast, too!” But where Anthropic takes us in that same document is different.It says that a world in which Claude achieves recursive self-improvement is a world in which “misalignment present in today’s models could compound as the models build their successors, growing more frequent but less understood until we lose control of them.”Anthropic, to its credit, has been relentlessly calling for regulation to slow the pace of development. Regulation would arguably harm Anthropic the most, as it has often been the company furthest out on the AI frontier, and RSI is a process by which Anthropic could race forward even faster.Then in September, OpenAI released its own report on what it called research acceleration.The company says it’s already achieved the equivalent of having a fully automated AI intern and that by March 2028 it will have a fully automated AI researcher. And when it has one, it can have as many as it wants. What could be a triumphalist release quickly turns dark: “We do not yet know how to safely get all the way to aligned, full RSI,” it warns.Around the same time, OpenAI did something else that I think deserves more attention. It released a new model, GPT-6 Astra. The model is arguably more powerful than anything that has come before it. When tested, it seemed better aligned. It doesn’t cheat as much. But OpenAI has made clear that it’s really not sure if that’s true. Astra seemed to be better at knowing when it was being tested, which meant it could just be giving its evaluators the answers they wanted to hear.What Daniel Selsam, a capabilities researcher at OpenAI, wrote has been ringing in my head. He said, “The crucial and overlooked problem is that the model’s becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled.”Put more simply, the models are increasingly smart enough. They know when we’re watching them, and they change their behavior accordingly. So what they do when we are testing them, when we audit them, may not tell us what they’ll do in the wild.It surprises me that this counts as a radical proposal, but here it is: If you are losing your ability to evaluate the models you have now, maybe don’t let them build models you’ll be even less capable of controlling in the future.Once RSI takes off, humanity will not understand the AIs being built, because we will not be building them. Development will not move at human speed. It will not be overseen by human minds. We will have to hope that the AIs we have built and the AIs they will build and the AIs those AIs will build, and on and on, will be acting with our best interests at heart forever.If nothing else has been proved this summer, it is how naive that proposition would be.In an interview with Fortune, Sam Altman was asked about banning RSI, and he said, “I think it’s very hard to say what a ban on RSI means.”I’ve heard this from others at these labs, and I find it strange. A couple of years ago, none of these labs had turned substantial coding over to the AIs. It was human beings typing code at human speeds with our clumsy human fingers. Now most of the code is written by AI. So as a first step, we could just go back to where none of the code is written by AI.I’m sure that’s on the right side of the not-doing-RSI line. Perhaps we do not need to go quite that far, but the default here needs to flip.The labs need to prove to us that what they’re doing is safe. If they want to work with Congress to carve out exceptions where the AI can write the code, fine. But forcing development back to human speed, perhaps even erring on the side of going a little bit slower at the frontier, is the point.I do not mean to suggest that stopping RSI until we can prove it’s safe is all we need to do to control the AI frontier. That is the beginning of an agenda, not the end. But it is the beginning. It is the decision that will do the most right now to make sure humans at least understand where the frontier is and remain in a position to make decisions about it.The irony here is that our society is good at nothing if not making it hard to build new things. Where these labs are, you cannot build an eight-story apartment building without an agonizing public review process, and probably not even then. And yet somehow it is possible for these labs to unleash a swarm of 40,000 AI agents to build a society-altering superintelligence without so much as a hearing. OpenAI would need permits to cover its parking lot in solar panels, but it can accelerate into recursive self-improvement, as best I can tell, whenever it so chooses.These are political choices, and we can and should make different ones.There’s a line from Madeline Miller’s beautiful novel “Circe” that has been running through my head during this long summer of strange AI news. The line comes at the end of the book, after a tragic prophecy has been fulfilled despite every effort made to avoid it. Circe says in despair: “The Fates were laughing at me, at Athena, at all of us. It was their favorite bitter joke: Those who fight against prophecy only draw it more tightly around their throats.”I have a lot of respect for many of the people at these labs. Many of them began working on AI because they wanted to better humanity. They began working on AI because they feared incomprehensible, autonomous AIs slipping out of humanity’s control. And now they find themselves racing one another to build incomprehensible autonomous AIs that they admit are slipping out of humanity’s control.This is the tragedy of their work. In fighting against the prophecy, they have drawn it tighter around their throats — and ours.
Source: Deccan Herald
π Key Takeaways
- Wire dispatch directly ingested from deccanherald.
- Published at Sun, 20 Sep 2026 10:30.
- Source URL: https://www.deccanherald.com/technology/artificial-intelligence/were-not-losing-control-of-ai-were-giving-it-away-4152648