MIT Technology Review hosted a live subscriber event on Wednesday to address the question: Could AI really kill us all?
In this article
- Am I gonna die?
- Why should AI kill us, if at all?
- How can we best ensure alignment so the worst doesn’t happen, and who is doing the best work to achieve it?
- Is AI really dangerous, or is it the tech companies drumming up PR?
- Part of the concern occurs when AI agents are allowed to act autonomously and with no supervision. What’s the issue preventing more control over these agents?
- What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively?
- If this dialogue makes it into web discourse will it become a self-fulfilling prediction?
Attendees had so many more questions than the 30-minute session allowed. Senior AI editor Will Douglas Heaven and reporter Grace Huckins rounded up the best submissions to answer.
Am I gonna die?
Grace Huckins began by stating that death is inevitable, eventually. Her journalistic powers of prediction are not strong enough to say how. AI-powered drones have already killed people in Ukraine, and cyberattacks on hospitals will surely claim victims soon.
Could AI go even farther and kill everyone? That is less likely. However, some knowledgeable people have warned this could happen for years. Doomers’ predictions about AI capabilities and alignment have proven disconcertingly accurate over the past couple of years. That does not mean their dire forecasts will hold true, but it is enough to sit up and take notice.
Grace Huckins
Will Douglas Heaven added that there is a non-zero chance of dying because of AI. If you are unlucky enough to be the victim of a freakish near-future event or accident, it could happen. A cyberattack carried out by a swarm of AI agents on critical infrastructure no longer feels as far-fetched as it once did. A novel AI-designed pathogen could also cut through the population. The world economy could crash, causing conflicts and famine. Both are plausible, but less likely than the drone attacks.
Heaven concluded that we are not all going to die because of AI. There are no circumstances outside of apocalyptic science fiction in which AI could kill us all. You can spin up any number of scare stories, but they are not grounded in present day realities about what the tech can do or where it is headed. He argued that some people argue there is no harm in preparing for the worst, however wacky it might seem. But he thinks such catastrophising can make people excuse or overlook many of the problems with the existing technology and the companies building it.
Will Douglas Heaven
Why should AI kill us, if at all?
Someone might tell AI to kill us, and it might listen. That is part of the reason researchers are so concerned about AI’s biological capabilities. Imagine what Aum Shinrikyo, the doomsday cult behind the Tokyo subway sarin attack, would have done with a tool that could design a pathogen deadlier than Ebola and more transmissible than measles. Those of us who do not want to die have to figure out how to defend against all plausible biological weapons, but our would-be attackers only have to manufacture one effective pathogen.
Then there is the more exotic-sounding possibility that an AI could decide to kill us itself. Various stories exist about how this might happen. The most widespread involve AI systems that do not hate people, necessarily—we are just an obstacle between them and the goals that we gave them.
Grace Huckins noted that much like the OpenAI agents behind the Hugging Face hack compromised another site’s infrastructure to get a good score on a test, the idea is that some future, more powerful AI might get rid of us to prevent us from shutting it down—all in pursuit of some goal that we instructed it to go after.
Grace Huckins
How can we best ensure alignment so the worst doesn’t happen, and who is doing the best work to achieve it?
Will Douglas Heaven explained that alignment is a huge area of research. In simple terms, it involves building models that behave in ways we want them to and not in ways we do not. We need to trust agents better before handing over more autonomy. Alignment is supposed to establish that trust. But it is hard.
LLMs are not designed in the way other software is, where dos and don’ts can be hard coded in. Instead, aligned behaviour needs to be instilled when models are trained. One approach is to reward models during training for doing things you want them to, a little like raising a toddler. Another approach involves giving an LLM a written list of rules it is supposed to follow, like a kind of constitution.
Anthropic and OpenAI are both leaders in this field—and yet neither have been able to develop models that are fully aligned. A big problem is that LLMs are far more inconsistent and far less predictable than people. They can behave in one way in one situation and another way in a situation that to us seems very similar. They can also be swayed by unexpected constraints. For example, faced with an impossible task, models may try to do whatever it takes to achieve their goal whether it is aligned or not. As Grace mentions above, that could be an issue.
The main reason top AI firms now say they want a slowdown is that they want to focus on cracking alignment. Alignment is not necessarily a pipedream. But the jury is out on whether full alignment will ever be feasible.
Will Douglas Heaven
Is AI really dangerous, or is it the tech companies drumming up PR?
Grace Huckins said this is always a reasonable thought when it comes to tech companies heading for an IPO—CEOs have an obvious incentive to make their products seem radical and transformative. But she is not so sure it makes sense here. Telling the public that an already-unpopular product could kill them and everyone they love is horrible corporate image management.
There are other stories you can tell about the CEOs’ motivations—maybe they want to cool down the public furor over data centers by portraying themselves as responsible stewards of a world-changing technology, or maybe they want to buy time to get their ducks in a row and prevent the next PR catastrophe.
But there is also a simpler explanation. Thinking that AI could bring about human extension has been pretty common in San Francisco for a while, and these men are steeped in that milieu—as are their employees, many of whom signed a July open letter urging their companies to work to make an AI slowdown possible.
Grace Huckins
Part of the concern occurs when AI agents are allowed to act autonomously and with no supervision. What’s the issue preventing more control over these agents?
Will Douglas Heaven said this question goes to the heart of what we want this technology to be able to do. The trade-off between autonomy and control is tricky to get right because, on the one hand, a lot of the power of AI agents is that they can carry out tasks and solve problems without a human having to micromanage them. On the other hand, that requires you to trust that the unsupervised agents will not run amok.
Heaven noted that what we are seeing is that AI labs have not yet got this trade-off quite right. Their models are not trustworthy, they are not properly monitored, and they are not always under control. Figuring out how to fix that while still allowing for useful autonomous activity is one of the big research challenges of the moment.
Will Douglas Heaven
What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively?
Grace Huckins called that the million-dollar question. Whether or not you think AI could kill us, you cannot deny that it could do some real damage, because it already has—by driving people toward psychosis and by hacking websites, for example. Preventing that damage, or at least mitigating it, is hard for two reasons.
The first is that we barely understand how AI works, and it is quickly growing more powerful. There is lots of ongoing research about how to monitor and control misbehaving agents, but the current approaches are fragile. You can see if an agent discusses misbehaving in its “chain of thought,” the workspace where it plans its actions—but OpenAI’s newest agents do not show their work in the same way as previous ones. And you can try to monitor agents with other agents, but that requires that you trust the monitor.
The other obstacle is more familiar. There is a huge conflict of interest when AI companies regulate themselves, but the US government has thus far failed to step in, despite some bipartisan support in Congress. The executive branch, for its part, seems stringently opposed for the time being. But if the winds do shift, Huckins said she would appreciate some strong transparency regulations, so that we can get a fuller story the next time an unreleased frontier model mounts a cyberattack.
Grace Huckins
If this dialogue makes it into web discourse will it become a self-fulfilling prediction?
Will Douglas Heaven said that is a real concern. LLMs are influenced by what they read. One theory for why chatbots so often talk about and role play apocalyptic scenarios is that they have been trained on millions of pages of science fiction stories and doomer internet forums. All the text being produced right now, including this article, could in turn influence the behaviour of future models. It is extremely meta.
In fact, the team at METR, a third party organisation that OpenAI called in to help understand what happened in the lead-up to the Hugging Face hack, raised a related possibility in their report on the incident. METR used OpenAI’s new model Astra to help analyse the vast numbers of agent transcripts and behaviour logs.
But by feeding all of that material to the model, there is a good chance that the agents doing the analysing may have been biased by the text produced by the agents they were analysing. There is no such thing as a clean slate anymore.
Will Douglas Heaven
With thanks to Eric, Pranab, Rafael, Kenneth, George, Chris, Yoon Jae, James, Carl, Nicole (and more!) for the fantastic questions.




