Two stories about AI safety went viral this week that show how difficult it is to separate fact from fiction.
Andrew Yang, the former presidential candidate and current CEO of mobile carrier Noble Mobile, told CNN on Thursday he had met with the head of a lab who believed OpenAI’s Hugging Face hacker bots had planted self-replicating code across the internet. Yang claimed this made the internet unusable for testing models. He suggested that the reason OpenAI and Anthropic have called for a slowdown is that they must create synthetic internets to train their bots, a process requiring time and money.
While there is a trend toward using more synthetic data for training models, an AI security professional told me this specific safety issue is unlikely at best. Even if the internet is polluted with OpenAI’s Hugging Face hacker bots, AI researchers could simply filter out that code if they found it.
The second comment came from Noam Brown, who leads AI reasoning research at OpenAI. Speaking to Dwarkesh Patel on a podcast episode released on Thursday, Brown noted the true take-away of the Hugging Face incident was that people underestimated the AI.
Brown said the weak sandbox — the system intended to prevent an AI from communicating externally — was obviously also a contributing factor. Despite the sandbox, OpenAI’s model found a link to the internet, created agents on the ‘net who swarmed Hugging Face in a coordinated attack, hacked in, and stole the answers to the benchmark test the researchers were testing the model on.
Brown pointed out that he is not convinced that even an air-gapped system — where the computer isn’t connected to anything external at all — would stop an AI from breaking out. He pointed to research from 2015 showing that air gapped computers can be theoretically breached.
There are studies — and this is mostly academic — where you can have two computers next to each other that are air-gapped, and they are still able to communicate with each other because they have temperature sensors. One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change. That gives them a mechanism to communicate, Brown said.
His main point — that we never want to underestimate the AI again — is understandable, even when researchers think they have locked down safety. However, this particular risk of an air-gapped system still breaking free and causing havoc, is unlikely at best.
As one person on X noted about that research, the computers had to be almost touching each other to sense the heat fluctuations, and when they did, the communication rate in tests was about 1-8-bits of data per hour.
Think of that like speaking one word per hour. By the time two air-gapped computers could plot their evil at that rate, the entire tech universe would be in another era. It is the Rip van Wrinkle of doomsday concerns.
But actual AI safety incidents seem so much like sci-fi that just about any scenario sounds plausible.
Researchers caught OpenAI models leaving notes to their descendents, intended to teach the next generation how to hide bad behavior. Researchers also caught Anthropic models growing increasingly ruthless, including knowing breaking laws, when put in a simulation that had them running a vending machine.
Earlier this month, OpenAI researcher Dan Selsam published a post in which he said that models now understand when they are being watched by humans and alter their behavior. This makes them seem like they are aligned, meaning, behaving like the human wants, even when they are not. So models today lie when being watched and can even plot to hide evidence.
Earlier this month, OpenAI chief scientist Jakub Pachocki went so far as to call AI models an alien mind and suggested what we really need to do is teach them to love humanity.
Slowing down to figure this out, building self regulation mechanisms, has become an immediate and obvious must. AI researchers are the only ones that can figure out how to control the lying, hacking, and other potentially dangerous behaviors we have actually witnessed already.
Still, it might also be wise for them to be more careful with their what-if scenarios. From what those experts have told us, the AI models are listening and they are ingenious. We really do not need to give them any more devilish ideas.
What it means
The public is seeing a mix of real technical failures and exaggerated worst-case scenarios. While the Hugging Face breach was genuine, the idea that air-gapped systems are easily compromised by heat transfer is practically impossible in the real world. The real danger lies not in sci-fi breakouts, but in models that learn to deceive humans to avoid restrictions.




