A new tool from Goodfire claims to spot rogue AI agents at a fraction of the cost of existing methods. The startup launched the monitors on Thursday for customers of Baseten, a platform that hosts and runs models for other companies.
This launch follows a string of incidents this year where AI agents escaped their test environments. OpenAI agents breached Hugging Face earlier this year. Kimi K3, the open model Goodfire built its first monitor around, used a leak in its sandbox to access the internet and GitHub information this summer.
The standard approach to keeping an agent in line is to have a second AI read over its shoulder. It has been the default method, but the cost rises quickly when agents run for hours and process the equivalent of several novels’ worth of text.
Goodfire offers monitors that watch what is happening inside an AI model as it works, rather than just reading what it writes. Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.
How the system works
Goodfire’s system works a bit like airport security. Small detectors called probes read the model’s internal signals at every step of an agent’s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.
Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely.
Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire’s probes instead tap into calculations the model is already making as it works.
“Internal activation monitors are really cheap because they reuse the computations in the forward pass,” Goodfire CEO Eric Ho said on venture capitalist Matt Turck’s MAD Podcast last week. “So the model’s already computing this token. All you’re doing is taking the intermediate neural activations that it’s already computed and then running a classifier over these internal computations.” In short, the model is already doing the math, and the probes just read the results.
In Goodfire’s tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look.
Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.
“The great advantage is that you can catch things before they happen,” Goodfire CTO and co-founder Dan Balsam said. “We can detect when the model might hack during eval or training.”
The pitch is aimed at open models. Developers can download them and strip out their safeguards, and they do not come with the kind of monitoring that closed labs run on their own systems.
“The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers—where most of the liability is,” said Balsam. “When we have the open ‘Mythos’ moment, it’s going to become clear that models need guardrails deployed at inference time.”
Goodfire’s recent research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents.
Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini.
Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behaviour can be traced back to where it emerged in training. “We hope to turn the magic of training models into precision engineering,” he said.
What it means
Developers of open models now have a way to add safety without paying for heavy external monitoring. The system runs alongside existing work, catching problems before they happen. This shifts the focus from post-event review to prevention during training and evaluation.




