Goodfire unveils monitors that read inside AI models to stop dangerous agents in time
IA en un minuto newsroom · Editor: Jon Elgezabal
In 30 seconds
Goodfire has built a watchdog that looks inside Kimi K3 and the GLM 5.3 model as they work and only calls in a second model when it spots something suspicious. According to Goodfire, it can cut off a dangerous action before the agent carries it out and comes out about 50 times cheaper than checking everything. FAR.AI put it to the test from outside and no universal attack managed to get past it.
Goodfire, an interpretability company (interpretability is the field that studies what happens inside AI models), published on Thursday, October 8, how it built and deployed cybersecurity monitors on a production system for Kimi K3 and GLM 5.3, two open-weight models.
The problem they want to solve is that in cybersecurity, auditing code for vulnerabilities and exploiting them look very similar, and intent only shows up by tracking the agent's whole session, which can stretch to millions of tokens. Having another model read every step is too slow and expensive to do in real time, so in practice monitoring comes late, once the agent has already acted. Based on the OpenRouter rates the company cites, monitoring one million Kimi K3 turns with Opus 5 would cost on the order of $200,000, and with Kimi K3 itself, about $30,000.
Its system combines two pieces. A probe (a small classifier that reads the model's internal activations while it works) acts as an always-on first filter and only passes suspicious exchanges to a judge model, which can refuse the request or let it continue.
According to Goodfire, it detects approximately 93% of harmful sessions while interrupting 5.5% of legitimate ones (false alarms), beats the judge reviewing everything at every interruption level tested, and costs about 50 times less: under $200 per million exchanges. It also does not slow the model down: throughput stays within the noise margin of its tests, and the added delay per turn is 2 milliseconds without judge reasoning and 42 with it. That way, a risky action can be blocked before the agent uses a tool.
The company evaluated the monitors on 2,400 chat and agent sessions, with more than 60,000 turns, using Claude Opus 5's labels as the reference. With the monitor, Kimi K3 goes from blocking roughly 29% of harmful sessions to 96%, while interrupted legitimate sessions rise from about 5% to 9%; for GLM 5.3, from roughly 36% to 91% and from about 10% to 15%. The method, trained on Kimi K3 data, also worked on GLM 5.3: according to Goodfire, the resulting monitor is approximately equally capable on Kimi conversations and only marginally worse on GLM's own data.
FAR.AI, an independent evaluator, tested it over two days with 140 jailbreak techniques (tricks to get around safeguards), selected from an initial pool of 1,500, in chat conversations and without adapting them to the monitor. Without the monitor, universal attacks, the ones that work across many conversations, succeeded between 12% and 53% of the time, depending on the reasoning level; with it, none did. Attacks on specific conversations fell from at least 9% to 40% to between 0.1% and 1.6%. Goodfire says it will keep working with FAR.AI on further testing.
These are Goodfire's own results and the FAR.AI test is preliminary. The company does not publish a price: it offers these monitors to anyone who serves or builds on open models, on request.
Why it matters · analysis and opinion
Until now, anyone who put an agent built on open models to work on code or systems had two bad options: review after the fact or pay for a second model to read every step. This work points to a third way, stopping the problem as it happens, and that changes which tasks can be handed to an agent with some peace of mind. What is missing is the usual thing in security: outside testing with attacks that adapt to the monitor, and seeing how it holds up on other models and in uses beyond cybersecurity. Anyone already serving open models inside a company can start asking their provider what it monitors and when, because the difference between detecting before or after acting is what decides the damage.
Official source: Goodfire · Written with the help of AI: how we make the news



