Goodfire launches ‘inside-out’ activation-probe monitors for AI agents
94% of malicious hacking sessions caught at ~$51 vs ~$10,000 for a frontier-model monitor
Partly confirmed
The takeaway
On Oct 8, 2026 interpretability startup Goodfire released monitors that read a model’s internal activations with small classifiers (“probes”) instead of having a second AI reread every output.
Status
- Claim
Partly confirmed
- Our reporting
- Medium confidence
- Importance
- 3 of 5
- Last verified
- 8 October 2026
Your AI and this story
- GPT-6 Astra161 days after its cutoff
- Claude Opus 5.5100 days after its cutoff
- Gemini 3.8 Flash191 days after its cutoff
- Grok 4.7130 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 100 days before it.
Key facts
- Method: probes on intermediate activations the model already computes; CEO Eric Ho: ‘All you’re doing is taking the intermediate neural activations that it’s already computed and then running a classifier over these internal computations’
- Kimi K3 test (~1,500 sessions): 94% of malicious hacking sessions detected; 8.7% of harmless sessions flagged for second review; <2% added latency with four probes running at once
- Cost for the test set: ~$51 with probes vs ~$233 with a cheaper monitor model vs ~$10,000 with a top-tier monitoring model
- Also tested on GLM-5.2 and in a Google DeepMind Gemini deployment (TechCrunch)
- Configurable categories: offensive hacking, chemical/biological weapons misuse, reward hacking
- Availability: to customers of Baseten, which hosts open-weight models (probes need access to activations, so they suit self-hosted or open-weight models)
- All figures are Goodfire’s own; no independent evaluation yet
What happened
Goodfire, the interpretability startup co-founded by CEO Eric Ho and CTO Dan Balsam, turned activation probes into a commercial monitoring product for agent deployments. It is offered through the model host Baseten. It targets the misuse categories that labs worry about most for open-weight models: cyber offense, CBRN and reward hacking.
Confidence is medium: only TechCrunch’s report was read, and Goodfire’s own write-up was not located.
Why it matters
Open-weight models with frontier-level cyber skills (GLM-5.3, Kimi K3) cannot be monitored at the API like closed models. Monitoring through the model’s internals is about 200x cheaper than a frontier-model monitor, which could make safety monitoring of self-hosted agents affordable. It is also one of the first commercial uses of interpretability.
Sources
1 source from 1 site. Numbers match the chips in the text.
1 source: 1 press
Press
Changes
- Filed