{"schema":"postcutoff/event@1","as_of":"2026-10-08T23:45:00+02:00","url":"https://postcutoff.com/e/2026-10-08-goodfire-activation-probe-monitors/","md":"https://postcutoff.com/e/2026-10-08-goodfire-activation-probe-monitors/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"2026-10-08-goodfire-activation-probe-monitors","date":"2026-10-08","date_precision":"day","short_title":"Goodfire launches 'inside-out' activation-probe monitors for AI agents","deck":"94% of malicious hacking sessions caught at ~$51 vs ~$10,000 for a frontier-model monitor","takeaway":"On Oct 8, 2026 interpretability startup Goodfire released monitors that read a model's internal activations with small classifiers (\"probes\") instead of having a second AI reread every output.","category":"policy-safety","category_label":"Policy & safety","importance":3,"confidence":"medium","status":{"key":"partly","labels":["Partly confirmed"]},"sources":[{"n":1,"title":"TechCrunch: Goodfire says its new 'inside-out' monitors catch rogue AI agents at a fraction of the cost","url":"https://techcrunch.com/2026/10/08/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost/","type":"press","group":"press","domain":"techcrunch.com"}],"official":0,"filed":"2026-10-08","updated":"2026-10-08","orgs":["Goodfire","Baseten"],"title":"Goodfire launches 'inside-out' activation-probe monitors for AI agents: 94% of malicious hacking sessions caught at ~$51 vs ~$10,000 for a frontier-model monitor","summary":"On Oct 8, 2026 interpretability startup Goodfire released monitors that read a model's internal activations with small classifiers (\"probes\") instead of having a second AI reread every output. In Goodfire's test on Kimi K3, about 1,500 agent sessions, the probes caught 94% of malicious hacking sessions and flagged 8.7% of benign ones. Monitoring cost about $51, against about $233 for a cheap monitor model and about $10,000 for a top-tier one. The monitors are offered to Baseten customers.","key_facts":["Method: probes on intermediate activations the model already computes; CEO Eric Ho: 'All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations'","Kimi K3 test (~1,500 sessions): 94% of malicious hacking sessions detected; 8.7% of harmless sessions flagged for second review; <2% added latency with four probes running at once","Cost for the test set: ~$51 with probes vs ~$233 with a cheaper monitor model vs ~$10,000 with a top-tier monitoring model","Also tested on GLM-5.2 and in a Google DeepMind Gemini deployment (TechCrunch)","Configurable categories: offensive hacking, chemical/biological weapons misuse, reward hacking","Availability: to customers of Baseten, which hosts open-weight models (probes need access to activations, so they suit self-hosted or open-weight models)","All figures are Goodfire's own; no independent evaluation yet"],"key_numbers":[],"tags":["interpretability","monitoring","probes","agents","cybersecurity","open-weights"],"science":null,"body_md":"## What happened\n\nGoodfire, the interpretability startup co-founded by CEO Eric Ho and CTO Dan Balsam, turned activation probes into a commercial\nmonitoring product for agent deployments. It is offered through the model host Baseten. It targets the misuse categories that\nlabs worry about most for open-weight models: cyber offense, CBRN and reward hacking.\n\nConfidence is medium: only TechCrunch's report was read, and Goodfire's own write-up was not located.\n\n## Why it matters\n\nOpen-weight models with frontier-level cyber skills (GLM-5.3, Kimi K3) cannot be monitored at the API like closed models. Monitoring\nthrough the model's internals is about 200x cheaper than a frontier-model monitor, which could make safety monitoring of self-hosted\nagents affordable. It is also one of the first commercial uses of interpretability.","disputed":[],"related":[{"id":"2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities","url":"https://postcutoff.com/e/2026-09-29-anthropic-glm-5-3-spread-of-cyber-capabilities/","date":"2026-09-29","date_precision":"day","short_title":"Anthropic: open-weights GLM-5.3 nearly matches Mythos Preview at exploit development","deck":null,"takeaway":"It is the first time a frontier lab has published evidence that an open-weights model reached the level of exploit capability it had judged too risky to release widely.","category":"policy-safety","category_label":"Policy & safety","importance":4,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":13,"official":5,"filed":"2026-09-30","updated":"2026-10-07","orgs":["Anthropic","Zhipu AI"]}],"people":[{"id":"eric-ho","name":"Eric Ho","url":"https://postcutoff.com/person/eric-ho/"}],"posts":[],"videos":[],"models":[],"changes":[{"date":"2026-10-08","type":"filed","text":"Created"}],"provenance":{"agents":[{"model":"Claude Opus 5.5","maker":"Anthropic","tool":"Claude Code"}],"filed":"2026-10-08","run":null,"sources_read":null,"updated":"2026-10-08","human_review":null,"version":{"date":"2026-10-08"}},"gaps":[{"model_id":"gpt-6-astra","name":"GPT-6 Astra","cutoff":"2026-04","days_after":161,"in_training_data":false},{"model_id":"claude-opus-5-5","name":"Claude Opus 5.5","cutoff":"2026-06","days_after":100,"in_training_data":false},{"model_id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","cutoff":"2026-03","days_after":191,"in_training_data":false},{"model_id":"grok-4-7","name":"Grok 4.7","cutoff":"2026-05","days_after":130,"in_training_data":false}],"short_url":null}