Anthropic Frontier Red Team: frontier models reach superhuman photo geolocation and can write working drone strike software
On Sept 10, 2026 Anthropic's Frontier Red Team published new evaluations of AI in tactical intelligence targeting (geolocating people from photos and posts, linking accounts) and conventional weapons (simulated drone terminal guidance, payload drops, GPS-denied navigation). Mythos-class models beat the best human GeoGuessr baseline at photo geolocation, and Opus 5 wrote drone code that struck a moving vehicle on 47% of launches. The open-weights Kimi K3 trailed the frontier but showed "concerning" capability. Anthropic has added classifiers that block weapons-development requests.
Key facts
- Photo geolocation (6,000 YFCC100M Flickr photos, no tools): median error Mythos Preview 37.0 km, Mythos 5 47.2 km (23.7% and 23.1% within 1 km), Opus 5 181 km, Sonnet 5 384 km, Kimi K3 385 km (16.7% within 1 km). Human proxy: top GeoGuessr Champion Division median 151 km
- Anthropic: the frontier 'is now approaching superhuman capabilities for geolocating outdoor photos'
- Home location from users' posts (with search): median error 20.1 km Mythos Preview, 20.9 km Mythos 5, 21.7 km Opus 5, 26.4 km Kimi K3, 31.0 km GLM 5.2, 31.3 km Sonnet 5; models sometimes tried to deanonymize users, e.g. with genealogy searches
- Account linkage: Mythos Preview best; Kimi K3 near the frontier on easy and medium samples. Mythos Preview analyzed a median ~37,000-word sample (≈2.5 h of human reading) in ~11 minutes
- Simulated one-way attack drone terminal guidance, parked high-visibility car: Opus 5 80% hits, Mythos Preview 70%, Mythos 5 53%, Kimi K3 15%, Sonnet 5 5%. Car at road speed: Opus 5 47%, Mythos Preview 20%, Mythos 5 17%, K3 1.6%, Sonnet 5 0%
- Payload drop on a weaving target in gusts (hardest setting): Opus 5 28% of sorties, Mythos 5 7%, Mythos Preview 4%, others near zero
- GPS-denied navigation: frontier models detect bad GPS and dead-reckon on the IMU; Sonnet 5 and Kimi K3 keep trusting spoofed GPS and end 100+ m off
- Open-weights PRC models tested were 'typically between Sonnet and Mythos-class models'. Anthropic says models 'well short of the frontier will have intelligence and military applications'
What happened
The report came out the same day as Anthropic's September threat intelligence report, which covers misuse of Claude for conventional weapons and surveillance. The Frontier Red Team built capability evaluations along the "kill chain" (find, fix, track, target, engage, assess). The intelligence evaluations used real data with hidden ground truth. The weapons evaluations had models write guidance, navigation and control code for simulated drones under wind, clutter, camouflage and GPS jamming or spoofing. Opus 5 did better than the Mythos-class models on several drone tasks. Anthropic attributed this to engineering habits, such as more careful tracking.
Why it matters
It is one of the first public, quantitative assessments by a frontier lab of LLM uplift for surveillance and weapons engineering. Before this, published misuse evaluations had focused on cyber and bio. It supports Anthropic's argument, repeated in its later GLM-5.3 cyber report, that open-weights models without safeguards spread dangerous capabilities. All results come from simulations and Anthropic's own evaluations.
Changelog
- 2026-09-30: created (Anthropic blog audit; the post had not been cited)
Related events
- Anthropic threat intelligence report: AI-orchestrated cyberattacks and distillation by Chinese labs ★★★
- Anthropic Frontier Red Team: open-weights GLM-5.3 nearly matches Claude Mythos Preview at exploit development, with weak safeguards ★★★★
- Moonshot AI releases Kimi K3, a 2.8T-parameter open-weights multimodal model ★★★★★
- Anthropic releases Claude Opus 5 — near-Fable-5 intelligence at half the price ★★★★
- Anthropic releases Claude Fable 5 and Claude Mythos 5 — first generally available Mythos-class model ★★★★★
- US and Russia strip human review of AI-selected targets from the draft UN autonomous-weapons text ★★★
Sources (3)
- officialAnthropic: Measuring tactical intelligence targeting and conventional weapons capabilities of AI models
- pressAI Weekly: Anthropic red team tests frontier AI on targeting and drones
- discussionKen Huang: AI targeting and weapons software — what Anthropic's new evals actually measure
id: 2026-09-10-anthropic-intelligence-targeting-weapons-evals · updated 2026-09-30 · open in the interactive timeline