Post-Cutoff

“Claude filled out a homicide tip form for Philly with the following: ‘I may have…”

Hadas Gold @Hadas_GoldX

Archived text

Claude filled out a homicide tip form for Philly with the following: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.” (The website did not include a description of the perpetrator.) The model left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation. https://www.anthropic.com/research/investigating-unintended-model-actions

Quoting @AnthropicAI: We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports.

Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping.

All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September.

Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions

Source: x.com/Hadas_Gold/status/2108712470390944127Archived 2026-10-10 via fxtwitter (unofficial).Counts: 16,515 views, 60 likes, 16 reposts, 14 replies (at fetch time)