Modular pretraining (GRAM) lets dangerous capabilities be switched off per module (AE Studio and Anthropic)
On July 8, 2026 researchers from AE Studio and Anthropic published "Modular Pretraining Enables Access Control". It introduces Gradient Routed Auxiliary Modules (GRAM), which route dual-use knowledge such as virology, cybersecurity and nuclear physics into separate modules during pretraining. The modules can be switched on or off, so one training run matches several data-filtered models.
Key facts
- 800M-parameter model with four dual-use domains: GRAM and LoRA matched data filtering at capability removal and beat it on retained domains
- Capability removal improved with scale across 50M–5B parameters
- Four modules give 16 configurations from one training run, vs one run per configuration for data filtering
- With only 50% of data labelled, GRAM isolated capabilities better than filtering or LoRA
- Authors include Cem Anil and Alex Cloud (Anthropic) and Judd Rosenblatt (AE Studio)
What happened
An extension of gradient routing. Dangerous knowledge is placed in removable modules during pretraining instead of being filtered out of the data.
Why it matters
It points to tiered access, where vetted users get a model with, say, the virology module and the public does not, without training separate models. This bears on how labs might ship open or differentiated weights safely.
Changelog
- 2026-10-01: created (leads run, from the Anthropic uncited-posts audit)
Sources (1)
id: 2026-07-08-anthropic-ae-studio-modular-pretraining-gram · updated 2026-10-01 · open in the interactive timeline