Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2022
  4. Anthropic introduces Constitutional AI (RLAIF)

Anthropic introduces Constitutional AI (RLAIF)

★★★★policy-safetyAnthropicconfidence: high

Anthropic's Constitutional AI trained a harmless-but-helpful assistant using AI feedback guided by a written set of principles (a 'constitution') instead of human harm labels.

Key facts

What happened

The model critiqued and revised its own outputs according to principles, and a preference model trained on AI judgments then guided RL.

Why it matters

Showed alignment could scale with AI supervision, making values explicit and auditable; RLAIF is now widespread.

Changelog

  • 2026-09-29: created

Related events

  1. InstructGPT: RLHF aligns language models to follow instructions ★★★★★
  2. Anthropic releases Claude ★★★★
  3. Anthropic launches with a focus on AI safety ★★★

Sources (2)

id: 2022-12-15-constitutional-ai · updated 2026-09-29 · open in the interactive timeline