Constitutional AI
Anthropic's training method that aligns models using a written set of principles (a constitution), having the model critique and revise its own outputs against those rules rather than relying only on human labels.
Analogy
Like giving a trainee a code of conduct to self-check their work against, instead of correcting every mistake by hand.
Why it matters
It is why Claude tends to refuse harmful requests and stay on-policy, which matters when it acts on client accounts.
In practice
Claude declining to write deceptive spam because it conflicts with its guiding principles.
