AJ Learning Hub logoAJ Learning Hub

Constitutional AI

Anthropic's training method that aligns models using a written set of principles (a constitution), having the model critique and revise its own outputs against those rules rather than relying only on human labels.

Analogy

Like giving a trainee a code of conduct to self-check their work against, instead of correcting every mistake by hand.

Why it matters

It is why Claude tends to refuse harmful requests and stay on-policy, which matters when it acts on client accounts.

In practice

Claude declining to write deceptive spam because it conflicts with its guiding principles.

Related terms:GuardrailsEvalsHallucination