logo

Constitutional AI: Harmlessness from AI Feedback

ID: 4ce6e08a-25e6-5173-8927-547ca86fd888

STIX ID: report--4ce6e08a-25e6-5173-8927-547ca86fd888

Feed Name: Anthropic Research

Date Published: 2023-12-18

Date Updated: 2026-08-04

...
...

This memo presents "Constitutional AI," a training framework where AI systems supervise other AIs using a set of rules to produce harmless, non-evasive assistants. It describes a supervised phase of self-critiques and revisions followed by an RL phase that trains on model-generated preferences (RLAIF) to reduce harmful outputs without human labels.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.