Constitutional AI: Harmlessness from AI Feedback
ID: 4ce6e08a-25e6-5173-8927-547ca86fd888
STIX ID: report--4ce6e08a-25e6-5173-8927-547ca86fd888
Feed Name: Anthropic Research
This memo presents "Constitutional AI," a training framework where AI systems supervise other AIs using a set of rules to produce harmless, non-evasive assistants. It describes a supervised phase of self-critiques and revisions followed by an RL phase that trains on model-generated preferences (RLAIF) to reduce harmful outputs without human labels.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
