Specific versus General Principles for Constitutional AI
ID: eeea98bf-6968-505b-8064-59d4c182049a
STIX ID: report--eeea98bf-6968-505b-8064-59d4c182049a
Feed Name: Anthropic Research
This abstract presents Constitutional AI, an approach replacing human feedback with AI-model feedback guided by written principles to prevent harmful or self-interested behaviors in conversational models. The authors test whether a single general principle (roughly "do what's best for humanity") can produce harmless assistants and find that large dialogue models can generalize from this short constitution, though more detailed constitutions help control specific harms.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
