Collective Constitutional AI: Aligning a Language Model with Public Input
ID: 59f8df90-b8ae-57ff-ba0b-eac4a99aa719
STIX ID: report--59f8df90-b8ae-57ff-ba0b-eac4a99aa719
Feed Name: Anthropic Research
This report describes Anthropic and the Collective Intelligence Project’s experiment engaging ~1,000 U.S. participants via the Polis platform to draft a public constitution for AI, how the public-sourced principles were translated into Constitutional AI training, the training and evaluation of two Claude Instant-sized models (Public vs Standard), key findings (comparable helpfulness/harmlessness, modestly reduced bias in the Public model), and lessons learned about participant selection, moderation, principle mapping, prompt database alignment, and evaluation challenges.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
