A General Language Assistant as a Laboratory for Alignment
ID: 5f4f4ee7-59fa-5f33-bb43-92509e368ca9
STIX ID: report--5f4f4ee7-59fa-5f33-bb43-92509e368ca9
Feed Name: Anthropic Research
This abstract summarizes research into aligning large language models by evaluating baseline techniques like prompting, comparing training objectives (imitation learning, binary discrimination, ranked preference modeling), finding that ranked preference modeling outperforms and scales more favorably than imitation learning or binary discrimination, and exploring a 'preference model pre-training' stage to improve sample efficiency when finetuning on human preferences.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
