logo

A General Language Assistant as a Laboratory for Alignment

ID: 5f4f4ee7-59fa-5f33-bb43-92509e368ca9

STIX ID: report--5f4f4ee7-59fa-5f33-bb43-92509e368ca9

Feed Name: Anthropic Research

Date Published: 2023-12-18

Date Updated: 2026-08-04

...
...

This abstract summarizes research into aligning large language models by evaluating baseline techniques like prompting, comparing training objectives (imitation learning, binary discrimination, ranked preference modeling), finding that ranked preference modeling outperforms and scales more favorably than imitation learning or binary discrimination, and exploring a 'preference model pre-training' stage to improve sample efficiency when finetuning on human preferences.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.