logo

Evaluating the Effectiveness of Reward Modeling of Generative AI Systems

ID: b27021d5-357c-55d1-bd16-4dfae27f1010

STIX ID: report--b27021d5-357c-55d1-bd16-4dfae27f1010

Feed Name: Schneier on Security

Date Published: 2024-09-11

Date Updated: 2026-04-19

Author: Bruce Schneier

...
...

This post summarizes SEAL, a study evaluating reward modeling in RLHF, introducing metrics (feature imprint, alignment resistance, alignment robustness) and reporting experiments with Anthropic preference data and OpenAssistant reward models that show strong target-feature imprints, sensitivity to spoiler features, ~26% alignment resistance where labelers disagreed with humans, and misalignment often due to ambiguous data—highlighting the need to scrutinize both reward models and alignment datasets.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.