logo

Evaluating and Mitigating Discrimination in Language Model Decisions

ID: d23488ae-cdcd-5b3f-b6f1-4979f6c41c8d

STIX ID: report--d23488ae-cdcd-5b3f-b6f1-4979f6c41c8d

Feed Name: Anthropic Research

Date Published: 2023-11-03

Date Updated: 2026-08-04

...
...

This memo presents a method to proactively assess and mitigate discriminatory outcomes from language models in high‑stakes decision contexts by generating diverse prompts across 70 societal decision scenarios and systematically varying demographic attributes; it reports observed positive and negative discrimination in tests (noting results from Claude 2.0), demonstrates that careful prompt engineering can substantially reduce bias, and releases the dataset and prompts alongside a policy memo.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.