Google study finds LLMs are embedded at every stage of abuse detection
ID: 11a34eac-3984-5bea-8774-109930103d3b
STIX ID: report--11a34eac-3984-5bea-8774-109930103d3b
Feed Name: Help Net Security
This survey examines the use of large language models at each stage of the content-moderation lifecycle—labeling, detection, review/appeals, and auditing—highlighting gains in contextual reasoning and scale alongside new biases, over-refusal and false positives, unfaithful explanations, demographic disparities, and the impracticality of running large models at full scale; it recommends hybrid architectures combining smaller specialists, retrieval-augmented policies, continuous red-teaming, and sustained human oversight.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
