logo

Salesforce study finds LLM agents flunk CRM and confidentiality tests

ID: 303c2a7a-55fc-55a6-a8d4-65031f2aecec

STIX ID: report--303c2a7a-55fc-55a6-a8d4-65031f2aecec

Feed Name: The Register (Security)

Date Published: 2025-06-16

Date Updated: 2026-04-26

Author: Lindsay Clark

...
...

Researchers led by Salesforce’s Kung-Hsiang Huang present CRMArena-Pro, a synthetic-data CRM benchmark showing LLM agents achieve ~58% on single-step tasks and drop to ~35% on multi-step tasks, while exhibiting low confidentiality awareness; the study highlights a significant gap between current LLM capabilities and enterprise needs, urging caution for organizations planning efficiency gains from AI agents.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.