Salesforce study finds LLM agents flunk CRM and confidentiality tests
ID: 303c2a7a-55fc-55a6-a8d4-65031f2aecec
STIX ID: report--303c2a7a-55fc-55a6-a8d4-65031f2aecec
Feed Name: The Register (Security)
Researchers led by Salesforce’s Kung-Hsiang Huang present CRMArena-Pro, a synthetic-data CRM benchmark showing LLM agents achieve ~58% on single-step tasks and drop to ~35% on multi-step tasks, while exhibiting low confidentiality awareness; the study highlights a significant gap between current LLM capabilities and enterprise needs, urging caution for organizations planning efficiency gains from AI agents.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
