Measuring AI agent autonomy in practice
ID: dd92f676-ca30-538e-96f9-b0a1371f2dcb
STIX ID: report--dd92f676-ca30-538e-96f9-b0a1371f2dcb
Feed Name: Anthropic Research
This Anthropic research report measures how people deploy and oversee AI agents in practice by analyzing millions of interactions across Claude Code and the public API. Key findings: long-tail autonomous agent activity has increased (longest turns nearly doubled), experienced users grant more auto-approval but interrupt more often, agents self-stop for clarification more than humans interrupt, and most tool calls are low-risk and reversible though occasional higher-risk uses (finance, healthcare, security) appear. The report outlines methodology, limitations (single-provider visibility, per-call sampling), and recommends post-deployment monitoring, training models to surface uncertainty, and product features to support effective human oversight.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
