logo

Claude 4 benchmarks show improvements, but context is still 200K

ID: 690b708a-445a-5be4-9347-125fb3072849

STIX ID: report--690b708a-445a-5be4-9347-125fb3072849

Feed Name: Bleeping Computer

Date Published: 2025-05-22

Date Updated: 2026-04-20

Author: Mayank Parmar

...
...

Anthropic introduced Claude 4 (Opus and Sonnet), emphasizing strong coding and complex task performance in benchmarks (e.g., 72.5% on SWE-bench, 43.2 on Terminal-bench) while retaining a 200K-token context window that trails Gemini 2.5 Pro and ChatGPT 4.1 offerings up to 1M tokens; the piece also lists pricing (including prompt caching and batch discounts) and includes a brief promotion for an IAM strategy guide.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.