Claude Opus 4.7 in 5 Minutes
Developers Digest · 2026-05-02 · community · 17,497 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, the presenter from the YouTube channel Developers Digest provides an overview and breakdown of Anthropic’s Claude Opus 4.7 release. He covers the official announcement details, comparative benchmark scores across coding and reasoning evaluations, changes to file-system memory handling, and new API and Claude Code features such as task budgets and effort levels.
What is shown
- [00:00] The official Anthropic announcement page ("Introducing Claude Opus 4.7", dated April 16, 2026) and announcement post on X.
- [00:44] The benchmark comparison table highlighting Opus 4.7 versus Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Claude Mythos Preview across evaluations including SWE-bench Verified, SWE-bench Pro, Humanity's Last Exam, GPQA Diamond, and CharXiv Reasoning.
- [01:31] Pricing details and early-access testimonials from Intuit and Augment Code on the announcement blog.
- [02:22] Blog post text detailing Opus 4.7's file-system-based memory and progressive disclosure capabilities.
- [03:03] A bar chart showing SWE-bench Multilingual and Multimodal accuracy comparing Opus 4.7 to Opus 4.6.
- [03:09] An X thread from Claude detailing new developer features: the
xhighreasoning effort parameter, task budgets (beta), Claude Code/ultrareview, and expanded auto mode for Max users. - [04:23] A scatter plot of agentic coding score versus token usage across effort levels (
low,medium,high,max), demonstrating the significant token consumption increase when usingmaxeffort on Opus 4.7.
Claims & numbers
- Benchmarks:
- The presenter notes that Opus 4.7 achieves 64.3% on SWE-bench Verified (up from 53.4% on Opus 4.6).
- On SWE-bench Pro, Opus 4.7 scores 87.6% (compared to 80.8% on Opus 4.6 and 80.6% on Gemini 3.1 Pro).
- On Terminal-Bench 2.0, Opus 4.7 scores 69.4% (Opus 4.6: 65.4%; GPT-5.4: 75.1%).
- On Humanity's Last Exam, Opus 4.7 scores 46.9% without tools and 54.7% with tools (Opus 4.6: 40.0% / 53.3%).
- On CharXiv Reasoning, Opus 4.7 scores 82.1% (91.0% with zoom), compared to 69.1% (84.7% with zoom) for Opus 4.6.
- On SWE-bench Multilingual, Opus 4.7 reaches 80.5% compared to 77.8% on Opus 4.6.
- Pricing & Availability:
- The presenter states pricing remains unchanged from Opus 4.6 at $5 per million input tokens and $25 per million output tokens.
- Opus 4.7 is generally available across the API, Claude Code, web, and desktop apps.
- Model behavior and features:
- Augment Code reports the model exhibits reduced sycophancy and offers more opinionated perspectives rather than blindly agreeing with developers.
- A new
xhigheffort setting sits betweenhighandmax. - At the
maxeffort setting on agentic coding evaluations, Opus 4.7 utilizes roughly 250,000 tokens per task compared to approximately 130,000 tokens on Opus 4.6.
Notable quotes
- [01:42] "It is still going to be $5 per million tokens of input and $25 per million tokens of output."
- [02:04] "...it's actually nice when a model will disagree with you."
- [04:35] "The number of total tokens that are used are substantially higher."
Assessment
This is an independent summary and commentary video reviewing Anthropic's blog post, social media announcements, and benchmark charts. The presenter does not run independent benchmarks or live code tests during the video, relying entirely on Anthropic's published release materials and tester testimonials.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.