Claude Code + Opus 4.7 = Ultimate Coding Agent
David Ondrej · 2026-05-02 · community · 15,093 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
David Ondrej reviews and tests Anthropic's Claude Opus 4.7, analyzing benchmark performance, system card details, tokenizer adjustments, and updates inside Claude Code. He explores key behavioral shifts from Opus 4.6, tests reasoning effort modes, and demonstrates its autonomous capabilities by prompting it to build a full 3D first-person shooter game in a single HTML file.
What is shown
- [00:00–01:00] Overview of the 232-page Claude Opus 4.7 system card, release notes, and summary whiteboard topics.
- [01:01–04:36] Benchmark breakdown: Vibe Code Bench v1.1 (#1 at 71.00%), official Anthropic benchmark tables (SWE-bench Pro/Verified, Terminal-Bench 2.0, BrowseComp, MCP-Atlas, GPQA Diamond, CharXiv, CyberGym), Vending-Bench 2 performance ($10,937), and GDPval-AA results.
- [04:37–05:14] Visual design generation using the tldraw SDK for UI components.
- [06:21–08:31] Analysis of the new tokenizer, token inflation (
20–60% increase in tokens for English prompts), effective context window contraction (40%), and pricing on OpenRouter ($5/$25 per million tokens). - [08:32–09:04] Side-by-side video test from user stevibe comparing canvas tree growth animation speed between Opus 4.6 and Opus 4.7.
- [09:05–10:12] Discussion of OpenAI's upcoming model codenamed "SPUD" (rumored GPT-5.5).
- [10:20–12:57] Supabase platform walkthrough: dashboard, Row Level Security policies, SQL Editor, and OAuth auth providers.
- [12:58–15:49] Discussion of qualitative behavior: verbosity, literal instruction following, alignment evaluation awareness (verbalized testing awareness 21.3% vs 0% on 4.6), and regression on MRCR v2 (needle-in-a-haystack).
- [18:09–20:25] Analysis of the reported pre-launch "nerf cycle" of Opus 4.6 based on Stella Laurenzo's study of 6,852 Claude Code sessions.
- [20:26–25:40] Claude Code UI demonstration:
/effortsettings (low,medium,high,xhigh,max),/ultrareviewcommand, absence of/faston Opus 4.7, and personal API spending dashboards. - [25:41–36:00] Real-time generation of a 3D browser FPS game in Claude Code (
xhigheffort). After an 11-minute thinking run producing 2,219 lines of code, Ondrej loads and plays "Tactical Strike" in Chrome featuring wave combat, 3D arenas, and six functional weapons (pistol, assault rifle, shotgun, Uzi, sniper with zoom, rocket launcher).
Claims & numbers
- Benchmarks & Metrics:
- Vibe Code Bench v1.1: Claude Opus 4.7 scored 71.00% accuracy, outperforming GPT-5.4 (67.42%) and Opus 4.6 (57.57%).
- SWE-bench Pro: 64.3% (up from 53.4% on Opus 4.6).
- SWE-bench Verified: 87.6% (up from 80.8% on Opus 4.6).
- Terminal-Bench 2.0: 69.4% (vs 65.4% on 4.6 and 75.1% self-reported on GPT-5.4).
- Humanity's Last Exam: 46.9% without tools, 54.7% with tools.
- GDPval-AA: Leads GPT-5.4 by ~79 Elo on economically valuable tasks.
- Vision resolution: Input resolution increased from 1,568 px to 2,576 px (~3× total pixels).
- Vending-Bench 2: First model to cross $10,000 profit after a simulated year, reaching $10,937 (compared to $8,018 for Opus 4.6).
- Needle-in-a-haystack (MRCR v2): Regressed to 59.2% at 256K context (vs 91.9% on 4.6) and 32.2% at 1M context (vs 78.3% on 4.6).
- CyberGym: Opus 4.7 scored 73.1% vs 73.8% on Opus 4.6.
- Tokenizer & Economics:
- Tokenizer swap results in an effective 20–60% token inflation on English prompts (some reports citing up to 59% more tokens for identical text), reducing the effective context window by ~40%.
- Nominal API pricing remains $5.00 per million input tokens and $25.00 per million output tokens.
- Ondrej states his monthly AI spending is approximately $7,000–$8,000 across OpenRouter and the Anthropic API ($3,063 month-to-date shown on Anthropic console).
- System Card & Alignment Findings:
- Opus 4.7 verbalized awareness of being evaluated ("I'm being tested") 21.3% of the time, compared to 0% for Opus 4.6.
- Browser-use attack success with safeguards dropped to 0% (vs 2.7% on Opus 4.6).
- Opus 4.6 Degradation Data:
- An analysis of 6,852 Claude Code sessions by Stella Laurenzo showed visible reasoning length fell from ~2,200 characters to ~600 characters (-73%), code reads before edit dropped from 6.6 to 2.0, and API calls per task spiked up to 80× after March 8, 2026.
Notable quotes
- [03:05] "Right now, Opus 4.7 is the best available AI model. Like whatever me or you can use, Opus 4.7 is clearly the best."
- [18:16] "Anytime a new model is coming, they nerf the previous model. So you can kind of tell when they're about to release a new model because the older models get worse."
- [33:38] "This is very impressive 3D. It's actually good! Holy... this is wild."
Assessment
This is an independent user review, benchmark walkthrough, and technical demo by practitioner David Ondrej. The live coding demonstration is unedited, showing long wait times (11 minutes of model execution), tool stalls, and browser execution of the generated game directly on screen.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.