Claude Opus 5 is a freak
AI Search · 2026-07-31 · community · 690,985 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This video is a comprehensive review and benchmark critique of Anthropic’s Claude Opus 5 model, presented by the tech channel AI Search. The creator tests Opus 5’s agentic and vibe-coding capabilities across full-stack browser application design, 3D asset generation, motion graphics video production, DAW music production, visual object detection, and biomedical reasoning, while comparing its real-world performance, speed, and cost against frontier models like GPT-5.6, Claude Fable 5, and Kimi K3.
What is shown
- Introduction & Overview [00:00 - 00:56]: Introduction of Anthropic’s Claude Opus 5 announcement page (dated July 24, 2026), its positioning within the Claude model lineup, and its intended deployment in autonomous agentic coding frameworks like Claude Code.
- Windows 11 Browser Replica [00:57 - 04:55]: A single prompt in Claude Code asking Opus 5 to create a functional web-based Windows 11 replica with working apps (Word, Excel with a formula engine, PowerPoint, Media Player, Discord/Slack simulations, and Spotify with synthesized audio). Opus 5 plans the architecture, writes multi-file JavaScript, uses a headless browser to detect errors, fixes layout bugs, and serves a fully interactive desktop environment inside Google Chrome.
- Context & Token Usage Inspection [06:50]: A review of the Claude Code terminal stats showing the Windows 11 generation consumed 366.2k tokens out of the 1M context window and took over an hour to execute.
- 3D Scene Reconstruction from 2D Reference [07:14 - 08:46]: An isometric office image prompt turned into an animated 3D Three.js HTML scene. After a critique about furniture placement and post-processing glow, Opus 5 refines camera elevation, object coordinates, and lighting to match the reference closely.
- Automated Financial Report Video Production [09:06 - 11:49]: Opus 5 autonomously web-scrapes Q4 2025 financial reports for Nvidia, Google, Meta, and Amazon, analyzes the metrics, writes a motion graphic animation using Hyperframes, generates voiceover audio using Gemini TTS, and renders a 16:9 presentation video.
- Sponsored Segment: Luma Agents & Luma Skills [11:50 - 13:55]: Demonstration of Luma AI’s multi-agent design platform, saving repeatable visual branding and runway fashion workflows into reusable "Skills."
- Blender MCP 3D Modeling & Animation [13:56 - 15:37]: Claude Code connects directly to Blender 5.2 via Model Context Protocol (localhost:9876) to programmatically model, texture, rig wing hinges, animate, and render an X-Wing fighter spaceship.
- End-to-End Music Composition in Waveform DAW [15:38 - 20:36]: Opus 5 scans the local Waveform DAW setup, searches GitHub/web for free VST plugins under 800 MB, downloads and installs the Surge XT synthesizer, arranges 18 MIDI tracks (kick, sub-bass, arpeggios, pads, risers), configures panning and automation, and renders a 5-minute melodic techno song.
- Visual Failure Cases (Camouflage & Medical CT) [21:04 - 23:08]:
- An image of leaves with a camouflaged frog is analyzed via 3x3 tile inspection; Opus 5 hallucinates a potential snake search and concludes no animal is present [21:50].
- A CT scan with 6 brain tumor slices is fed to the model; Opus 5 misclassifies or misses the lesion in all 6 slices [22:54].
- Deep Biomedical Research [23:09 - 24:11]: Opus 5 synthesizes atherosclerosis pathophysiology, creating interactive HTML/SVG flowcharts, plaque diagrams, and clinical trial tables.
- Leaderboards, Pricing & Guardrail Analysis [24:44 - 32:03]: Comparative analysis of Opus 5 across Frontier-Bench, GDPval-AA, ARC-AGI-3, LiveBench, Vals Index, DeepSWE, Artificial Analysis speed/cost charts, and safety fallback mechanisms.
Claims & numbers
- Release date: Claude Opus 5 was released by Anthropic on July 24, 2026 (the presenter shows the announcement page).
- Context window & specs: Features a 1 million token context window, capable of ingesting roughly 700,000 words or entire codebases (the presenter states).
- Pricing: The presenter states Opus 5 is priced on the API at $5 per million input tokens and $25 per million output tokens; citing the Artificial Analysis blended cost index, Opus 5 costs $2.03 per unit compared to $1.04 for GPT-5.6 Sol and $2.75 for Claude Fable 5 (with fallback).
- Execution speed: The presenter cites Artificial Analysis measuring Opus 5 at 53 output tokens per second, noticeably slower than Fable 5 (71 tps), GPT-5.6 Sol (66 tps), and open-weight models like gpt-oss-120b (273 tps).
- Benchmark results cited:
- Frontier-Bench v0.1 (terminal coding): Opus 5 scores 43.3% vs. Fable 5 at 33.7% and GPT-5.6 Sol at 34.4%.
- DeepSWE v1.1: Opus 5 achieved 74% pass@1 (average task cost $11.84), narrowly leading GPT-5.6 Sol at 73% ($8.39) and Fable 5 at 71% ($21.63), though the presenter notes confidence intervals overlap.
- LiveBench: Opus 5 ranks #3 overall at 80.3, behind Claude Fable 5 (82.0) and GPT-5.6 Sol Max Effort (82.4).
- Vals Index: Opus 5 achieves 74.82% accuracy ($8.54/test) behind Claude Fable 5 (75.14% at $11.00/test) and slightly ahead of Kimi K3 (74.70% at $2.34/test).
- ARC-AGI-3: Anthropic reports a 30.2% score for Opus 5, but the presenter cites independent testing by researcher Guanghan Ning showing Opus 5 succeeds on familiar puzzle genres (scoring 43.4 ± 3.2 on Witness-style puzzles) but regresses below Opus 4.8 on completely novel rule sets.
- Artificial Analysis Omniscience Hallucination Rate: Opus 5 scores a 50.07% hallucination rate, roughly on par with Kimi K3 (50.94%), while open models like GLM-5.2 achieve 28.13%.
- Safeguards & Fallbacks: The presenter notes Opus 5 intervenes ~85% less often on cybersecurity prompts than Fable 5, allowing source code vulnerability scanning while blocking binary exploit generation; flagged queries fall back to Claude Opus 4.8.
Notable quotes
- [10:17]: "Again, the awesome thing about Opus 5 is that it can autonomously verify its generation and then fix any errors that it sees."
- [24:41]: "I feel like it's twice as slow as Kimi K3 or GPT-5.6, which are already really slow. And also, Opus 5 is much more expensive."
- [31:04]: "In fact, in 100% of my personal workflows, I don't actually need to use Opus 5. I can just go with GPT-5.6 or Kimi K3 or even the much cheaper GLM-5.2..."
Assessment
This is an authentic, independent hands-on review and critique video. The presenter demonstrates real, unscripted model executions through Claude Code and local tool harnesses (Blender MCP and Waveform DAW), openly showing severe model failures (failing camouflage detection and medical scan diagnosis) alongside successful complex coding runs. Long-running tasks taking over an hour are appropriately fast-forwarded via timelapses.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.