I Tested NEW Sonnet 5 with 25 Coding Prompts
AI Coding Daily · 2026-07-31 · review · 6,950 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Povilas Korop from AI Coding Daily tests Anthropic’s Claude Sonnet 5 on his 5-project, 25-prompt LLM coding benchmark suite. He evaluates the model across React, Laravel API, Fluent Validation, Filament Admin, and CSV import tasks, comparing its performance and execution costs directly against Claude Sonnet 4.6 and other frontier models.
What is shown
- [00:06] The initial LLM Coding Leaderboard before adding Sonnet 5, showing Claude Opus 4.8 at #1 (24.5/25) and Sonnet 4.6 at #11 (16.4/25, $0.49/prompt).
- [01:15] Anthropic announcement tweet regarding the redeployment of Claude Fable 5 with updated cybersecurity classifiers.
- [01:47] Evaluation of Project 1 (React & TypeScript): Sonnet 5 scores a perfect 5/5.
- [02:28] Evaluation of Project 2 (Laravel API): Sonnet 5 achieves 4/5, failing 1 of 5 attempts due to incorrect product ordering (22/23 tests passed, $0.83 average run cost).
- [03:41] Evaluation of Project 3 (Laravel Fluent Validation): Sonnet 5 scores 3/5, failing 2 attempts on syntax/parameter mismatches and N+1 query assertions.
- [04:50] Evaluation of Project 4 (Filament Admin Panel with PHP Enums): Sonnet 5 scores 0/5 (down from Sonnet 4.6's 3/5). At [06:11], Povilas reproduces the bug in the browser UI, revealing an unhandled
MassAssignmentExceptionbecause Sonnet 5 forgot to define$fillableproperties on the Eloquent model and failed to generate automated tests to catch it. - [08:44] Evaluation of Project 5 (Harden Contact CSV Importer): After passing the first two runs (29/29 and 28/29 tests), subsequent runs crash at [09:12] because the account hit Anthropic's 5-hour usage limit on the $20/month subscription ([09:24]).
- [10:23] Setup and purchase of extra usage credits (€5 minimum) on Claude.ai to finish the remaining runs.
- [11:40] Resumed CSV Importer runs, scoring 29/29 on all three final runs, yielding an overall 4.5/5 score for Project 5.
- [12:31] The updated LLM Coding Leaderboard placing Sonnet 5 (Medium) at #11 with 16.5/25 total points, an average execution time of 2:01, and an average prompt cost of $0.72.
- [13:07] Anthropic's pricing announcement page showing introductory rates of $2/M input and $10/M output through August 31, 2026, rising to $3/$15 in September 2026.
- [13:37] Community reactions and benchmark comparisons on X criticizing Sonnet 5's cost-to-performance ratio for coding tasks.
Claims & numbers
- Presenter's benchmark results for Claude Sonnet 5 (Medium effort):
- Total score: 16.5 out of 25 maximum points across 5 projects (scoring 5 in React, 4 in Laravel API, 3 in Fluent Validation, 0 in Filament Enum, and 4.5 in CSV Import).
- Score comparison: Marginally higher than Sonnet 4.6 (16.4/25) but significantly behind Opus 4.8 (24.5/25) and Chinese open/proprietary models like GLM-5.2 (17.7/25) and MiniMax M3 (18.5/25).
- Speed and cost: Average execution time was 2:01 per prompt; average cost was $0.72 per prompt (a 47% increase compared to Sonnet 4.6 at $0.49, nearing Opus 4.8 at $0.74).
- Usage limits: The presenter exhausted 100% of his 5-hour paid usage session on the $20/month plan after executing only 22 agentic prompts ([10:14]).
- Anthropic pricing: Sonnet 5's introductory token pricing is $2 per million input tokens and $10 per million output tokens through August 31, 2026, after which it increases to standard pricing of $3 input and $15 output per million tokens ([13:17]).
Notable quotes
- [01:06] "Sonnet 5 results kind of confused me: why did they even release that model in the first place?"
- [08:08] "And this is the classical example of models saying to you 'everything works' where it doesn't work."
- [15:02] "So I would not recommend using Sonnet for coding in basically any shape or form."
Assessment
This is an independent benchmark review demonstrating live terminal test runs, browser reproductions of runtime failures, and real account billing interfaces. Everything presented is supported by transparent automated test suites, execution logs, and live code inspection without deceptive cuts or unverified hype.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.