Public benchmarks do not measure your day-to-day work. Here is how I built eight tests for Opus 5 and Fable 5, automated the scoring, blind-reviewed the subjective task, and diagnosed the questions I designed badly.
An SEO audit found no urgent keyword problem. The real issues were missing social cards, 12 MB of cover images, inconsistent canonical URLs, and a thin discovery layer.
I ran Opus 5 and Fable 5 through the same eight tests: math, logic, coding, bug hunting, instruction following, extraction, decision writing, and agentic ETL. The final score was 52.63 to 53.25—a much smaller gap than the price suggests.