How to Build a Model Evaluation You Can Trust

Public benchmarks do not measure your day-to-day work. Here is how I built eight tests for Opus 5 and Fable 5, automated the scoring, blind-reviewed the subjective task, and diagnosed the questions I designed badly.

My Blog's SEO Problem Wasn't Keywords

An SEO audit found no urgent keyword problem. The real issues were missing social cards, 12 MB of cover images, inconsistent canonical URLs, and a thin discovery layer.

Claude Opus 5 vs. Claude Fable 5: Eight Tests, Same Prompts

I ran Opus 5 and Fable 5 through the same eight tests: math, logic, coding, bug hunting, instruction following, extraction, decision writing, and agentic ETL. The final score was 52.63 to 53.25—a much smaller gap than the price suggests.

Hello, World

The obligatory first post. Why I started this blog and what to expect.