AI Workflows: How to Choose Models and Agents
A practical AI workflows guide for choosing models, designing agents, controlling costs, and evaluating real work without trusting benchmarks alone.
7 posts
A practical AI workflows guide for choosing models, designing agents, controlling costs, and evaluating real work without trusting benchmarks alone.
Fresh CodexRadar data shows when GPT-5.6 Sol, Terra, or Luna makes sense, separating repeatable task coverage from latest-run cost and speed.
DeepSeek V4 Flash and GPT-5.6 Luna theoretically match Sol xhigh's coverage for about $0.42, but only if a verifier can identify the correct patch.
GPT-5.6 Sol passed fewer repository tasks at max than xhigh while costing 46% more. This benchmark maps the full Codex effort curve across 112 tasks.
Personal GPT-5.6 Juice value observations for Sol, Terra, and Luna across reasoning effort levels in Codex, plus prompts used to reproduce them.
Build a custom LLM evaluation with hidden tests, mechanical scoring, blind review, and checks for variance, leakage, and framework confounders.
Claude Opus 5 and Claude Fable 5 tied on most of eight tests. Fable won the decision memo, while Opus delivered the same objective score at half the price.