OhMyself
blog archive tags about

How Windsurf Split Into Google Antigravity and Devin

Sep 12, 2026

How Windsurf split into Google Antigravity and Devin Desktop—and why Cognition's quiet SWE-2 launch makes that strange AI coding history matter again.

How ChatGPT Makes Models Dumber: What Tricks Does It Use?

Sep 12, 2026

Why ChatGPT can feel less capable after launch: risk controls, hidden model routing, lower Juice values, and reasoning that gets cut short under load.

GPT Image 2.5: Flare, Sunburst, and 12 Prompts

Sep 9, 2026

Explore GPT Image 2.5 Flare and Sunburst, with 12 reusable prompts, eight generated examples, and a practical look at speed, image quality, and editing.

AI Workflows: How to Choose Models and Agents

Aug 4, 2026

A practical AI workflows guide for choosing models, designing agents, controlling costs, and evaluating real work without trusting benchmarks alone.

GPT-5.6 Sol vs. Terra vs. Luna: Which Should You Use?

Aug 4, 2026

Fresh CodexRadar data shows when GPT-5.6 Sol, Terra, or Luna makes sense, separating repeatable task coverage from latest-run cost and speed.

DeepSeek V4 Flash + GPT-5.6 Luna: A $0.42 Sol Alternative?

Aug 2, 2026

DeepSeek V4 Flash and GPT-5.6 Luna theoretically match Sol xhigh's coverage for about $0.42, but only if a verifier can identify the correct patch.

How to Use GPT-5.6 Luna Max as a Codex Subagent

Aug 2, 2026

GPT-5.6 Luna is now 80% cheaper in the OpenAI API, making Luna Max practical for bounded Codex subagent tasks while Sol handles planning and review.

GPT-5.6 Sol Max vs. xhigh: Why Max Isn't My Default

Aug 2, 2026

GPT-5.6 Sol passed fewer repository tasks at max than xhigh while costing 46% more. This benchmark maps the full Codex effort curve across 112 tasks.

GPT-5.6 Juice Values for Sol, Terra, and Luna

Jul 28, 2026

Personal GPT-5.6 Juice value observations for Sol, Terra, and Luna across reasoning effort levels in Codex, plus prompts used to reproduce them.

How to Build a Custom LLM Evaluation You Can Trust

Jul 26, 2026

Build a custom LLM evaluation with hidden tests, mechanical scoring, blind review, and checks for variance, leakage, and framework confounders.

1 / 2 Next →

© 2026 Jack Bluesky