Sonnet 5.5 and Sol 6.1 move Saggar's defaults again
GPT-6.1 Sol now powers Saggar's Codex Everyday and Hard recommendations. Sonnet 5.5 has the highest score on Bug Hunt Bench at max, but its cost and runtime keep Opus 5.5 in Everyday for Claude.

We were still introducing Everyday, Quick, and Hard when two new model releases landed. The benchmark data keeps moving, so Saggar's defaults and recommendations will move with it. These results are early, but they already change the shape of the table.
Sonnet 5.5
On max, Sonnet 5.5 is the highest-scoring setting on Bug Hunt Bench. That's a sharp change for Anthropic's "balanced workhorse", a model I found poor at coding tasks in earlier testing.
The catch is price and time. Sonnet 5.5 max sits far to the right of the other top scores on the linear chart:

At high, Sonnet 5.5 looks comparable to Opus 5.5 at medium, which makes it a good Everyday candidate. It takes much longer than Opus, though, so Opus 5.5 remains Claude's Everyday default for now.
Sol 6.1
Sol 6 launched to a lot of fanfare, yet benchmarks put Luna variants ahead for most Quick or Everyday work and Astra ahead for Hard work. 6.1 changes that.
At max, GPT-6.1 Sol now looks better suited to Hard work than GPT-6 Astra. At high, it replaces GPT-5.6 Luna at high while finishing much faster. The Codex rows now look like this:
| Codex | Claude | |
|---|---|---|
| Everyday | GPT-6.1 Sol at high | Opus 5.5 |
| Quick | GPT-6 Luna | Opus 5.5 at low |
| Hard | GPT-6.1 Sol at max | Opus 5.5 at max |
Quick stays with GPT-6 Luna. You can still override any of these choices in Settings ▸ Providers, or for one session in the new-session sheet.
A benchmark, not a verdict
Obviously here im just sharing data from one of many model benchmarks, but I found these visualizations particularly clear.
I'll keep updating Saggar's defaults as new models land and the data settles. For now, Sonnet 5.5 is the expensive ceiling, while Sol 6.1 is the more useful change: it gives Codex a faster Everyday choice and a stronger Hard one.