Skip to main content

Picking a model and effort for Codex and Claude

Saggar now starts each Codex and Claude model on its own default effort, chosen for the best balance of results, cost, and time, rather than the provider's blanket default.

ModelSaggar's default
GPT-6 Astraxhigh
GPT-6 Lunamax
GPT-5.6 Lunamax
GPT-6 Solhigh
GPT-5.6 Solmedium
Opus 5.5medium
Fable 5.1max
Sonnet 5high

An effort you pick yourself still wins. Set one per model in the Default effort column on the provider's card in Settings ▸ Providers, or for a single session in the new-session sheet. If a Codex model doesn't offer the default effort, Saggar falls back to the one Codex publishes. Saggar's default is passed on the command line, so it also takes precedence over an effort set in Codex's config.toml or Claude Code's own settings.

Everyday, Quick, and Hard​

You don't need to pick from the list at all. The new-session menu now offers each provider as three grades of task, and every model is still under All models:

CodexClaude
EverydayGPT-5.6 LunaOpus 5.5
QuickGPT-6 LunaOpus 5.5 at low
HardGPT-6 AstraOpus 5.5 at max

A tier opens its model on the effort in the table above, unless it names one, and an effort you set in Settings still wins. The rest of this post is why each of those picks earned its place.

Where the defaults come from​

We based them on Bug Hunt Bench, Pawel Huryn's benchmark that hides 105 real bugs in two production codebases. Each model gets one round per repo in its own vendor's CLI, and a judge grades every diff blind against a withheld answer key. The figures below are rough, read from the board in late September. Opus 5.5 was run three times at each effort; every other figure comes from a single run. Rows get re-run, so check the live board before quoting one.

Codex​

  • GPT-6 Astra at xhigh finds about 43 bugs for about $24. That's roughly 8 more than high for about $3 more. Max adds about 2 more bugs for about $9, which is within noise.
  • GPT-6 Luna at max finds about 18 bugs for about 52¢. Stepping up from xhigh adds 4 bugs, which could be noise, but it costs only about 13¢ more, so max is the default.
  • GPT-5.6 Luna at max finds about 31 bugs for about $3.20. That's level with GPT-6 Sol at max, at a third of the cost, and ahead of every other Sol setting. Below max, GPT-6 Luna beats it.
  • GPT-6 Sol at high finds about 20 bugs for about $4. At every effort, another model does as well or better for less.
  • GPT-5.6 Sol at medium finds about 29 bugs for about $16. The gaps between its efforts are small, so this default is a lighter call than the others. Astra at max scores higher at about a third of the cost of 5.6 Sol at max.

Claude​

  • Opus 5.5 at medium finds about 30 bugs for about $16. That's where extra effort stops paying: low to medium adds 8 bugs for about $7, but medium to high adds only 2, which is a tie. Xhigh finds about 6 more than medium for about $19 more and 25 more minutes, so step up when a bug is worth that.
  • Fable 5.1 at max finds about 43 bugs for about $87. Below max, a cheaper Opus 5.5 effort matches or beats every Fable effort, so max is the only reason to pick Fable.
  • Sonnet 5 at high finds about 9 bugs for about $17. At max it found the same 9 for about $27 in twice the time.

Which model to reach for​

For Codex:

  1. GPT-6 Astra has the highest ceiling. Use it for the work that matters most.
  2. GPT-6 Luna is the best choice under $1. Use it for quick, cheap tasks and parallel side work.
  3. GPT-5.6 Luna is strong value at max effort, and only there. It's the Everyday pick for Codex, and no longer hidden.
  4. GPT-6 Sol never wins outright, but it's cheap enough to be a fallback.
  5. GPT-5.6 Sol costs the most and is never the best choice.

For Claude:

  1. Opus 5.5 is the best Claude model at every level of spend. At max it finds about 42 bugs, level with Fable 5.1 at max, for about two-thirds of the cost.
  2. Fable 5.1 is worth it only at max, when you want its top score and cost doesn't matter.
  3. Sonnet 5 is hidden by default now. Opus 5.5 at low finds more than twice as many bugs for less money in a third of the time. Bring it back from the hidden models on the Claude card in Settings ▸ Providers.

Across both providers, GPT-6 Astra at xhigh finds as many bugs as Opus 5.5 or Fable 5.1 at max for well under half the cost.

Higher effort is slower​

Within a model, more effort costs more tokens and more time. On the September board, Astra took about 28 minutes at medium, 40 at high, 59 at xhigh, and 79 at max. Opus 5.5 took about 12 minutes at low, 17 at medium, 24 at high, 43 at xhigh, and 67 at max. Opus 5.5 at low is the fastest setting on the board, at about 22 bugs for $8. When you're waiting on the answer, try it, Astra at medium, or GPT-6 Luna at high first. GPT-5.6 Luna at max is cheap but slow, at about 86 minutes. GPT-5.6 Sol at max took 164.

On a Claude or ChatGPT subscription, a heavier effort costs you usage limits rather than dollars, but it's slower either way.

A benchmark, not a verdict​

Bug Hunt Bench is a single-author benchmark whose repos and answer keys are private, which keeps them out of training data but means no one can reproduce a run independently. Most settings are scored from one run. Repeat runs of Grok 4.5 scored 16, 13, and 17, and Opus 5.5's middle efforts varied by one or two bugs between runs, so treat gaps smaller than five as ties. Fable 5.1's single runs are noisy enough that medium scored below low, and xhigh below high. Costs mix real bills with list-price estimates, so comparing efforts within one model is safer than comparing models. Your own codebase is the benchmark that counts, and you can change any default in Settings.