KB: Claude Sonnet 5 + model routing (Opus 4.8 / Sonnet 5 / Haiku 4.5) — and the Codex parallel¶
Research note
This page is background research kept for the project's own reference. It is not user documentation, it is not kept up to date with every release, and model or product details in it may be out of date. For how anti-hall works today, start at the home page.
Reference KB. Benchmark numbers rest on third-party agreement (the Anthropic system-card PDF was unparseable at time of writing) and are tagged with sources; treat exact figures as directional until checked against the official model card. Model IDs and pricing are from Anthropic docs. Dual-platform note: anti-hall routes both Claude and Codex — the Codex-model table is in §8; orchestration parity is OMC ↔ OMX (see
KB-omc.md/KB-omx.md).[CORRECTION 2026-08-22 — title/lineup one full generation stale; every "Opus 4.8" reference below (22 occurrences) is now legacy, not current. Verified current lineup: Claude Fable 5 (
claude-fable-5, current flagship, 1M ctx/128k out, $10/$50) and Claude Mythos 5 (claude-mythos-5, same underlying "Mythos-class" model, gated to approved orgs) now sit above Opus 5 (claude-opus-5, enterprise/agentic, 1M/128k, $5/$25 — Opus 5's price/context/output match what this doc lists for Opus 4.8; Opus 5 is Opus 4.8's direct successor, shipped 2026-06-09). Opus 4.8 is now DEPRECATED / legacy — API-valid but superseded; the vendor migration guide covers Opus 4.8 → Fable 5. Sonnet 5 is unchanged as general-purpose EXCEPT pricing: the §2 table's "$3/$15 ($2/$10 intro → Aug 31 2026)" is now WRONG — the scheduled Sept 1 price increase was CANCELLED, so Sonnet 5 stays $2/$10 indefinitely, not just as an intro rate. Haiku 4.5 is unchanged (200k ctx — disqualified as a main-agent downshift target perKB-model-modes.md§13, added 2026-08-22). None of §3-§6's Opus-4.8-branded benchmarks/decision guidance are re-verified for Fable 5/Mythos 5/Opus 5 this pass — read every "Opus 4.8" below as "the flagship-class model at the time this KB was compiled," not a current pin. §7's routing table is the one section with operational consequence — see the inline correction there: the actual shipped policy (plugins/anti-hall/skills/MODEL-POLICY.md) already routes by tier token (opus/sonnet/haiku/fable), resolved to the newest model in-family at runtime, so it was NOT stale — only this KB's hardcoded "Opus 4.8" labels in the seat map are. Per this KB's own §10 "Discrepancies / caveats" pattern, left as stacked annotation, not rewritten.]**[UPDATE 2026-09-03 — Fable point release.
claude-fable-5-1(released 2026-09-01) supersedesclaude-fable-5as current Fable; this is a point release, not a new generation, so nothing else in the lineup above changes. Pricing/context forclaude-fable-5-1are not independently verified here — seeKB-fable-5.md's 2026-09-03 update.fabletier-token routing is unaffected.]**
1. TL;DR¶
- Sonnet 5 (
claude-sonnet-5, 2026-06-30) is a clear win for implementation — near-Opus coding at ~30–40% real savings (the ~40% sticker discount is eroded by a 1.3–1.4× tokenizer). - Sonnet 5 @xhigh ≈ Opus 4.8 @medium-to-high (per §4's DataCamp citation — the source hedges across both tiers, not just @medium) — so it's a great cheaper secondary planning/review tier, but does not replace Opus on the hardest work.
- Opus 4.8 stays ahead on deep reasoning (GPQA +15.6), olympiad math (USAMO +17.2), and deep multi-file coding (SWE-bench Pro +6.0). Keep it for root-cause, architecture, top-level planning.
- Haiku 4.5 stays the cheap, fast default (91–104 t/s, $1/$5) for trivial/mechanical/high-volume.
2. Model facts (from Anthropic docs)¶
| Opus 4.8 | Sonnet 5 | Haiku 4.5 | |
|---|---|---|---|
| Model ID | claude-opus-4-8 |
claude-sonnet-5 |
claude-haiku-4-5 |
| Released | 2026-05-28 | 2026-06-30 | ~2025-10 |
| Price $/M (in / out) | $5 / $25 | $3 / $15 ($2/$10 intro → Aug 31 2026) | $1 / $5 |
| Context / max output | 1M / 128k | 1M / 128k (300k batch) | 200k / 64k |
| Thinking | adaptive, scales with effort |
adaptive, scales with effort |
extended thinking (explicit budget) |
| Claude Code/API default effort | high |
high |
n/a (no effort support) |
| Knowledge cutoff | Jan 2026 | Jan 2026 | Feb 2025 |
Correction (2026-07-01): an earlier version of this KB claimed
xhighwas the default effort for Opus 4.8/Sonnet 5. That was wrong — verified against the live official effort docs (platform.claude.com/docs/en/build-with-claude/effort, fetched 2026-07-01):highis the default ("produces exactly the same behavior as omitting theeffortparameter entirely");xhighis an available, non-default tier. Corrected below; seedocs/KB-token-usage-models.mdfor the full effort-tier + billing research this correction came from.
Notes: the full API effort taxonomy is five tiers — low, medium, high (default), xhigh,
max — not a Sonnet-specific set. xhigh support is narrow: only Opus 4.8, Opus 4.7, and Sonnet 5
among the models in this KB (official). Haiku 4.5 does not support the effort parameter (or
adaptive thinking) at all — it's legacy manual thinking:{budget_tokens} only. Sonnet 5 does
not support temperature / top_p / top_k. Opus 4.8 and Sonnet 5 share the newer tokenizer
introduced at Opus 4.7 (official range: ~1.0–1.35× tokens vs pre-4.7 for the same text, "~30%"
aggregate; independent testing found up to ~1.47× for some content types — see
docs/KB-token-usage-models.md §3 for the full tokenizer research).
3. Benchmarks (source-tagged; directional)¶
| Benchmark | Opus 4.8 | Sonnet 5 | Haiku 4.5 | Winner |
|---|---|---|---|---|
| SWE-bench Verified | 88.6 | 72.7 (official) / 82.1 (ext. budget) | 73.3 | Opus |
| SWE-bench Pro | 69.2 | 63.2 | 39.5 | Opus +6.0 |
| GPQA Diamond | 93.6 | 78.0 | — | Opus +15.6 |
| USAMO 2026 | 96.7 | 79.5 | — | Opus +17.2 |
| HLE (no tools) | 49.8 | 43.2 | — | Opus +6.6 |
| HLE (with tools) | 57.9 | 57.4 | — | ~tie |
| OSWorld-Verified | 83.4 | 81.2 | — | ~tie (Opus +2.2) |
| Terminal-Bench 2.1 | 74.6* | 80.4 | — | Sonnet +5.8 |
| GDPval-AA v2 (knowledge work) | 1615 | 1618 | — | Sonnet +3 |
| MMLU-Pro / AIME | unconfirmed | unconfirmed | 72.4 / 80.7 | — |
| Output speed (t/s) | moderate (unmeasured) | 57 @high / 78.9 @max | 91–104 | Haiku |
*Terminal-Bench Opus figure is harness-dependent, not effort-dependent: 74.6 raw model vs 78.9 with
Claude-Code harness (see docs/KB-codex-vs-opus-coding.md §Benchmark table). The 82.7% figure
sometimes seen alongside this benchmark is GPT-5.5's (Interesting Engineering, source [13] in that
doc), not Opus's — do not conflate.
4. Effort-tier behavior¶
low (renames/greps/classification) · medium (coding, small refactors, tests) · high (default;
multi-file refactors, complex debugging) · xhigh (long agentic, exploratory) · max
(architecture, race conditions, security — confirmed for Opus; unconfirmed for Sonnet 5).
Key finding: "At its maxed-out xhigh reasoning level, Sonnet 5 performs roughly in line with
Opus 4.8's medium-to-high setting" (DataCamp). Effort alone does not close Sonnet 5 → Opus 4.8
to parity on the hardest tasks. Sonnet 5 @max: generation ~78.9 t/s but TTFT ~163s (deep
reasoning phase) — expensive and slow; avoid inside loops.
5. Decision matrix (task → model + effort)¶
| Task type | Model | Effort | Why |
|---|---|---|---|
| Trivial edit (rename, grep, format) | Haiku 4.5 | — | 104 t/s, no reasoning overhead |
| Mechanical impl (CRUD, boilerplate) | Haiku 4.5 | — | within Haiku's tier; ~80× cost edge on Opus |
| Standard impl (1–3 files, known stack) | Sonnet 5 | medium/high | 63.2 SWE-Pro covers it; cheap |
| Complex impl (multi-file, arch-aware) | Sonnet 5 | xhigh | closes most of the Opus gap; escalate only on failure |
| Deep debugging (races, subtle logic, security) | Opus 4.8 | high | GPQA 93.6 vs 78; reasoning depth matters |
| Olympiad / proof-level math | Opus 4.8 | high/max | USAMO +17.2 decisive |
| Architecture planning / design review | Opus 4.8 | high/xhigh | HLE-no-tools +6.6 (deep constrained reasoning) |
| Correctness / safety review | Opus 4.8 or Codex | high | GPQA edge; cross-model for independence |
| Long-horizon agentic (multi-turn, 1M ctx) | Sonnet 5 | high | beats Opus on GDPval-AA v2 + terminal-bench, 40% cheaper |
| High-volume production agents (latency SLA) | Haiku 4.5 / Sonnet 5 low |
— | throughput |
6. Switch thresholds¶
- Haiku → Sonnet 5 when: >2 inference hops; Haiku hallucinates field names / gets lost past ~50k ctx; needs post-Feb-2025 knowledge; needs >64k output or >200k input.
- Sonnet 5 → Opus 4.8 when: GPQA/USAMO-class scientific/math reasoning is the bottleneck; formal proof / security audit / multi-repo refactor with subtle invariants; Sonnet 5 @xhigh failed twice on the same task; error blast-radius is high (prod security, compliance, correctness- critical path).
- Do NOT use Opus for: knowledge work, long-horizon agentic pipelines, terminal tasks, customer-facing chat — Sonnet 5 matches/beats it there at ~40% lower price.
- [ADDED 2026-08-22] Downshift (usage-limit conservation), not capability-based switching:
when a session must move off the flagship to conserve limits, the main agent needs 1M context —
Sonnet 5 is the correct downshift target (1M ctx, $2/$10, general-purpose); Haiku 4.5 is
disqualified for the main agent (200k ctx, not 1M), though it's still correct for
trivial/leaf subagent work. The agent cannot self-switch models (
/modelis a user action) — a downshift surfaces only as a recommendation. Full treatment:KB-model-modes.md§13.
7. anti-hall routing (Claude side)¶
[CORRECTION 2026-08-22] This table's "Opus 4.8" / "Sonnet 5" labels describe the model class at each seat, not a pinned version — verified against the live
plugins/anti-hall/skills/MODEL-POLICY.md, which routes by tier token (opus/sonnet/haiku/fable), resolved to the newest model in-family at runtime (its own stated rule: "Never pin a model version"). So the "opus" seats below now actually resolve to Opus 5 (claude-opus-5), not Opus 4.8, and MODEL-POLICY.md additionally routes the deadly-loop Reviewer seat to Fable first when available (fable-availability.js, re-enabled 2026-07-12), falling back to Sonnet then Opus — a routing detail this table does not show. Treat this table as directionally correct but not the live seat map; MODEL-POLICY.md is ground truth.
Effective seat map (mirrored in skills/MODEL-POLICY.md):
| Seat | Model | Effort |
|---|---|---|
| Main coordinator | Opus 4.8 | — |
| Implementation (code-apply) | Codex primary → Sonnet 5 failover | Sonnet 5 @high on failover |
| Planning — L / top-level | Opus 4.8 | xhigh |
| Planning — M / secondary | Sonnet 5 | xhigh |
| deadly-loop Reviewer | Sonnet 5 | xhigh |
| deadly-loop Auditor | Opus 4.8 | high |
| deadly-loop Critic | Codex (Opus/Sonnet 5 if Codex implemented the diff) | — |
| Root-cause / deep debug | Opus 4.8 | high |
| Trivial | Haiku 4.5 | — |
Rules: (1) the code implementer and its correctness reviewer are always different models
(cross-model second opinion, no self-review); (2) Codex is primary implementer because it draws its
own limit (conserves the Claude weekly bucket) — failover to Sonnet 5 on unavailable /
rate-limited / placeholder return; backoff, never retry-loop; (3) never run Sonnet 5 at max
inside a loop (TTFT ~163s + cost); (4) the sonnet tier token now resolves to claude-sonnet-5.
8. Codex-model parallel (gpt-5.x) — for the Codex port¶
anti-hall ships a Codex-native port, so routing has a Codex mirror. GPT-5.6 (Sol/Terra/Luna)
went GA 2026-07-09, superseding gpt-5.5/gpt-5.4 as the recommended tiers — see
KB-gpt-5.6.md for full sourcing/confidence-labeling; this section reflects
that migration. gpt-5.4-mini is kept as the cheap-default tier (unchanged; still the cheapest
confirmed-priced option) alongside the new gpt-5.6-luna as a selective alternative.
Model facts (KB-gpt-5.6.md §1–2, PRIMARY-CONFIRMED for pricing/tier-ID/positioning; context
window, knowledge cutoff, and full effort-tier taxonomy were not independently re-confirmed for
GPT-5.6 and are marked accordingly — do not carry over the old gpt-5.5/gpt-5.4 numbers as fact):
| | gpt-5.6-sol (frontier) | gpt-5.6-terra (flagship) | gpt-5.4-mini | gpt-5.6-luna (alt.) |
|---|---|---|---|---|
| Price $/M (in / out) | $5 / $30 | $2.50 / $15 | $0.75 / $4.50 | $1 / $6 |
| Cached input $/M | $0.50 (confirmed) | [unverified] | $0.075 | [unverified] |
| Context | [unverified for GPT-5.6 — not re-confirmed, see KB-gpt-5.6.md §9] | [unverified] | 400k | [unverified] |
| Knowledge cutoff | [not stated in confirmed GPT-5.6 sourcing] | [not stated] | Aug 2025 | [not stated] |
| Max effort | documented max/"ultra" tier, Sol-only (2-1 vote, not unanimous — KB-gpt-5.6.md §1) | "standard reasoning-effort range," no documented max | high (likely) | "standard reasoning-effort range" |
| Claude equivalent | Opus 4.8 | Sonnet 5 | Haiku 4.5 | Haiku 4.5 |
Reasoning effort: model_reasoning_effort = minimal | low | medium | high | xhigh (Responses
API only; medium was the confirmed default for the gpt-5.5/gpt-5.4 generation). GPT-5.6-tier
effort defaults are not independently confirmed — treat the same minimal/low/medium/high/xhigh
scale as a working assumption pending direct verification, not a confirmed GPT-5.6 fact. Rough
parity: Codex minimal/low ≈ Claude thinking-off, medium ≈ Claude default, high ≈ ~8k budget,
xhigh ≈ 16k+ budget.
Benchmarks (directional; vendor vs standardized-harness noted). These are the last confirmed
figures for the superseded gpt-5.5/gpt-5.4/gpt-5.4-mini generation — GPT-5.6's own reported
benchmarks (Artificial Analysis Coding Agent Index: Sol scores 80; SWE-Bench Pro: Sol 64.6%, both
press-reported per KB-gpt-5.6.md §6) use different suites/formats and are not directly
comparable to the rows below; do not merge them into one table:
| Benchmark | gpt-5.5 | gpt-5.4 | gpt-5.4-mini |
|---|---|---|---|
| SWE-bench Verified | 88.7 (vendor) / 80.6 (LMC std) | 76.9 (LMC) | — |
| SWE-bench Pro | 58.6 | 59.1 (Scale SEAL) | 54.4 |
| GPQA Diamond | 93–94 | 93.3 | 88 |
| Terminal-Bench 2.0 | 84.7 | 81.8 | — |
| OSWorld-Verified | 78.7 | 75.0 | — |
Codex decision matrix: trivial → gpt-5.4-mini @minimal · mechanical impl → gpt-5.4-mini @low · standard impl → gpt-5.6-terra @medium · complex impl → gpt-5.6-terra @high or gpt-5.6-sol @medium · debugging / planning → gpt-5.6-sol @high · correctness-review → gpt-5.6-sol @xhigh · long-horizon agentic → gpt-5.6-sol @medium · parallel subagent → gpt-5.4-mini @low–medium (or gpt-5.6-luna where available).
Switch thresholds: mini → gpt-5.6-terra when multi-file reasoning / mini errors 2× / needs
high. gpt-5.6-terra → gpt-5.6-sol when concurrency/cross-system debugging, correctness-critical
xhigh, or Sol's deeper reasoning tier matters. Downgrade for budget sprints (gpt-5.6-terra default;
gpt-5.6-sol for review/unblock).
Cross-platform equivalence: gpt-5.6-sol ≈ Opus 4.8 · gpt-5.6-terra ≈ Sonnet 5 (affordable-flagship class) · gpt-5.4-mini / gpt-5.6-luna ≈ Haiku 4.5 (subagent/latency tier). This equivalence mirrors the prior gpt-5.5≈Opus / gpt-5.4≈Sonnet 5 / gpt-5.4-mini≈Haiku framing by tier position, not by re-verified benchmark parity — GPT-5.6's own benchmark numbers (§ above) were not cross-compared against Claude figures in this pass.
Orchestration (OMX ↔ OMC): OMX (oh-my-codex) is the Codex analogue of OMC. scalarian/oh-my-codex
was archived 2026-05-25; Yeachan-Heo/oh-my-codex is the active fork. OMX does not auto-select
models — the user/orchestrator sets model + model_reasoning_effort in .codex/config; skills
recommend effort in-prompt but can't switch models mid-session. Pattern: orchestrator on flagship,
$team parallel workers on mini. See KB-omx.md / KB-omc.md.
9. Sources¶
Official Anthropic: (1) platform.claude.com models overview · (2) platform.claude.com pricing. Third-party (2026): (3) morphllm.com Claude benchmarks · (4) MarkTechPost — Sonnet 5 vs Sonnet 4.6 vs Opus 4.8 · (5) codingfleet.com — Sonnet 5 vs Opus 4.8 · (6) llm-stats.com — Sonnet 5 vs Opus 4.8 · (7) cosmicjs.com — Sonnet 5 benchmarks · (8) datacamp.com — Claude Sonnet 5 · (9) artificialanalysis.ai — Sonnet 5 · (10) artificialanalysis.ai — Haiku 4.5 · (11) anthonymaio.substack.com — effort levels.
Codex (gpt-5.x) — official OpenAI: (12) developers.openai.com/codex/models · (13) developers.openai.com/api/docs/pricing · (14) developers.openai.com/codex/config-reference. Third-party: (15) LM Council benchmarks · (16) morphllm SWE-bench Pro. OMX: (17) github.com/Yeachan-Heo/oh-my-codex.
10. Discrepancies / caveats¶
- Terminal-Bench 2.1 Opus: 74.6 (raw model) vs 78.9 (Claude-Code harness) — harness-dependent, not
effort-dependent (see
docs/KB-codex-vs-opus-coding.md). The 82.7% Terminal-Bench figure circulating in some sources is GPT-5.5's, not Opus's. - Sonnet 5 SWE-bench Verified: 72.7 (official announcement) vs 82.1 (third-party, extended thinking).
- GDPval: "Sonnet beats Opus" holds on v2 only (1618 vs 1615); Opus led the original v1.
- MMLU-Pro / AIME / LiveCodeBench / TAU-bench: no published Opus-4.8/Sonnet-5 figures found as of 2026-07-01.