KB.md — Canonical Knowledge Base for anti-hall¶
Source of truth for developing and maintaining the anti-hall plugin. This is the map. The deep research lives in the section docs linked below; this file is the authoritative index, the current-plugin ground truth, and the staleness ledger that sits over all of them. When code and a doc disagree, code wins — and the disagreement gets logged in §4 Staleness ledger.
0. How to use this KB¶
What it's for. Anti-hall is a verify-first / anti-hallucination Claude Code plugin. This KB is the dev-facing knowledge layer the plugin is built and maintained from: the prompting/hook/orchestration research it encodes, the design rationale behind each mechanism, and the current ground-truth state of the shipped plugin.
How it's organized.
| Layer | What it holds | Where |
|---|---|---|
| Index + ground truth | This file. Current plugin facts, navigation, freshness + staleness ledgers. | KB.md (you are here) |
| Living reference docs | Deep, citation-backed research. Still authoritative for their topic. | §2 |
| Folded-in source docs | Research that has been distilled into the living KB synthesis; kept for provenance. | §2 (marked folded) |
| Historical artifacts | Dated audits, reviews, and one-time plans. Frozen records — never edited to match current code. | §5 History |
How to keep it fresh (the rule):
- On any significant change (new hook, renamed file, version bump, changed nudge/footprint/count, new skill), update §1 Current plugin ground truth in the SAME change. The CHANGELOG is the authority for what changed; this KB mirrors the current resting state.
- Cite date + ref for every new entry. Use the date already in the source
(CHANGELOG entry, commit, source doc header). Never invent a date — if
unknown, write
(date unknown). - Never silently rewrite a fact you can't verify. Flag it in
§4 Staleness ledger with
doc:lineand the reason. - Code is the tiebreaker. Any prose claim about plugin behavior must be
checkable against a file in
plugins/anti-hall/. If it isn't, it's an aspiration, not a fact — label it.
This file is the source of truth for plugin-development decisions. Design debates reference the living docs for evidence; the current state and the open discrepancies are settled here.
1. Current plugin ground truth¶
[UPDATE 2026-09-24, v0.108.0] Re-verified against the working tree on 2026-09-24:
plugin.jsonversion is0.108.0. Hooks: 73.jsfiles underplugins/anti-hall/hooks/(incl. shared library modules; 62 scripts registered inhooks.json, 44 of them also incodex/hooks/hooks.json). New hooks this release:auto-handover.js(UserPromptSubmit),auto-handover-pause-nag.js(Stop),repair-on-reload.js(SessionStart + UserPromptSubmit); unreleased on top of that (post-0.108.5):silent-agent-nudge.js(Stop). Claude skills: 29 directories underplugins/anti-hall/skills/(+MODEL-POLICY.md, not itself a skill; new:settings); Codex skills: 32 directories underplugins/anti-hall/codex/skills/(new:anti-hall-settings).PostToolUsewired: 3 Bash-matcher handlers (output-verify-guard.js,devswarm-parent-reply-tracker.js,devswarm-child-drain.js).PostToolUseFailurewired: 1 Bash-matcher handler (failure-root-cause-nudge.js). Counts re-derived fromls plugins/anti-hall/hooks/*.js,ls -d plugins/anti-hall/skills/*/,ls -d plugins/anti-hall/codex/skills/*/, andhooks.json— not carried forward from the stale blocks below. The earlier blocks are retained for their per-version prose history, not overwritten (per this doc's own §6 recommendation for a generated ground-truth block instead of hand-maintained counts — still not implemented). The complete condensed component catalog (every hook, skill, script, setting) isllms.txt.[UPDATE 2026-08-22] Re-verified against the working tree on 2026-08-22: committed
plugin.jsonversion is0.75.1; the working tree carries0.76.0(release in flight, uncommitted per release-prep convention — do not assert 0.76.0 has shipped). Hooks: 53.jsfiles +hooks.json= 54 (added this session:defect-nudge.js,claude-cli-version.js,claude-cli-version-refresh.js,repo-self-drift.js). Claude skills: 16 (+MODEL-POLICY.md, not itself a skill); Codex skills: 19.PostToolUsewired: 3 Bash-matcher handlers —output-verify-guard.js(hooks.json:370),devswarm-parent-reply-tracker.js(:380),devswarm-child-drain.js(:390).PostToolUseFailurewired: 1 Bash-matcher handler —failure-root-cause-nudge.js(hooks.json:402). The 2026-06-04 snapshot below is retained for its per-version prose history, not overwritten.Verified against the working tree on 2026-06-04 (commit context: post
0.24.1;0.22.0shipped api-guard (mechanical API-hallucination guard) + the eval harness,0.22.1/0.22.2Windows-CI fixes,0.23.0api-guard v2 (opt-in 3rd-party after a deadly-loop proved default-on = edit-time RCE),0.24.0extended git-guard to block AI self-credit in gh pr/issue/release bodies,0.25.0added the always-on SCOPE & FIDELITY discipline + the opt-in mcp-reaper companion (macOS + Linux),0.26.0made rule G's structured-return discipline concrete + S4-measured (~5× smaller than verbose prose, lossless),0.27.0replaced feature-launch with the leanship-itworkflow (plan-mode + Workflow-swarm + deadly-loop, scaled S/M/L) and added the deadly-loop D1.5 verification gate (fresh-evidence + vacuous-test guard) — all in ground truth). This block is the canonical snapshot — update it on every significant change. Ref:plugins/anti-hall/.claude-plugin/plugin.json,plugins/anti-hall/hooks/,plugins/anti-hall/skills/,CHANGELOG.md.
| Fact | Value | Source / verify |
|---|---|---|
| Version | 0.108.0 |
plugins/anti-hall/.claude-plugin/plugin.json version — the single authority (the Codex manifest .codex-plugin/plugin.json is bumped with it). Marketplace entry carries NO version. What changed per release: CHANGELOG.md. The per-version narrative this row used to carry (0.58.0–0.71.1) is kept in §5.1. |
| Runtime | Pure Node, built-ins only; requires Node.js ≥ 22 on PATH | plugin.json description; hooks launched as node <hook>.js |
| Hook language | All hooks are .js (NOT .sh) |
ls plugins/anti-hall/hooks/ |
| Hooks shipped (73 files) | 73 .js files under plugins/anti-hall/hooks/ (incl. shared modules); 62 scripts registered in hooks.json. One row per hook, current behaviour: llms.txt → Hooks. Per-hook detail: GUIDE.md. The per-hook history this row used to carry is in §5.1. |
ls plugins/anti-hall/hooks/*.js, hooks/hooks.json |
| Skills shipped (29) | root-cause, orchestration, ship-it, deadly-loop, deadly-loop-multi, doctor, install-statusline, update, activate, simplify, debt, devswarm, handover, system-briefing, defects, jev, settings (+ shared MODEL-POLICY.md). Codex: 32 anti-hall-* skills (same set plus context-conserve, model-policy, omc, omx). One-line purpose of each: llms.txt → Claude skills / Codex skills. |
plugins/anti-hall/skills/, plugins/anti-hall/codex/skills/ |
| Cadence — full protocol | Injected at SessionStart via TWO hooks (0.60.0 SPLIT so each clears the ~10k per-hook injection cap): verify-first-full.js = the FOUNDATION (Iron Law + rationalization table + Positive Rules + Scope & Fidelity + the DISCIPLINES/SKILLS index) and verify-first-orch.js = the ORCHESTRATION DISCIPLINE ruleset (rules A–N + DevSwarm-Primary rule W). Also injected at SubagentStart via verify-first-subagent.js (0.39.0: every spawned subagent now receives the Iron Law; the SubagentStart hook deliberately omits the orchestration/delegate block to prevent nesting recursion; shared core in verify-first-core.js). verify-first-full.js — includes the always-on SCOPE & FIDELITY discipline (0.25.0: simplest sufficient solution; intent over letter; confirm before expanding scope; match rigor to blast radius; finish what was asked / drop nothing), named in the "ALWAYS APPLY" disciplines list alongside root-cause/orchestration/anti-sycophancy. Orchestration (0.25.0) also requires the coordinator to independently verify delegated work — a subagent's "done/passing" is an unverified claim re-checked against ground truth before marking complete (rule L), and now defaults delegated heavy/parallel work to the background (the coordinator passes run_in_background so the user needn't background it manually) while still verifying each on completion — never fire-and-forget (rule F). SessionStart re-fires after compaction with source="compact", so the same no-matcher entry covers the post-compact reset. The only PreCompact hook is precompact-snapshot.js (0.108.0): a side-effect-only file snapshot that never prints, because a PreCompact hook's exit 2 or decision: "block" blocks compaction (official hooks reference; see KB-handover-research §4). |
hooks/hooks.json, hooks/verify-first-full.js Cost-trim Phase 3 (default context.protocolLevel=compact): the SessionStart text is a compact core (≈3.0k chars, every load-bearing clause inline) that points at the generated plugins/anti-hall/PROTOCOL.md; the orchestration rules A–N still arrive in full at SessionStart (default context.orchFullOn=auto=session; hook chars 10.7k instead of 16.6k). The experimental, opt-in orchFullOn=spawn sends ORCH_COMPACT (≈1.0k) at SessionStart and the full text once per context epoch on the first Agent/Task/Workflow spawn (orch-on-spawn.js, Claude Code only), unverified live; it costs ≈+1k chars per spawning epoch beyond one full send. context.protocolLevel=full restores today's text byte for byte. |
| Message-bloat prevention (#45, 0.25.0; rule G concretized 0.26.0) | Rule G is SYNTHESIZE, NEVER RELAY: the coordinator never pastes a subagent's raw return into the user thread (verbatim relay = the #1 cause of message-context bloat); subagents return tight summaries under an OUTPUT BUDGET. For a SUBSTANTIAL return (review/audit/research dump, many claims) rule G now specifies a concrete structured shape — {claim, evidence:"file:line", verdict, blockers/uncertainty, next} — which an S4 study measured ~5× smaller than verbose prose with zero decision-relevant loss (claim/evidence/uncertainty/blockers/next rubric, N=8, directional). The prose-for-tiny caveat is kept: a single prose line wins for a SMALL result (schema only ~1.4× denser there, and JSON overhead can make tiny outputs LARGER). Reconciliation: the earlier ~1.43× pilot figure measured small, already-summarized content; the ~5× figure is for verbose returns — both hold, different inputs. To enforce rather than request, pass a schema to the Agent/Task tool (validated structured return); the biggest levers remain the output budget + no-raw-relay rule, the schema is the multiplier on large returns. Shipped as prompt discipline (rule G text + nudge #57), not a schema-enforcement system. A PostToolUse-on-Task size-flag hook was evaluated and is NOT feasible: per KB-claude-codex §1.4 PostToolUse stdout/additionalContext never reaches the model and it cannot block — no hook was faked. [UPDATE 2026-08-21: this delivery claim now reads misleadingly next to the fact that anti-hall ships 3 real PostToolUse hooks (output-verify-guard.js, devswarm-parent-reply-tracker.js, devswarm-child-drain.js) and 1 PostToolUseFailure hook (failure-root-cause-nudge.js) — see KB.md:71 and docs/KB-claude-code-harness-features.md §2. Clarification: those shipped hooks are ADVISORY/SIDE-EFFECT-ONLY (they observe and record, e.g. tracking a reply or nudging on failure) — they do not rely on additionalContext reaching the model to function, which is consistent with this row's claim. The underlying delivery claim itself — whether PostToolUse/additionalContext reaches the model on the current CLI (2.1.238) — was NOT re-tested this session. Mark NEEDS-CONFIRMATION; do not assert it either way pending a fresh probe.] [CORRECTION 2026-08-22: the NEEDS-CONFIRMATION probe above is now resolved — WRONG on the current CLI (2.1.238). Directly observed live this session: turns received PostToolUse:Bash hook additional context: … and PostToolUse:Write hook additional context: … system-reminders in the transcript (first-party observation, not a doc reading). Corroborating: the plugin ships four PostToolUse hooks emitting hookSpecificOutput.additionalContext — output-verify-guard.js, devswarm-child-drain.js, failure-root-cause-nudge.js (PostToolUseFailure) and devswarm-parent-reply-tracker.js (side-effect-only) — written expecting delivery. CAVEAT: the official docs (code.claude.com/docs/en/hooks.md) list additionalContext as a supported field but do not explicitly state it reaches the model for PostToolUse the way they do for UserPromptSubmit — so this is cited as live-observed 2.1.238 behavior, not doc-confirmed, and should be re-verified after a CLI upgrade. Corrected conclusion: PostToolUse cannot BLOCK a gate-style decision, but it CAN inject advisory context — this re-opens the PostToolUse-on-Task size-flag design option that this row previously ruled out as infeasible.] |
hooks/verify-first-full.js (rule G), hooks/verify-first.js |
| Cadence — per-turn nudge | verify-first.js on UserPromptSubmit emits ONE short nudge, deterministically chosen by SHA-1 of the stdin envelope mod NUDGES.length. |
hooks/verify-first.js |
| Nudge count | 20 rotating one-liners (0.25.0 added 2 scope-fidelity + verify-delegated-work + background-default + synthesize-never-relay nudges; 0.30.0 added the done-bar/AGREED-acceptance-criteria nudge — DONE = verified against agreed acceptance, un-mechanically-verifiable fidelity = PENDING OWNER REVIEW, self-issued hedge hard-blocks 'done' + auto-merge — bringing the count to 18 at that point; 0.36.0 added 2 more — rule M shallow+wide/no-re-delegate and rule N distribute-models-per-seat — bringing the count to the current 20; NOT "5") | NUDGES array in verify-first.js (% NUDGES.length) |
| SessionStart injection footprint | SPLIT (measured): the single verify-first-full.js payload had grown to 15,323 chars (baseline) / 15,400 (DevSwarm Primary) — over the ~10k-char per-hook injection cap (KB-claude-codex.md §1.6). That cap is spill-to-file, NOT truncation: a >10,000-char injection delivers only ~2,000 chars inline + a file path the model must choose to open, so rules A–N + the DISCIPLINES/SKILLS block reached NObody's context inline. Because the cap is per hook COMMAND, the doctrine was SPLIT across two SessionStart hooks, each <10,000 so BOTH land 100% inline with zero content deleted: verify-first-full.js = the verify-first FOUNDATION (Iron Law + rationalization table + Positive Rules + Scope & Fidelity) + the DISCIPLINES-vs-SKILLS index → 7,672 chars; verify-first-orch.js = the ORCHESTRATION DISCIPLINE ruleset (rules A–N + DevSwarm rule W for a Primary) → 7,728 chars baseline / 8,141 Primary. The earlier reorder that hoisted rule L out of alphabetical order "to survive the cap" was REVERTED — with the split there is no truncation, so L is restored to its natural slot (between K and M); and rule W (squeezed to 76 chars under the old cap) was restored to its operative form (workspace = top tier, the node scripts/devswarm.js spawn <branch> -p "<brief>" command, the choice rule, the failure mode). Regression net: tests/hooks/injection-cap.test.js executes every context-injecting hook and asserts each emitted payload <=10,000 (fails against the pre-split 15,323 payload, passes after). |
Measured 2026-07-14; hooks/verify-first-full.js, hooks/verify-first-orch.js, tests/hooks/injection-cap.test.js |
| AGENTS.md mirror | Present at repo root (Codex/clone-based governance). NOT bundled by /plugin install. |
AGENTS.md |
| Model policy | Cross-model debate TRIO: Reviewer = Fable when available (model:"fable"), else Sonnet 5 (model:"sonnet") — policy-disabled 0.43.2, RE-ENABLED 2026-07-12, see CHANGELOG; Auditor = latest Opus (model:"opus", divergent regression/coupling lens); Critic = latest OpenAI Codex at xhigh reasoning (Opus adversarial-persona fallback). "Latest" is resolved at runtime — spawn paths use ONLY tier tokens (fable/opus/sonnet/haiku; Codex = latest the installed CLI reports). No versioned/dated model IDs in executable spawn snippets. API call sites are the sole exception (the API has no evergreen tier alias; claude-haiku-4-5 is the alias-form for its tier). |
skills/MODEL-POLICY.md (duplicated to skills/deadly-loop/references/MODEL-POLICY.md, byte-identical; 2 copies, not 3 — symlinks stripped on install) |
| Model tier tokens resolve latest-at-call-time | model:"fable" / "opus" / "sonnet" / "haiku" in Agent/Task tool and Workflow agent() resolve to the newest available build at call time — they are NOT pinned version IDs. Consequence: spawn snippets in docs/skills MUST use tier tokens only; a versioned ID (e.g. claude-opus-4-8) in a snippet would pin a snapshot and age. Verified live 2026-06-10 (step0-probe-record.md P1: model:"fable" accepted, spawned general-purpose (fable)). |
tests/fixtures/step0-probe-record.md P1 |
lastModelUsage has NO timestamps |
The live ~/.claude.json lastModelUsage field carries cumulative token/cost counters per model key — NO lastUsedAt / timestamp fields. Observed sample: {inputTokens, outputTokens, cacheReadInputTokens, cacheCreationInputTokens, webSearchRequests, costUSD} only. JSON.stringify(lastModelUsage).includes('lastUsedAt') === false verified 2026-06-10. Consequence: no reliable parent-model signal exists in hook stdin; strict mode is unconditional (not heuristic). The statusline-rich.js max-lastUsedAt loop reads a phantom field (ts always 0, masked by last-key fall-through) — latent bug fixed in 0.32.0 to last-key explicitly. |
tests/fixtures/step0-probe-record.md P6 |
PreToolUse additionalContext observed reaching model |
2026-06-10 (multiple instances, current build): hookSpecificOutput.additionalContext emitted by a PreToolUse hook appeared as model-visible context blocks (PreToolUse:Agent hook additional context: …) and was acted on by the model. This CONTRADICTS docs/KB-claude-codex.md:47 (official-docs sourced, older harness claim that only UserPromptSubmit/UserPromptExpansion/SessionStart deliver context). Both dated probes retained — see KB-claude-codex.md:44-47 annotation. The advisory rows in model-routing-guard rely on this channel; on a harness that does not deliver it, advisories are inert no-ops (fail-open, nothing breaks). |
tests/fixtures/step0-probe-record.md P2; gate A3-1 |
| Test net | Zero-dependency node:test E2E suite (black-box, process-level), 701 passing (+2 platform-skipped), 703 total — guards, api-guard (stdlib/builtin/3rd-party + shadowing/portability/RCE-no-execute), git-guard gh-PR self-credit, speculation, tasklist-guard (+ history-ledger reminder), task-guard idle-neglect (actionable-now + no-agents → sharp block; agents-running/blocked/owned → generic; dedupe + cap), scope-fidelity + verify-delegated-work + background-default + synthesize-never-relay protocol/nudge regression, etc. |
tests/, E2E-TESTING.md, CHANGELOG |
| Companion (opt-in) | companion/mcp-reaper.js (+ install-reaper.js) — NOT a hook; an interval job (macOS LaunchAgent / Linux systemd --user timer) that kills ONLY orphaned MCP-server processes (parent already died). Safety invariant: reaps only when the command matches a generic MCP signature AND the parent is a reaper/init (launchd / init / systemd --user), so a live MCP (live spawner parent) can never be killed. Windows is not supported. Recognizes Python MCPs too (uvx/uv + underscore mcp_server_* forms). Limitation: an MCP run as a LaunchAgent / systemd --user unit / OS service shares init as a parent (like a leaked orphan) and can be reaped — exclude it via ANTIHALL_REAPER_EXCLUDE='name|name'. Install: node companion/install-reaper.js (--uninstall). Env: MCP_REAP_DRYRUN=1, MCP_REAP_GRACE, ANTIHALL_REAPER_MATCH, ANTIHALL_REAPER_EXCLUDE. |
plugins/anti-hall/companion/ |
Why these matter: the historical docs (PLUGIN-REVIEW.md,
ULTRAPLAN.md) describe the plugin before the cadence redesign —
they reference .sh hooks, "5 nudges", a missing PreCompact, and a missing
AGENTS.md. All of those were resolved on the way to 0.19.0. Read those docs as
history, not as current spec (see §5).
2. Living reference docs¶
These hold the deep, citation-backed knowledge. Topic, date, status, and the
authoritative reference are below. Folded means the content is also distilled
into the KB-claude-codex.md synthesis (kept standalone for provenance + depth).
| Doc | Topic | ~Size | Status | Date | Primary ref |
|---|---|---|---|---|---|
KB-claude-codex.md |
Backbone synthesis — hooks, plugins, prompting, Codex, orchestration, anti-hallucination evidence (§1–§14) | 672 ln | Living — primary | (compiled 2026-05) | 8 parallel research streams; cited inline + Sources §15 |
TASK-WORK.md |
Task discipline (TaskCreate/TaskUpdate vs legacy TodoWrite); event-driven, no-timer freshness; basis for the tasklist-guard feature |
241 ln | Living | (date unknown; references Claude Code v2.1.142 / SDK 0.3.142) | Anthropic long-running-agent guidance + hook model; Sources at doc end |
TASKLIST-GUARD.md |
Usage guide for the tasklist-guard Stop hook + per-turn freshness note: when it blocks, the per-session progress file (.anti-hall/progress/<date>/<session-id>.md), the per-session fix-ledger reminder (.anti-hall/history/<date>/<session-id>.md, append each completed task with Cause/Fix/Verified), the INDEX.md per convention, env knobs, escape hatch, good workflow |
— | Living | 2026-07-01 | this repo (tasklist-guard.js / task-tracker.js); design in TASK-WORK.md |
CONTRACT-1.0.md |
1.0 contract (draft) — what semver freezes: settings keys, CLI verbs and machine-read output, hook contracts, stable state paths, Claude/Codex parity | — | Living | 2026-10-03 | derived from code on dev (settings-schema.js, devswarm.js, hooks.json, migrations.js) |
BENCHMARK-METHOD.md |
Pre-registration of the with/without plugin benchmark (claude plugin eval): hypotheses, metrics, paired cluster-robust analysis, decision rule, amendments |
— | Living | 2026-10-03 | written methodology with numbered sources (S1 to S16) |
E2E-TESTING.md |
How the zero-dep node:test hook suite works; I/O contract per event; env-isolation gotcha |
129 ln | Living | 2026-06-02 (mtime) | Claude Code hooks contract |
AH-ENGINE.md |
The ah-engine (resident Rust hook engine, off by default): architecture, features implemented and planned, commands, metrics, configuration, troubleshooting | - | Living | 2026-10 | ah-engine/ (source), ah-engine/REFERENCE.md (generated) |
DEVELOPMENT.md |
Developer guide: prerequisites, build, every test suite, local install, debugging, porting workflow, branch and release flow, layout, coding standards | - | Living | 2026-10 | ah-engine/scripts/doc-check.sh (runs every command in it) |
HOOK-LATENCY.md |
Measured hook latency: wall p50/p95 and CPU per hook, and the per-tool-call total for each event (scripts/hook-latency.js) |
- | Living | 2026-10 | scripts/hook-latency.js |
REPO-PIPELINES.md |
GitHub workflows (trigger, purpose, gating), policy files, repo security settings | - | Living | 2026-10 | .github/workflows/ |
index.md |
Docs site page: Docs site home: what anti-hall is, install, what you'll notice in your first session. | - | Living | 2026-10 | mkdocs.yml |
start/install.md |
Docs site page: Install for Claude Code and Codex, keep .anti-hall/ out of git, check it works. |
- | Living | 2026-10 | mkdocs.yml |
start/update.md |
Docs site page: Update on Claude Code and Codex. | - | Living | 2026-10 | mkdocs.yml |
start/uninstall.md |
Docs site page: Remove the plugin, the statusline and the optional companions. | - | Living | 2026-10 | mkdocs.yml |
features/guards.md |
Docs site page: Each guard, its message, its setting, skipping one, and what guards do not do. | - | Living | 2026-10 | mkdocs.yml |
features/tasks.md |
Docs site page: Task tracking: task-guard, tasklist-guard, the progress file and the fix ledger. | - | Living | 2026-10 | mkdocs.yml |
features/handovers.md |
Docs site page: Automatic and manual handovers, resuming, keeping context small. | - | Living | 2026-10 | mkdocs.yml |
features/skills.md |
Docs site page: Every skill, when to use it, and the Codex skill names. | - | Living | 2026-10 | mkdocs.yml |
features/statusline.md |
Docs site page: The optional two-line statusline: install, consolidate, remove. | - | Living | 2026-10 | mkdocs.yml |
features/devswarm.md |
Docs site page: The optional DevSwarm integration and its companions. | - | Living | 2026-10 | mkdocs.yml |
how-it-works/index.md |
Docs site page: how it works, the Rust engine as the core component and the temporary Node fallback. | - | Living | 2026-10 | mkdocs.yml |
settings/index.md |
Docs site page: Changing settings: the skill, /config, the CLI, precedence, safety settings. |
- | Living | 2026-10 | mkdocs.yml |
troubleshooting.md |
Docs site page: The doctor, common problems, what the messages mean, turning a check off. | - | Living | 2026-10 | mkdocs.yml |
contributing.md |
Docs site page: Contributing links and how to build the docs site. | - | Living | 2026-10 | mkdocs.yml |
background.md |
Docs site page: Introduces the research notes (KB-*) shown under Background on the site. | - | Living | 2026-10 | mkdocs.yml |
opus-4-8-features.md |
Latest-Opus feature reference (context window, effort param, thinking, pricing) | 299 ln | Living — snapshot | Released 2026-05-28; research 2026-05-29 | platform.claude.com whats-new-claude-4-8 |
opus-4-8-swarm.md |
Multi-agent orchestration on the latest Opus; Dynamic Workflows (research preview), Managed Agents (beta) | 319 ln | Living — snapshot | 2026-05-29 | Official release notes + community (cited inline) |
KB-claude-workflow-orchestration.md |
When/how to use the Workflow tool vs a single/shallow agent — programmatic orchestration vs ad-hoc spawns, decision criteria, core patterns, ~15×/90.2%/depth-7× cost numbers, "utilize it more" guidance |
— | Living — snapshot | Compiled 2026-06-29 | 14 sources (8 official Anthropic); reconciled vs in-repo live Workflow run |
KB-codex-platform-hooks-plugins.md |
Codex platform KB — hooks, plugins, skills, MCP, AGENTS.md, config, permissions, marketplace structure; Codex equivalent for the platform parts of KB-claude-codex | — | Living — primary Codex | Compiled 2026-06-30 | Official Codex manual + local installed plugin manifests |
KB-codex-workflow-orchestration.md |
Codex workflow/orchestration KB — subagents, worktrees, cloud/noninteractive/SDK/GitHub Action workflows, OMX mapping; equivalent for Claude Workflow/swarm docs | — | Living — primary Codex | Compiled 2026-06-30 | Official Codex manual + OMX metadata |
KB-omx.md |
oh-my-codex (OMX) — Codex orchestration companion, OMC equivalent, skills/hooks/plugin/runtime mapping | — | Living — primary Codex | Compiled 2026-06-30 | OMX package metadata + official Codex docs |
KB-codex-vs-opus-coding.md |
Codex (GPT-5.5) vs Opus 4.8 for coding — benchmark consensus, per-model strengths, the "Codex=apply/Opus=think" framing (partial), recommended division of labor + cross-check pattern | — | Living — snapshot | Compiled 2026-06-29 | 13 sources (Feb–Jun 2026, 2 official); sourced not re-run |
KB-cmux.md |
cmux — terminal for AI coding agents — native macOS terminal (manaflow-ai, libghostty + socket CLI) for running Claude Code in parallel; disambiguates the separate craigsc/cmux worktree CLI; notification wiring (OSC 777 / cmux notify), teammates-as-panes, memory + orphan-process gotchas |
— | Living — snapshot | Compiled 2026-06-29 | 11 sources (4 official); WEB-only, sourced not re-run |
KB-sonnet-5.md |
Sonnet 5 + model routing — 3-way Claude benchmark tables (Opus 4.8 / Sonnet 5 / Haiku 4.5) + the parallel Codex table (migrated to gpt-5.6-sol / gpt-5.6-terra / gpt-5.4-mini per the 2026-07-09 GPT-5.6 GA — see KB-gpt-5.6.md), effort behavior, pricing, task→model decision matrix, switch thresholds, anti-hall seat routing, cross-platform equivalence |
— | Living — snapshot | Compiled 2026-07-01; corrected 2026-07-01 (default effort was wrongly stated as xhigh; is high); Codex tiers migrated 2026-07-11 | 17 sources (2 official Anthropic + 3 official OpenAI); directional (system cards unparsed) |
KB-gpt-5.6.md |
GPT-5.6 (Sol/Terra/Luna) — OpenAI's three-tier GPT-5.6 family: pricing, API/Codex CLI availability, this machine's local Codex CLI picker gap (issue #31873), third-party coding benchmarks, confidence-labeled (PRIMARY/SECONDARY/PRESS) sourcing | 223 ln | Living — snapshot | Compiled 2026-07-11 | 26 sources (6 OpenAI docs pages primary-fetched HTTP 200 + OpenAI's own GitHub release page + GitHub's own Copilot changelog); 3 OpenAI announcement pages 403'd to direct fetch, corroborated via secondary sources |
KB-token-usage-models.md |
Token usage & cost mechanics — how tokens actually BILL/CONSUME (not benchmarks): thinking/reasoning-token billing (output-rate, both platforms), full effort taxonomy incl. Haiku-no-effort-support, tokenizer effects, Workflow/multi-agent cost multipliers (official 15×/7×/4× figures), ultracode's official definition, Codex's 272K long-context premium, first-party anti-hall session telemetry | — | Living — snapshot | Compiled 2026-07-01 | 28 sources (14 official Anthropic + OpenAI); some facets (cross-effort token volume) explicitly flagged thin |
KB-model-modes.md |
Model operating modes — Opus 4.8 / Sonnet 5 / Haiku 4.5 effort levels + adaptive vs legacy thinking; Claude Code Plan Mode mechanics; the Workflow tool + "ultracode" (honesty-checked: officially documented, not rumored); /code-review ultra (ultrareview) billing/GitHub requirement; Codex/GPT-5.x reasoning-effort tiers; Codex CLI approval/sandbox modes; anti-hall routing implications (incl. a real Auditor/Critic prompt-text terminology drift found by cross-checking ship-it.workflow.js) |
— | Living — snapshot | Compiled 2026-07-03 | 47 sources (19 official Anthropic + 17 official OpenAI); 3 spot-checked sources caveated inline, not dropped |
KB-overengineering.md |
Overengineering — YAGNI/essential-vs-accidental-complexity/worse-is-better/premature-abstraction-and-optimization; AI/LLM-agent-specific causes (RLHF length-bias reward hacking, benchmark misalignment, vendor counter-guidance); empirical bloat measurement (SlopCodeBench, real-world GitHub + GitClear/DORA data); anti-hall implications for SCOPE & FIDELITY, ship-it S/M/L tiering, and debt/simplify |
— | Living — snapshot | Compiled 2026-07-03 | 15 sources (2 official vendor + 13 academic/blog/community); no second spot-check pass, disclosed |
KB-omc.md |
oh-my-claudecode (OMC) — multi-agent orchestration layer for Claude Code (plugin marketplace, skills, agents, hooks, state); canonical launch, model routing, key skills (/team, /ralph, /autopilot, /ultrawork, /ccg); the combined cmux+OMC+Claude Code stack; agnostic |
— | Living — snapshot | Compiled 2026-06-29 | OMC shipped docs + public repo; agnostic |
KB-devswarm-hivecontrol.md |
DevSwarm & the hivecontrol CLI — multi-workspace AI IDE (workspace = git worktree + agent) + its bundled, publicly-undocumented CLI (v2.3.3): full command surface (workspace/repo/health/open), .devswarm/config.json zod schema, DEVSWARM_* env + DEVSWARM_SOURCE_BRANCH-empty=Primary role detection, async message-passing coordination (create→monitor→check-merge→merge-into-source), and a 2-tier anti-hall orchestration integration (L1 Primary→child workspaces; L2 child→Workflow/subagents, no child-child; OMC+OMX parity) |
— | Living — snapshot | Compiled 2026-07-04 | 20 sources (15 official) + 5 primary local-evidence (CLI v2.3.3 exec, app.asar, live Primary+child env probe); CLI surface + role-detection signal verified-by-execution |
gsd-distilled.md |
GSD phase model distilled; the lightweight phase loop ship-it borrows from | 270 ln | Living — folded (KB-claude-codex §12) | Research 2026-05-29 | .gsd/*, gsd-*-phase SKILLs (cited in header) |
superpowers-planning.md |
Distillation of 7 superpowers skills; Iron-Law + rationalization-table pattern; minimal plan-first loop | 234 ln | Living — folded (KB-claude-codex §13) | (date unknown) | superpowers skill set (read-only study) |
keynote-prompting-claude.md |
Distilled notes from two Anthropic prompting talks (Prompting 101 + Prompting for Agents) | 267 ln | Living — folded (KB-claude-codex §9) | Captured 2026-05-29; talks dated 2025-05-22 | youtube ysPbXH0LpIE |
keynote-transcript.md |
Best-available reconstruction of the Prompting 101 talk (no verbatim transcript exists — explicitly flagged) | 342 ln | Living — reference | Talk 2025-05-22 (published 2025-07-31) | DEV recap + youtubesummary + sinyblog (cited) |
CONTEXT-PRESERVATION-KB.md |
Consolidated research KB on slowing main-agent context growth — caching, sub-agent isolation, compaction, pruning, JIT retrieval, memory externalization (12 technique families) | ~640 ln | Living — research | Swarm-synthesized 2026-06-02 | 130 selected sources (123 read) across Anthropic/OpenAI/Google docs + arXiv + practitioner sources |
KB-fable-5.md |
Fable 5 knowledge base — identity, pricing, context window, features, tier-token routing | 158 ln | Living — snapshot | Compiled 2026-06-10 | 14 sources (researcher agent draft; spot-checked) |
2026-06-10-v0.32.0-fable5-model-routing-plan.md |
v0.32.0 design plan — Fable 5 support, model-routing guard, TRIO roster update, deadly-swarm-converged | 482 ln | Historical — design plan | 2026-06-10 | This repo (main@86fb79a baseline) |
KB-claude-code-harness-features.md |
Harness feature surface vs anti-hall usage — full Claude Code hook/tool/scheduling/plugin-component enumeration cross-referenced against what anti-hall actually uses, with a gap list feeding the adoption plan; now also covers cross-session/agent-to-agent messaging (SendMessage/ListAgents/agent teams) and an evaluated-and-rejected-for-now comparison against the DevSwarm mesh |
— | Living — snapshot | Audited 2026-08-01; extended 2026-08-21 | Official code.claude.com/docs tree (incl. cross-session-messaging, agent-teams, plugin-eval) + this repo's hooks.json/skills//monitors.json |
KB-claude-code-hooks.md |
Claude Code's hook system, reference-only — factual map of the hook surface the harness exposes today (events, blocking vs advisory, additionalContext); no analysis of anti-hall's own hooks and no change proposals, deliberately out of scope for this doc |
— | Living — reference | As of 2026-09-04 | code.claude.com/docs/en/hooks |
KB-devswarm-app-db.md |
The DevSwarm desktop app's own SQLite database — anti-hall reads it, never writes; ground truth for workspace state (open/archived, titles, sidebar order, session, PR) via companion/lib/devswarm-app-db.js; consumed by the per-turn parent-inbox table, roster, identity checks, doctor, and the supervisor sync |
— | Living — snapshot | v0.108.0 | this repo's companion/lib/devswarm-app-db.js + live app DB field evidence |
KB-session-handover.md |
AI-agent session handover — what belongs in a resumable handover vs a memory file, two-tier index+detail-file structure, fixed-schema (SBAR-derived) design, failure modes that destroy continuity; backs the handover skill + handover-resume SessionStart hook |
272 ln | Living — snapshot | Compiled 2026-08-07 | 24 sources (7 official Anthropic, 6 community, 8 OSS tools, 5 GitHub issues) + clinical-handoff research |
KB-handover-research.md |
Handover research refresh — compaction loss (constraints, compaction cliff), context rot, trigger points, Claude Code + Codex hook/compaction facts (PreCompact can block; Codex SessionStart compact source), receiver read-back; gap review behind the v0.108.0 safety net, token ceiling and content rules |
196 ln | Living — snapshot | Compiled 2026-09-24 | official Anthropic + OpenAI docs, arXiv papers, clinical (I-PASS), practitioner posts |
KB-claude-monitor-tool.md |
[APPENDED 2026-08-21] Monitor tool — event-driven wake mechanics, background watcher registration, DevSwarm mailbox-delta wake-watch usage |
855 ln | Living — snapshot | (date unknown; not re-verified this pass) | this repo's monitors/monitors.json + companion/lib/devswarm-wake-watch.js |
KB-false-completion.md |
[APPENDED 2026-08-21] False task completion — reward hacking/specification gaming, claimed-vs-verified benchmark gaps, verification-before-completion mitigations | 360 ln | Living — snapshot | (date unknown; not re-verified this pass) | 21 sources per §3 topic-map entry |
KB-goal-setting.md |
[APPENDED 2026-08-21] Goal-setting / acceptance-criteria wording — classical theory + AI goal-misspecification as a reward-hacking root cause | 372 ln | Living — snapshot | (date unknown; not re-verified this pass) | see §3 topic-map entry |
CODEX-KB-MIGRATION-MAP.md |
[APPENDED 2026-08-21] Codex KB migration map — cross-reference between Claude-side and Codex-side KB docs | 32 ln | Reference | (date unknown; not re-verified this pass) | this repo |
2026-06-06-context-opt-test-design.md |
[APPENDED 2026-08-21] Context-optimization test design — dated design artifact | 65 ln | Historical — design doc | 2026-06-06 | this repo |
Reading order for a new contributor: KB.md → KB-claude-codex.md (the
synthesis) → the topic doc you need. The keynote-* and superpowers/gsd
docs are background; their actionable content is already in the synthesis.
3. Topic → doc map (where each subject lives)¶
| If you're working on… | Read |
|---|---|
A new or changed hook (event taxonomy, blocking vs injection, additionalContext gating) |
KB-claude-codex §1; current state in §1 |
| plugin.json / marketplace / version precedence | KB-claude-codex §2; §1 version row |
| Prompting the protocol text / nudges | KB-claude-codex §3, §6, §9; keynote-prompting-claude; superpowers-planning (Iron-Law form) |
| Codex / AGENTS.md governance | KB-claude-codex §5; AGENTS.md at root |
| Codex platform/plugin/hook porting | KB-codex-platform-hooks-plugins; KB-claude-codex for Claude contrast |
| Codex workflow/swarm equivalents | KB-codex-workflow-orchestration; KB-omx; CONTEXT-PRESERVATION-KB for context rationale |
| OMX / oh-my-codex | KB-omx |
| Orchestration / swarm / subagents | KB-claude-codex §7, §11; opus-4-8-swarm |
| Slowing main-agent context growth (caching, sub-agent isolation, compaction, pruning, JIT retrieval, memory externalization) | CONTEXT-PRESERVATION-KB.md — consolidated research KB |
| deadly-loop / ship-it phase model + debate roster | KB-claude-codex §12, §13; gsd-distilled; superpowers-planning; skills/MODEL-POLICY.md |
| DevSwarm multi-workspace orchestration (driving the hivecontrol CLI; workspace-role detection; async parent/child coordination; making anti-hall orchestrate across workspaces vs in-process subagents) | KB-devswarm-hivecontrol (hivecontrol v2.3.3 command surface + .devswarm/config.json schema + DEVSWARM_*/sourceBranch role signal + 2-tier L1-workspaces/L2-subagents integration design, OMC+OMX) |
| DevSwarm-aware workspace-tier orchestration (the original design that made anti-hall orchestration workspace-topology-aware — Workspace-Primary/-Child role detection, devswarm-guard, the create→poll→serialized-merge→surface loop) | docs/archive/superpowers/specs/2026-07-05-devswarm-orchestration-design.md (approved design) + docs/archive/superpowers/plans/2026-07-06-devswarm-orchestration.md (implementation plan); related: KB-devswarm-hivecontrol |
| DevSwarm liveness supervisor (the opt-in, OPTIONAL companion that recovers a wedged/idle DevSwarm workspace session — workaround for claude-code#39755) | docs/archive/superpowers/specs/2026-07-08-devswarm-liveness-supervisor-design.md (design) + docs/archive/superpowers/plans/2026-07-08-devswarm-liveness-supervisor.md (implementation plan); related: KB-devswarm-hivecontrol |
| Harness-feature adoption plan (phased build plan derived from the harness-vs-usage gap list; Fable-reviewed 2026-08-02) | docs/archive/superpowers/specs/2026-08-01-harness-feature-adoption.md; companion: KB-claude-code-harness-features |
| Goal-setting / acceptance-criteria wording for plans, phases, and task prompts (classical theory, AI goal-misspecification as a reward-hacking root cause, cross-vendor task-specification practice) | KB-goal-setting (Locke & Latham + SMART + Definition of Done + OKRs; DeepMind/Anthropic/METR/OpenAI reward-hacking evidence; Claude Code + Codex "Done when" guidance; concrete ship-it Step 1/2 goal-drift gap finding) |
| Task discipline (tasklist-guard) | TASKLIST-GUARD (usage); TASK-WORK (design/research) |
| Testing the hooks | E2E-TESTING |
| Model selection / effort / thinking (which model when: Opus 4.8 / Sonnet 5 / Haiku 4.5 + Codex gpt-5.x) | KB-sonnet-5 (benchmark tables + decision matrix + switch thresholds + cross-platform, migrated to the gpt-5.6-sol/terra tiers); KB-gpt-5.6 (the GPT-5.6 Sol/Terra/Luna family itself — pricing, availability, sourcing); opus-4-8-features; skills/MODEL-POLICY.md |
| Model operating modes (effort-level behavior, Plan Mode, Workflow tool/"ultracode", ultrareview, Codex CLI approval/sandbox modes) | KB-model-modes (all 4 Claude Code product surfaces + both platforms' effort taxonomies + anti-hall routing implications) |
| Token billing / cost mechanics / effort-tier taxonomy / Workflow orchestration cost | KB-token-usage-models (thinking-token billing, all effort tiers both platforms, tokenizer effects, 15×/7×/4× multi-agent multipliers, ultracode definition, first-party session telemetry) |
| Overengineering / scope creep (why it happens, AI-agent-specific causes, measured bloat, SCOPE & FIDELITY / ship-it / debt / simplify implications) | KB-overengineering |
| Anti-hallucination evidence base (peer-reviewed) | KB-claude-codex §8 + "Design implications" |
| False task completion (reward hacking / specification gaming, claimed-vs-verified benchmark gaps, verification-before-completion mitigations, the STATE.json enforcement gap) | KB-false-completion (21 sources: reward hacking/scheming research, claimed-vs-verified benchmark gap, mitigation patterns, anti-hall implications) |
| [APPENDED 2026-08-21] Monitor tool / event-driven wake (background watcher registration, DevSwarm mailbox-delta wake-watch) | KB-claude-monitor-tool (855 ln); usage in monitors/monitors.json + companion/lib/devswarm-wake-watch.js |
| [NEW 0.78.0] Defect channel (agent-filed bug reports against anti-hall itself, maintainer rulings, derived-state/no-index/write-verified design) | No dedicated KB doc yet — component reference in Plugin README §Features (defect-nudge.js, hooks/lib/defect-store.js, scripts/defect.js); design rationale in the feat(defects) commit message (604314d) |
| [NEW] Jev classifier (TypeSafe System One opt-in classifier backend for speculation-guard.js, regex fallback, default OFF, benchmark evidence, fallback semantics, data/privacy, observability) | KB-jev-classifier.md |
4. Staleness ledger¶
Suspected-stale or code-contradicting claims found in the living docs, flagged for review. Per the freshness rule, these are flagged, not silently rewritten — a maintainer should verify and either fix the source doc or confirm it's fine. Format:
doc:line — claim — why suspect — current truth.
Living docs — open flags:
opus-4-8-features.md:5— Header pinsModel ID: claude-opus-4-8. Why suspect: a hardcoded model ID dates fast and can read as policy. Status: acceptable as a dated snapshot — the doc header carriesReleased: 2026-05-28andResearch date: 2026-05-29, andMODEL-POLICY.mdresolves "latest" at runtime, so policy is NOT pinned. Keep as a snapshot; do not cite this ID as the model to use.opus-4-8-swarm.md:5— "Opus 4.8 was the latest at time of writing … always use the newest available." Why noted: correctly framed already — left as a model for how snapshots should self-date. No action.TASK-WORK.md(Overview) — "Task tools became the default in Claude Code v2.1.142 / TS Agent SDK 0.3.142." Why suspect: a pinned client version that may have moved. Status: unverified against current Claude Code; treat the version numbers as historical, the behavior (Task tools default, setCLAUDE_CODE_ENABLE_TASKS=0to fall back) as current. Verify before quoting the exact version.KB-claude-codex.md:22— "official docs … enumerate ~28–32 events; community says 27+." Why noted: the doc already flags this as version-dependent and declines to pin it. No action — this is the correct handling of a moving count.KB-claude-codex.md:211— "Empirically (librarian v0.6.0) 7 rounds caught 30+ bugs." Why noted: references an external project by name/version as provenance for the deadly-loop origin. The claim is anecdotal-but-cited; keep, but it is not a plugin fact.keynote-transcript.md:1— "No verbatim transcript is publicly available … close reconstruction, not verbatim." Why noted: the doc self-flags as reconstruction. Correct handling; no action — do not treat its quotes as exact.
[APPENDED 2026-08-21] 9-doc staleness audit (this pass):
KB.md(this file) — stale items found and FIXED this pass: §1 provenance date (:51, was 2026-06-04), version row (:68, was0.71.1), hook count (:71, was 46 files), skill count (:72, was 12, missingdevswarm/handover/system-briefing), §2 living-doc table missing 8 docs (now appended), §4 ledger itself missing 2026-07+ entries (now appended, this block), §5:213/:216/:235stale version citations, §6:246"13 docs" claim. NOT fixed this pass (flagged NEEDS-CONFIRMATION, out of scope): thePostToolUse/additionalContextdelivery claim at:74. [UPDATE 2026-08-22: this item is now resolved — see the[CORRECTION 2026-08-22]annotation on the "Message-bloat prevention" row (live-observed on CLI 2.1.238;PostToolUsedelivers advisoryadditionalContextbut still cannot block).]docs/KB-claude-code-harness-features.md— stale items found and FIXED this pass:PostToolUse/PostToolUseFailuremis-listed as NOT USED (:38, now corrected — both ship since v0.69.0),plugin.jsonversion cited as v0.68.2 (:37), skill counts 14/17 (:33, should be 15/18,handovermissing), the "absent from every prior doc" header claim (:177, false for theisolation:"worktree"row),claude plugin evalmis-labeled "Org-flag gated" (:184, actually available locally). NOT re-verified this pass: thellms.txt404 and the unfetched-docs list (:134-135), now marked UNVERIFIED-since-2026-08-01 rather than asserted current.KB-claude-codex.md— CONTENT not edited this pass (out of scope per this session's instructions — a later wave handles model KBs and this backbone doc). Only its ledger presence here is new. Known live issue already self-documented inline at:47via its own[UPDATE …]annotations (PreToolUse/SubagentStart delivery confirmations) — no new staleness found or claimed against it this pass.KB-claude-monitor-tool.md— CONTENT not edited this pass (out of scope). Newly indexed in §2/§3 this pass (was previously invisible from the index despite being 855 lines). Not independently re-verified for staleness.KB-model-modes.md— CONTENT not edited this pass (out of scope, later wave). Routes model-selection guidance against Opus 4.8 per its header. [UPDATE 2026-08-22: RESOLVED — see the 2026-08-22 model-KB re-audit block below.]KB-sonnet-5.md— CONTENT not edited this pass (out of scope). Routes against Opus 4.8 / Sonnet 5 / Haiku 4.5 per its header. [UPDATE 2026-08-22: RESOLVED — see the 2026-08-22 model-KB re-audit block below.]KB-fable-5.md— CONTENT not edited this pass (out of scope). [UPDATE 2026-08-22: RESOLVED — see the 2026-08-22 model-KB re-audit block below.]KB-token-usage-models.md— CONTENT not edited this pass (out of scope). [UPDATE 2026-08-22: RESOLVED — see the 2026-08-22 model-KB re-audit block below.]KB-codex-vs-opus-coding.md— CONTENT not edited this pass (out of scope). Cross-cutting finding, NOT resolved this pass:Opus 5/claude-opus-5has 0 occurrences repo-wide (grepped this session), while every model KB above routes its Reviewer/Auditor guidance against "Opus 4.8." This is flagged NEEDS-CONFIRMATION — it is unclear whether "Opus 5" is simply not yet released/available, a naming-scheme assumption that doesn't hold, or a genuine gap in these docs. Do not resolve this here; a later wave owns the model-KB content pass. [UPDATE 2026-08-22: RESOLVED — confirmed real:claude-opus-5is Opus 4.8's actual successor (shipped 2026-06-09, same $5/$25 price/context class), not a naming-scheme guess. See the 2026-08-22 model-KB re-audit block below.]
[APPENDED 2026-08-22] Model-KB re-audit (this pass): the six model/routing
docs (KB-model-modes.md, KB-sonnet-5.md, KB-fable-5.md,
KB-token-usage-models.md, KB-codex-vs-opus-coding.md,
plugins/anti-hall/skills/MODEL-POLICY.md + its Codex mirror) were re-audited
against a verified current model lineup and corrected via stacked dated
annotations ([UPDATE 2026-08-22] / [CORRECTION 2026-08-22]), not silent
rewrites — each file's own correction convention was followed (or, where none
existed, the same stacked-annotation pattern as this ledger). Findings:
Opus 4.8 is now DEPRECATED / legacy, superseded by Opus 5
(claude-opus-5, same $5/$25 price/context class, shipped 2026-06-09); the
true current flagship is Claude Fable 5 (claude-fable-5, $10/$50,
1M/128k) with Claude Mythos 5 (claude-mythos-5) gated to approved orgs —
neither had any mention across KB-model-modes.md or KB-sonnet-5.md before
this pass. Sonnet 5's scheduled Sept-1-2026 price increase to $3/$15 was
cancelled — it holds at $2/$10 indefinitely, contradicting the "intro,
→Aug 31 2026" framing in KB-sonnet-5.md §2 and KB-token-usage-models.md
§1. KB-codex-vs-opus-coding.md additionally predates the 2026-07-09 GPT-5.6
(Sol/Terra/Luna) migration on the Codex side. MODEL-POLICY.md (both Claude
and Codex variants) needed no correction — both already route by tier
token (opus/sonnet/haiku/fable), resolved to the newest model
in-family at runtime, per each file's own stated "never pin a model version"
rule; only the prose docs had hardcoded stale model names. Also added: a
"downshift guidance" section (KB-model-modes.md §13, cross-linked from
KB-sonnet-5.md §6) covering main-agent usage-limit conservation — Sonnet 5
is the correct 1M-context downshift target; Haiku 4.5 (200k) is disqualified
for the main agent though still correct for leaf subagent work. The
plugins/anti-hall/hooks/lib/repo-audit-baseline.js MODEL_KB_AUDIT_DATE was
bumped from 2026-05-29 to 2026-08-22 to reflect this re-audit (threshold
unchanged at 60 days).
[APPENDED 2026-09-03] Fable point-release re-audit: claude-fable-5-1
(released 2026-09-01) supersedes claude-fable-5 as current Fable — a point
release, not a new generation, so Opus 5 / Sonnet 5 / Haiku 4.5 lineage above
is unaffected. docs/KB-fable-5.md, docs/KB-sonnet-5.md,
docs/KB-model-modes.md, docs/KB-codex-vs-opus-coding.md, and
docs/KB-token-usage-models.md were corrected via stacked dated annotations
([UPDATE 2026-09-03]), not silent rewrites. Pricing/context/effort-scale
specifics for claude-fable-5-1 are not verified in this repo — flagged in
each updated doc rather than assumed unchanged from Fable 5's $10/$50, 1M/128k.
MODEL-POLICY.md (both Claude and Codex variants) needed no correction —
already routes by tier token (fable), resolved at call time.
plugins/anti-hall/hooks/lib/repo-audit-baseline.js MODEL_KB_AUDIT_DATE
bumped 2026-08-22 → 2026-09-03. No routing logic found hardcoding a Fable
model id; the sole hardcoded-model-id exception repo-wide remains
speculation-judge.js's direct Anthropic Messages API call
(claude-haiku-4-5, unrelated to Fable, unchanged this pass).
Historical docs — pre-redesign claims (do NOT fix; they are frozen records): These describe the plugin before the cadence redesign and are intentionally stale. Listed here so no one mistakes them for current spec.
PLUGIN-REVIEW.md:14,29,57,90,113etc. — referenceshooks/verify-first.sh,git-guard.sh,graphify-session.sh. Current: all hooks are.js.PLUGIN-REVIEW.md:22–25— "No PreCompact re-injection … add a PreCompact hook." Current: resolved differently — SessionStart re-fires oncompact; no PreCompact hook (and the review's own §1.2 in KB shows PreCompactadditionalContextis inert).PLUGIN-REVIEW.md:34— "No AGENTS.md mirror → plugin is Claude-only." Current:AGENTS.mdexists at repo root.PLUGIN-REVIEW.md:46— "ships five strong skills." Current: 29 skills [UPDATE 2026-08-21: corrected from 12 — see §1 "Skills shipped" row] (ls plugins/anti-hall/skills/; see §1 "Skills shipped" row).ULTRAPLAN.md:71,316,341— "rotates 5 one-liners" / "spread across the 5 nudges." Current: 20 nudges (% NUDGES.length).ULTRAPLAN.md:184—version: 0.3.0, bump to0.4.0. Current: [UPDATE 2026-08-21: corrected from0.21.0— that was itself stale] committed0.75.1; working tree0.76.0(release in flight, uncommitted).AUDIT-REPORT.md:123,129,132/AUDIT-REPORT-2.md— version0.7.0,0.11.x, "Codex GPT-5.5" pin. Current: all superseded; these record the fixes as applied at the time. The "12 entries" note atAUDIT-REPORT.md:132is correct and matches current code.
5. History — historical artifacts¶
One-time records: dated audits, the plugin review, and the consolidated plan. Frozen. Never edited to match current code — their value is the timestamped snapshot of what was true and what was decided then.
| Artifact | What it is | Date / version context | Status now |
|---|---|---|---|
AUDIT-REPORT.md |
4-auditor review (2 Opus + 2 Codex); confirmed issues + fixes | 2026-06-01 (mtime); v0.7.0-era |
Superseded; findings applied |
AUDIT-REPORT-2.md |
Double deadly-loop final gate; sudo-bypass fix et al. |
2026-06-01; v0.11.1 → v0.11.2 |
Superseded; findings applied |
PLUGIN-REVIEW.md |
KB-driven plugin audit (P0–P2); the doc that prescribed the cadence redesign (Iron-Law form, SessionStart primacy, AGENTS.md, skills primer) | 2026-06-01; pre-redesign (.sh hooks, 5 nudges) |
Superseded — its P0s are now shipped (see §4) |
ULTRAPLAN.md |
Single consolidated reconciliation plan; planning artifact only | 2026-05-31; v0.3.0-era |
Superseded — executed; resting state is [UPDATE 2026-08-21: corrected from 0.20.3] committed 0.75.1 (working tree 0.76.0, release in flight, uncommitted) |
archive/devswarm-layered-recovery-history.md |
DevSwarm layered-recovery version history (v0.54–v0.107); moved verbatim out of GUIDE.md in v0.108.0 so the guide reads as current state |
Moved 0.108.0; content spans v0.54–v0.107 | Frozen — dated changelog/defect narrative, never edited to match current code; current behavior lives in GUIDE.md/KB-devswarm-hivecontrol.md/KB-devswarm-app-db.md/CHANGELOG.md |
Origin note: the deadly-loop discipline that anti-hall ships as a skill was
born from a real 7-round iteration that caught 30+ bugs solo review missed
(KB-claude-codex §6/§7, cited there). AUDIT-REPORT*.md are that discipline
applied to anti-hall itself — dogfooding.
5.1 Per-version ground-truth history (moved from §1)¶
Moved verbatim out of the §1 table in 0.108.0 so §1 reads as current state. Frozen.
Version row (0.58.0–0.71.1 narrative):
v0.71.1: DevSwarm ingest daemon busy-spin fix — the success-path loop now sleeps elapsed-aware (~1 iteration per intervalSec, clamped to MAX_PACE_MS) instead of spinning as fast as the OS would schedule it, stopping a data.kalloc.1024 macOS kernel-allocator leak and ~11% idle CPU (~362x fewer fork/exec spawns per second in a controlled harness). Delivered to running installs: the daemon stamps its plugin codeVersion into the heartbeat, doctor-repair restarts an alive daemon found running stale code, and the updater force-restarts the ingest daemon after update.js so an already-running unbounded-loop daemon re-execs onto the paced code. See CHANGELOG.md 0.71.1 for full detail. v0.71.0: Append-only reply-state redesign — recordReply (companion/lib/devswarm-reply-state.js) moved from a lockfile read-modify-write of a merged JSON object to an append-only JSONL log (one O_APPEND write, no lock); readReplyState now folds the log on read, with a fail-closed newline separator against partial records, a 480-byte per-record cap, and an Object.create(null) fold accumulator so a __proto__-named sender survives the fold instead of polluting the prototype. Ships with a loss-safe forward migration (migrateReplyState) wired into both update.js and doctor-repair.js, with an accepted, documented residual in the final write window; structurally eliminates the disclosed steal-branch TOCTOU rather than patching around it. Rule K's emoji-as-signal wording (from 0.70.0) is now ALSO injected at SubagentStart (verify-first-subagent.js) and in the Codex orchestration skill, not just the orchestrator's SessionStart. Test-store-leak hardening: 4 doctor tests' HOME-default landmine (an unset override silently fell back to the real home) fixed; a new READ-ONLY store audit/classifier (REAL/GARBAGE/UNKNOWN) plus a leak-report CLI ships whose --out is guarded (realpath canonicalization against the store root, O_EXCL/O_NOFOLLOW write, a distinctive report marker) so it can never overwrite a production devswarm.db; ambiguous/unreadable evidence degrades to UNKNOWN, never GARBAGE — detection-only, never deletes. register-primary's --session now defaults to CLAUDE_CODE_SESSION_ID (was the workspace hash), so Primary registry rows resolve their transcript for liveness reads; a KB doc's broken $CLAUDE_SESSION_ID reference was also corrected. See CHANGELOG.md 0.71.0 for full detail. v0.70.0: DevSwarm mesh/store hardening — foldArchivedRegistryRows (scripts/devswarm.js) now folds ALL same-worktree registry rows sharing an archived id and picks the forward survivor by LIVENESS, fixing a message-loss P0 where a real question could forward into a dead partition; ships as a dual-path migration wired into BOTH update.js and doctor --fix's migrationFix('fold-archived-rows', ...) (idempotent/fail-open-honestly/no-delete). A new read-side filter (archivedOnlyIds, companion/lib/devswarm-store.js) excludes archived-only workspaces from the live per-turn projection WITHOUT needing a doctor run — archived+unread still surfaces as an orphan, no lost signal; cmdArchive gained an inode-decided descriptor-conflict self-heal unblocking re-archive of a stale archived/<id>.json leftover; the per-turn parent-inbox STOP imperative is now advisory wording for the normal tier (loud tier unchanged). Plus 5 doctor/limit test files hardened against contention-killed subprocesses (no assertion weakened) and the orchestration rule K emoji-signal glyphs (✅/❌/⚠️) made explicit. See CHANGELOG.md 0.70.0 for full detail. v0.69.0: Harness Phase-1 hooks output-verify-guard.js (PostToolUse/Bash, advisory) and failure-root-cause-nudge.js (PostToolUseFailure/Bash, advisory) plus a sandbox doc section in AGENTS.md — from the Fable-reviewed docs/archive/superpowers/specs/2026-08-01-harness-feature-adoption.md. DevSwarm parent decide+reply gate: devswarm-parent-gate.js's Stop-gate now requires an OBSERVED reply, not just a read question — a new devswarm-parent-reply-tracker.js (PostToolUse/Bash, Primary only) records a successful send --to <id> --question via a new durable per-project reply-state store (companion/lib/devswarm-reply-state.js, repoKey-keyed, lock-protected recordReply); the forced-ack cap now escalates once instead of going silently quiet. New persisted shape needs_reply on mesh rows ships an additive/idempotent/fail-open/no-delete ALTER TABLE ADD COLUMN migration in devswarm-store.js's ensureMessagesMeshColumns, applied on every store open (covers both the update/migrate-state.js path and doctor's self-tests, which both transitively open the store). DevSwarm child self-continue directive in devswarm-child-turn.js tells a child to keep working across rounds of a multi-round task instead of idling between them. Windows support DROPPED from the CI matrix and from support claims (macOS + Linux only now; pure-Node code unchanged, Windows untested). NOTE (staleness disclosure): this row's exhaustive per-version prose was last fully maintained at 0.67.1 — 0.68.0/0.68.1/0.68.2's changes (deliverable direct messages read fix, doctor install-divergence detection + full-tree walk) are NOT individually narrated below; see CHANGELOG.md for those. This release only bumped the version number + appended the note above rather than backfilling the missing historical detail — flagged as a known gap, not silently dropped. plugin.json version — the single authority. Marketplace entry carries NO version (avoids the silent-precedence trap). v0.58.0 (released): DevSwarm mesh-only messaging — the per-project mesh store (repoKey primitive in companion/lib/devswarm-repokey.js, shared store at store/<repoKey>/ replacing the per-worktree store) is now the SOLE agent-initiated messaging transport, mechanically enforced by command-guard blocking hivecontrol message-child/message-parent (dequoted to shell-effective argv, closing quote-based bypasses) plus a per-turn communication override; every other hivecontrol feature (create/list/check-merge/merge) is kept and thinly wrapped. Includes the mesh CLI (send --to <meshId>|--to-primary|--broadcast [--urgency]/roster [--ack]/mesh read/heartbeat --summary on scripts/devswarm.js), a ONE-per-project ingest daemon with reap-before-drain of legacy per-worktree units, a doctor orphan-sweep + rollback path, a non-destructive hash→repoKey store migration, and the #36-STRUCTURAL cross-project scoping fix in devswarm-parent-gate.js/devswarm-parent-inbox.js. Supervisor escalates (never kills) on urgent/high unread; heartbeat sender identity is ownership-validated; emitted coordination commands use an absolute CLI path. DevSwarm coordination remains entirely OPTIONAL — dormant with zero behavioral change outside a DevSwarm session. Codex: the guard-block is shared and fires on Codex; the five per-turn override/reassert hooks (devswarm-child-role.js, devswarm-child-turn.js, devswarm-parent-inbox.js, devswarm-parent-gate.js, devswarm-child-gate.js) are now also registered, unmodified, in codex/hooks/hooks.json (corrected — an earlier claim that they were Claude-only rested on a since-disproven assumption that their gating DEVSWARM_* env vars were Claude-specific; see §8.7). The liveness supervisor and on-demand recovery CLI target remain genuinely Claude-only (identity-bind to claude --resume processes) — do not describe the Codex port as mesh-capable beyond the guard-block + these five hooks. See docs/KB-devswarm-hivecontrol.md §8.7 for the full reference. Note: this row is kept current at each release; the hook enumeration below was backfilled 2026-07-03 (was 0.39.0-era, missing fable-availability.js/progress-prune.js/session-history-index.js) to the then-current 32-.js-file ground truth (task #18), and updated again 2026-07-11 to the current 36-.js-file ground truth (edit-guard.js/coordinator-detect.js/codex-availability.js/devswarm-child-role.js added since). 0.54.1 modified 3 EXISTING hook files' behavior (devswarm-child-gate.js over-nag fix, devswarm-child-turn.js unread-inbox surfacing, devswarm-parent-inbox.js live workspace table) and added one new companion installer (companion/install-devswarm-ingest.js) — no new hook files, so the 40-.js-file count is unchanged. 0.55.0 added doctor repair mode: a new shared module hooks/lib/doctor-repair.js (like devswarm-detect.js, in hooks/lib/ — NOT counted in the top-level hooks/ enumeration) plus doctor.js flag parsing (--check/--fix/--repair/--dry-run) AND a new top-level hook hooks/inbox-read-guard.js (PreToolUse Read, Claude-only, blocks a direct Read-tool open of a raw DevSwarm inbox/store file) — bumping the file count to 41 .js files + hooks.json = 42 (updated below). Plain doctor now diagnoses AND repairs: AUTO-SAFE fixes (state migrations, statusline-if-missing, idempotent supervisor/codex refresh) always; GATED daemon fixes (ingest install / v0.54.1 wrong-path rebind / stale-script / supervisor first-install) only when isDevswarmActive(env) AND resolveWorktree(cwd) is a git worktree, else it reports the exact manual command. --check is the pure read-only path (CI); the test suite was re-pointed to --check. 0.56.0 shipped NO new top-level hook files — the archive-request send-only verb, the ack-ownership guard (callerIdentity + --ack-as-owner), the Primary own-unread gate/inbox parity, the DEVSWARM_BUILDER_ID heartbeat-key fix, and ingest-daemon durability (healIngestDaemon) all landed inside existing scripts/devswarm.js / hooks/devswarm-parent-gate.js / hooks/devswarm-parent-inbox.js / hooks/devswarm-child-turn.js / skills/update/scripts/update.js — so the 41-.js-file count (42 incl. hooks.json) is unchanged from 0.55.0. 0.58.1 shipped NO new top-level hook files — a hard timeout backstops the ingest daemon's hivecontrol workspace monitor spawn (cooperative -t + 10s, fixing an indefinite wedge observed as 4,319 consecutive lock refusals over ~15h) and SIGTERM/SIGINT now release the ingest lock before exit (companion/devswarm-ingest.js); reconcile now runs automatically as a gated doctor repair (hooks/lib/doctor-repair.js) and as a post-update step (skills/update/scripts/update.js) instead of requiring a manual command, remaining idempotent and worktree-safe. v0.67.0: human-readable DevSwarm workspace names — spawn sets a title after create via a SEPARATE best-effort hivecontrol workspace update-title -b <branch> "<title>" call derived from the -p brief (a caller-supplied -t/--title is never overridden); spawn's pass-through argv to workspace create stays untouched; new shared fs name cache companion/lib/devswarm-names.js (atomic tmp+rename, fail-open read — a companion lib, NOT a top-level hook, so the hook file count is unchanged) is written by devswarm.js and read by devswarm-parent-inbox.js, whose per-turn table now renders name (shortid) instead of a bare UUID and never spawns hivecontrol on that hot path; reconcile caches hivecontrol's existing label for pre-existing workspaces but never invents one. Review-seat integrity (BEHAVIOR CHANGE): ship-it's per-phase gate previously ignored dead seats, so fewer live seats produced fewer findings and a silently PASSING converged:true; the gate now carries totalSeats/liveSeats/deadSeats/degraded/seatReports and requires deadSeats === 0, and also honors args.codexAvailable mirroring deadly-loop (incl. the Opus adversarial-persona fallback) — codex-availability.js now instructs the coordinator to thread its result into Workflow invocations, since workflow scripts have no fs access and cannot read the JSON themselves; three docs claiming an "enforced codexUp probe" (none existed) were corrected. Model routing is now version-agnostic: pinned model versions removed across skills/hooks/docs/Codex port in favor of tier tokens (opus/sonnet/haiku/fable), plus a standing MODEL-POLICY rule never to pin a version — sole exception speculation-judge.js's direct Messages API call (requires an exact id, no "-latest" alias; overridable via ANTIHALL_JUDGE_MODEL). Fixes: a raw NUL byte in scripts/devswarm.js replaced with the \x00 escape (grep no longer treats the 245KB file as binary); a hasFlag redeclaration collision that broke --yes/--confirm detection in reconcile-active/reap-stale. v0.67.1: four stacked fixes to the DevSwarm supervisor escalation path, which had never delivered end-to-end — a missing deriveSummary call after the append (the parent-facing projection stayed stale), openStore called without a hash (wrote to a legacy bucket instead of the repoKey store), parentId derived from the CHILD's own worktree path rather than resolved via resolveMainWorktree before primaryWorkspaceId (escalations landed in the child's own bucket), and two fold/rehome paths that skipped their projection refresh. A guarded require of devswarm-repokey.js (companion/lib/recovery.js:49-56) now protects devswarm-supervisor.js:41 and scripts/devswarm.js:129 — a throwing module there can no longer crash both consumers before their own fail-open engages. Cross-repo hijack hardening on the roster's native fold (fe1b987, NOT a repair of a dead feature — fetchNativeChildren passed env unmodified in every prior shipped release): hivecontrol workspace list children resolves its scope entirely from DEVSWARM_REPO_ID/DEVSWARM_BUILDER_ID, never cwd, so a process holding a foreign repo's env got that repo's children back with exit 0 and valid JSON; each record's repositoryId is now cross-checked against a separate cwd-anchored, env-stripped list all ground truth, with mismatches dropped + logged and the fold failing open unfiltered when no ground truth is available. New devswarm.js skip <guard> [--ttl N] CLI verb (cmdSkip, devswarm.js:2637) plus an edit-guard skip-key fix. Known remaining gap: projection delivery additionally requires the Primary to have self-registered from the true main worktree — otherwise the escalation lands in orphans[] (companion/lib/devswarm-store.js:1405-1424: a partition with real unread and no live registry row goes to orphans[], never workspaces[]). The informational devswarm-parent-inbox hook surfaces orphans; the blocking devswarm-parent-gate hook does not (grep -c orphans on that file = 0), so a Primary can Stop unblocked on an escalation the inbox hook is displaying.
Hooks row (per-hook history):
agent-watchdog, api-guard, codex-availability, codex-nudge, command-guard, coordinator-detect (shared module, not a hook), devswarm-child-gate, devswarm-child-role, devswarm-child-turn, devswarm-parent-gate, devswarm-parent-inbox, devswarm-parent-reply-tracker, doctor, edit-guard, fable-availability, failure-root-cause-nudge, git-guard, inbox-read-guard, limit-conserve-inject, limit-conserve (shared helper, not a hook), merge-gate, model-routing-guard, omc-detect, output-verify-guard, phase-tracker, progress-prune, ship-it-guard, skip-guard, speculation-guard, speculation-judge, swarm-guard, task-guard, task-tracker, tasklist-guard, session-history-index (shared module, not a hook), verify-first-core (shared module, not a hook), verify-first-full, verify-first-orch, verify-first-subagent, verify-first, version-alert, version-alert-refresh, plus hooks.json (45 .js files + hooks.json = 46). v0.69.0 additions: devswarm-parent-reply-tracker (PostToolUse/Bash, Primary only) — observes a successful devswarm.js send --to <id> --question and records it via companion/lib/devswarm-reply-state.js (a new companion lib, not itself a top-level hook) so the Stop-gate can tell "read" apart from "replied". output-verify-guard (PostToolUse/Bash, advisory) and failure-root-cause-nudge (PostToolUseFailure/Bash, advisory) — Harness Phase-1 hooks, both fail-open/never-block. verify-first-orch (0.60.0) = SessionStart hook, the companion to verify-first-full: it carries the ORCHESTRATION DISCIPLINE ruleset (rules A–N + DevSwarm-Primary rule W) that was SPLIT off verify-first-full because the combined ~15.3k-char payload exceeded the ~10k per-hook injection cap (which spills the overflow to a file — see the SessionStart-injection-footprint row). Each half now lands 100% inline; the split deleted zero content and reverted two cap-era workarounds (rule L rehoisting, rule W squeezing). Registered on both the Claude plugin and the Codex port. inbox-read-guard (0.55.0) = PreToolUse Read, OPTIONAL/feature-gated, Claude-only (not registered on the Codex port): blocks a direct Read-tool open of a raw DevSwarm inbox/store file via the same devswarm-inbox-paths.js classifier command-guard.js uses — inbox/** denied unconditionally, the store db/journal denied once a Primary read path exists (inbox messages — shipped), everything else under the DevSwarm root allowed. Fail-open; own devswarm-read-guard skip name. fable-availability (0.43.0) = SessionStart hook: reads ~/.claude.json's modelAccessCache/additionalModelOptionsCache (the same entitlement cache Claude Code's own /model selector uses) ONCE per session, fail-open, silent unless Fable 5 is available; sets args.fableAvailable for ship-it/deadly-loop Workflow invocations. Per 0.43.2, Fable routing was policy-disabled for the Reviewer seat (over-restrictive/refusal-prone per community feedback) — the hook and its cache stayed in place for visibility only, not acted on. RE-ENABLED as of 2026-07-12 (owner call, Fable 5 now available): the Reviewer seat once again tries Fable first when args.fableAvailable === true, falling back to Sonnet 5 then Opus. progress-prune (0.43.0) = SessionStart hook, per-cwd 24h-throttled: archives stale per-session progress files (past UTC-date folder, mtime >6h stale — a session still running across midnight is never touched mid-flight) by appending the file's full content into that session's own history ledger under an "Archived progress" heading before deleting it; never deletes if the archive append fails. session-history-index.js = shared module (not a hook, not registered in hooks.json), exports appendIndexLineIfAbsent(): single-line atomic idempotent append (fs.appendFileSync with 'a', never read-modify-rewrite) used to maintain the per-session progress/history INDEX.md files; consumed by tasklist-guard.js. verify-first-subagent (0.39.0) = SubagentStart hook: re-injects the Iron Law + rationalization table + positive rules + scope-fidelity into each spawned subagent; deliberately omits the orchestration/delegate block (subagents are workers; re-injecting it recreates deep nesting). verify-first-core.js = shared module (not a hook) that is the single source of truth for the Iron Law content shared by verify-first-full.js and verify-first-subagent.js — prevents drift between the two hooks. (SubagentStart is a confirmed Claude Code event; see KB-claude-codex.md §1.1.) limit-conserve-inject (0.38.0) = UserPromptSubmit, limit-conservation mode: injects a token-conservation nudge when context usage ≥ ANTIHALL_LIMIT_THRESHOLD (default 85). ANTIHALL_LIMIT_CONSERVE: auto (default, reads OMC usage cache at ~/.anti-hall/omc-usage-cache.json) / on / off. Auto requires OMC; without it, operates in manual on/off mode only. Skip-guard hatch: limit-conserve. limit-conserve.js = shared helper (not a hook) consumed by limit-conserve-inject.js (reads OMC usage cache, applies threshold logic). version-alert (0.37.0) = SessionStart, NON-BLOCKING: reads running version vs ~/.anti-hall/version-check.json; if behind, emits a one-line "vX available — /anti-hall:update"; when cache absent/stale spawns a DETACHED unref'd version-alert-refresh.js (git ls-remote --tags) — SessionStart never blocks or does synchronous network. Off-switch ANTIHALL_VERSION_ALERT=off; skip-guard hatch. model-routing-guard (0.32.0) = PreToolUse Agent/Task, anti-waste routing: classifies spawn descriptions (mechanical vs complex) and blocks/advises toward the cheapest fitting model; strict by default (v0.35.0+), advisory opt-out (ANTIHALL_MODEL_ROUTING=advisory via project-scoped env — global-export blast radius: blocks omitted-model mechanical spawns in every project; remedy: explicit cheap model or set advisory). omc-detect (0.32.0) = shared helper (not a hook), exported isOmcLoopActive(): detects whether an oh-my-claudecode autonomous loop is active and fresh; consumed by task-guard + tasklist-guard to defer Stop-blocks to advisory when an OMC loop is running. merge-gate (0.31.0) = OPT-IN PreToolUse Bash, default OFF; only when ANTIHALL_MERGE_GATE ∈ {1,true,yes,on} does it block (exit 2) an auto-merge intent (gh pr merge incl. --auto, gh pr review --approve, git merge --no-ff/--ff into main/master/develop) when the recent assistant transcript tail (bounded 128 KB) carries an UNRESOLVED self-hedge ("pending review"/"first-pass"/"do not merge"/"needs your eyes"/…) — i.e. mechanizes the checkable part of the v0.30.0 "false done" failure (hedged then merged anyway). A hedge is RESOLVED (merge allowed) when a resolution token follows it ("owner signed off"/"fidelity verified"/"verified against"/"resolved:"/…). HONEST limits: keyword-heuristic, bypassable (alt merge syntax/heredoc/UI/API), fail-open on every error, cannot hard-loop (PreToolUse single-shot, no state). A backstop, NOT a guarantee. api-guard (0.22.0) = PreToolUse Write/Edit/MultiEdit, blocks fabricated stdlib/builtin APIs in code (verified against installed python3/node); bench tools/eval/api-guard-bench.js. ship-it-guard (0.28.0, .planning/ support removed 2026-07-03 — GSD discontinued) = OPT-IN PreToolUse Write/Edit/MultiEdit, default OFF; only when ANTIHALL_SHIPIT_GATE ∈ {1,true,yes,on} does it block (exit 2) a CODE edit on a hard-risk path (migration/auth/.github-workflows/security) when no PLAN.md exists (repo root only). Also does CONFORMANCE ADVISORY (never blocks): when a PLAN.md's "## Phases" declares per-phase files: lists, a Write/Edit/MultiEdit target matching none of them gets an advisory. HONEST limits: enforces artifact-EXISTENCE only (not plan quality), bypassable via Bash heredoc, conservative (never gates ordinary edits), fail-open. task-guard (0.29.0) = Stop hook that now detects IDLE NEGLECT: it classifies open tasks into ACTIONABLE-NOW (status pending + unowned/main-owned + no OPEN blockedBy) and checks ~/.anti-hall/agents/*.json for a FRESH (<~20 min ts/mtime) heartbeat; if ≥1 actionable-now task AND no agents running, it blocks with a SHARP, task-naming reason demanding parallel dispatch — otherwise (agents in flight, or only blocked/owned/in_progress tasks) it falls back to the gentler generic nudge. A task carrying an explicit owner-blocked marker (metadata.blockedOn/blockedOn ∈ {owner,user,human,external}, or an “OWNER:”/“OWNER DECISION” subject prefix — isOwnerBlocked()) is treated as non-dispatchable and excluded from ACTIONABLE-NOW without needing a fake blockedBy; toggle guards.taskGuardOwnerBlockedMarker (default on). Loop-safe: idle-neglect dedupes on (actionable-set + "no-agents"), absolute MAX_BLOCKS=5 cap across both modes → cannot hard-loop; fail-open. task-tracker (0.29.0) = UserPromptSubmit hook providing the per-turn complement to that Stop block: on EVERY prompt it reconstructs tasks (now capturing owner/blockedBy; status-only TaskUpdate doesn't clear them), applies the SAME classifyOpen (pending + unowned/main-owned + no OPEN blockedBy) and the SAME ~/.anti-hall/agents/*.json fresh-heartbeat check; when ≥1 actionable-now task AND no agents running it injects a SPECIFIC "TASK REVIEW (every turn): N … dispatch a background agent for EACH now, in parallel … : <up to 4 names>" line into additionalContext (subjects control-char-stripped + JSON.stringify'd = inert), else the existing generic discipline + open-tasks freshness note. Bounded (256 KB tail) + fail-open. codex-nudge (0.36.0) = Stop, advisory: nudges once/session for an independent Codex second-opinion review when substantial code shipped with no Codex review; off-switch ANTIHALL_CODEX_NUDGE=off. command-guard (DevSwarm redirect 0.53.0) = besides the coordinator heavy-command block, under a DevSwarm-active session it also redirects DESTRUCTIVE native inbox reads in ALL contexts (coordinator AND subagent — a delegated read drains the queue identically) with its OWN skip name devswarm-read-guard, evaluated BEFORE command-guard's own skip/coordinator gate: hivecontrol workspace monitor blocks UNCONDITIONALLY (no-timeout long-poll that hangs the shell + consumes the queue); hivecontrol workspace read-messages blocks ONLY with durable-layer evidence (ANTIHALL_DEVSWARM_INBOX_CMD non-empty OR a ~/.anti-hall/devswarm/workspaces/*.json descriptor with truthy inboxPath) else ALLOW (harmless single-consumer read; fail-OPEN-to-allow). Matching reuses the heavy path's quote-neutralization + bash -c/eval/$()/backtick recursion so quoted DATA (grep '…read-messages…') does NOT false-positive while smuggled forms still block; closed-vocabulary reason never echoes input, does not claim a message-count 0 proves absence, and warns not to delegate the read. edit-guard (0.50.0) = PreToolUse Write/Edit/MultiEdit/NotebookEdit: blocks a COORDINATOR from editing files directly, requiring delegation to a subagent (subagents always pass — payload agent_id/agent_type or CLAUDE_CODE_ENTRYPOINT=agent_tool); DevSwarm-aware block wording (sub-orchestrator phrasing when a DevSwarm liveness supervisor is active). Built-in root-anchored allowlist (CLAUDE.md/AGENTS.md/GEMINI.md, .claude/**, .omc/**, .anti-hall/**, root-only PLAN.md/STATE.json) — bare-filename patterns match ONLY a root-level file, never a same-named file nested elsewhere; extensible via ANTIHALL_EDIT_GUARD_ALLOW (:/,-split globs, unrestricted depth). Skip-guard hatch name edit-guard (NOT in the DESTRUCTIVE set, so a broad all skip covers it). Shares coordinator/subagent detection with command-guard.js via coordinator-detect.js. Fail-open. coordinator-detect.js (0.50.0) = shared module (not a hook, not registered in hooks.json) — extracted from command-guard.js, exports isCoordinator()/isSubagent(): the single coordinator-vs-subagent discriminator now consumed by both command-guard.js and edit-guard.js, preventing detection-logic drift across guards. codex-availability (0.51.0) = SessionStart hook: OS-agnostic PATH probe (Windows PATHEXT-aware) for a real codex executable, writes ~/.anti-hall/codex-availability.json ({available, checkedAt, source:"path-probe"}) once per session so coordinators/skills read the cached fact instead of re-probing; when available, emits an additionalContext nudge toward codex:codex-rescue for the deadly-loop/ship-it Critic seat. Proves reachability only, NOT authentication/readiness — a runtime spawn can still fail even when available:true. Registered on both the Claude plugin (SessionStart) and the Codex port (SessionStart). Fail-open. devswarm-child-role (SessionStart, OPTIONAL/feature-gated) = Layer 1 of the DevSwarm layered recovery model: for a DevSwarm CHILD workspace only (devswarm-detect.js active AND devswarm-role.js reports child via non-empty DEVSWARM_SOURCE_BRANCH), injects a reminder to proactively self-report idleness (hivecontrol workspace message-parent) instead of sitting unnoticed off the parent's task list; silent no-op for Primary/non-DevSwarm sessions, byte-identical to dormant behavior. devswarm-parent-inbox / devswarm-parent-gate / devswarm-child-turn / devswarm-child-gate (0.53.0, OPTIONAL/feature-gated, mirror the same dormant-unless-DevSwarm model as devswarm-child-role) = the Phase-1 mechanical triggers for the "Primary neglects child workspaces" failure (claude-code#39755). devswarm-parent-inbox (UserPromptSubmit, Primary only) surfaces the real unread/idle state of active workspaces each turn + recommends archiving any archive_ready workspace; as of 0.54.1 it ALSO injects a compact live table EVERY turn (one row per active workspace: status/finish-rate/unread/last-activity, attention-needing rows sorted first, capped at 12 rows with a logged +N more); as of 0.60.0 a row idle beyond ANTIHALL_DEVSWARM_IDLE_MS (default 6h, ms) is relabeled active→idle in the table (view-only demotion — no delete/gate/archive; never overrides escalated/stale/archive-ready); as of v0.70.1 a row whose newest known activity signal is at least ANTIHALL_DEVSWARM_DORMANT_MS old (default 30 min, ms) is instead labeled dormant (sorts last, below even active) — because a mesh/registry row outlives its workspace (closing one in the DevSwarm app deletes nothing), and only heartbeat/verdict age reliably separated a live workspace from a closed one across measured cases; this demotes, never hides (a dormant row still renders with its unread count), never overrides escalated/stale/archive-ready, and roster (scripts/devswarm.js) carries the identical hint so the two surfaces can't disagree. devswarm-parent-gate (Stop, Primary only, capped) blocks the Primary from ending its turn while a child has unread backlog past its cursor OR the supervisor judged it stale/escalated OR (0.56.0) the Primary's OWN summary-projected unread is nonzero, surfaced with the same imperative "STOP and read them FIRST" wording as the child gate; devswarm-child-turn (UserPromptSubmit, child) writes a turn-authored heartbeat + reminds the child to report to its parent, and as of 0.54.1 also surfaces (PARTIAL — surfacing only, does not itself drain the native queue; full child-side reception is a v0.54.2 follow-up) a non-destructive unread count from the child's own durable inbox when >0; devswarm-child-gate (Stop, child, capped) forces a self-report before idling — as of 0.54.1, the brief v0.54.0 heartbeat-freshness silencing was REVERTED (it false-silenced a child that worked <5 min then stopped without reporting), so the gate ALWAYS demands a report per unchanged blocking state, bounded by the per-window cap MAX_BLOCKS = 2 plus a never-resetting MAX_BLOCKS_PER_SESSION = 6 lifetime cap (v0.97.0); a benignly-dropped heartbeat --summary still counts as an attempted report via a session/nonce-authenticated local attempt record (summary-attempts/<repoKey>.ndjson) rather than a re-prescribed loop. Hooks NEVER open the store — they read the summary.json projection + fs/cursor signals (no git/computeLiveness on the hot paths). Substrate they sit over (not hooks): companion/lib/devswarm-store.js = the persistent write/derive side (ONE API, TWO backends chosen by feature-detecting node:sqlite → WAL sqlite else append-only NDJSON journal; dependency-free, green Node 18/20 through 22/24; derives summary.json atomically + the configurable-gate archive_ready state); scripts/devswarm.js = THE structured CLI (register/ensure, heartbeat, inbox count/read/ack, workspaces list, gate, nudge, archive, archive-ignore/-unignore, migrate — with a root-anchored command-guard LIGHT_EXCEPTION so the guard doesn't block its own wrapper); companion/devswarm-migrate.js + companion/devswarm-ingest.js = the auto-safe migration (idempotent, non-destructive, single-consumer-locked, count-verified) + the one supervised monitor→store ingest daemon (lockfile-enforced single native consumer); as of 0.54.1, companion/install-devswarm-ingest.js auto-installs/refreshes that daemon on /anti-hall:update inside an active DevSwarm session (same no-offer/no-ask posture as install-devswarm-supervisor.js) — macOS LaunchAgent KeepAlive / Linux systemd --user Restart=always .service (cron fallback ticks every minute, ~60s worst-case revive gap on a cron-only Linux host), idempotent, distinct label/log from the supervisor. See docs/KB-devswarm-hivecontrol.md §8.7. | plugins/anti-hall/hooks/
6. Recommendations (consolidation outcome)¶
- No source docs deleted. All 13 remain in place; this KB references and
classifies them. [UPDATE 2026-09-24:
docs/now holds 45.mdfiles at top level (ls docs/*.md | wc -l), all indexed indocs/README.md. Earlier: 41.mdfiles total (ls docs/*.md | wc -l) — the "13" was the original consolidation-era count; the growth is newly-added living docs (§2 now lists 8 rows appended this pass) plus historical/planning artifacts, not a contradiction of "no source docs deleted."] - Archived. The historical artifacts
(
AUDIT-REPORT.md,AUDIT-REPORT-2.md,PLUGIN-REVIEW.md,ULTRAPLAN.md, the dated design plans, and thesuperpowers/specs and plans) live indocs/archive/and are linked from §5. - Folded docs stay standalone.
gsd-distilled,superpowers-planning,keynote-*are distilled intoKB-claude-codex.md§9/§12/§13 but retained for depth + provenance. Marked folded in §2. - Single canonical file chosen over a
KB/tree. The heavy content already lives in well-structured source docs (3,500 lines across 13 files). Re-flowing them into section files would duplicate content and create a second staleness surface to maintain. The higher-leverage artifact is this thin authoritative layer — index + current ground truth + staleness ledger — sitting over the existing docs. One file, one place to keep fresh. - [APPENDED 2026-08-21] Cross-cutting finding: version/hook/skill counts are
hand-maintained in TWO places, and both drifted independently this pass.
KB.md§1 anddocs/KB-claude-code-harness-features.md§2/§1 each carry their own hand-typed version number, hook file count, and skill count — this pass found all three stale in BOTH files, independently and by different amounts (KB.mdsaid 46 hooks/12 skills/v0.71.1; the harness-features doc said v0.68.2/14+17 skills). Flagging as a candidate for a single generated ground-truth block (e.g. a small script deriving version/hook-count/ skill-count fromplugin.json+ls hooks/*.js+ls skills/at doc-build or CI time) that both docs transclude or cite, instead of two independently hand-maintained numbers that can silently diverge. Not implemented this pass — recorded as a recommendation only.