Skip to content

Changelog

All notable changes to the anti-hall plugin are documented here. The plugin pins an explicit version in plugin.json only (the authority); the marketplace entry carries no version to avoid the silent-precedence trap where plugin.json wins silently. Every behavioral change MUST bump plugin.json version or installed users will not receive the update.

0.300.2 (2026-10-10)

Engine update tool, security fixes and repo tidy.

Added

  • ah-update: one command (hooks/ah-update.sh, no Node) moves the engine binary between sources: offline from a file (--from), the latest stable release (--channel stable) or the latest dev build (--channel dev), with sha256 checks and --rollback. The new engine.autoUpdate setting (off, stable, dev; default off) can run it for you at most once a day (#141).

Fixed

  • Security: the markdown escapers in the handover and pre-compact snapshot hooks now escape the backslash first, so a backslash can no longer undo an escape (#133); the remaining engine workflow actions are pinned to commit SHAs (#134).
  • Engine settings parity: the settings CLI knows engine.autoUpdate, and the parity goldens run in CI whenever the engine config or settings schema changes (#153).

Docs

  • OpenSSF Best Practices badge in the README (#147).
  • Engine docs match the shipped engine; new section on verifying release binaries (#135, #138).
  • GUIDE, AH-ENGINE, PRIVACY and llms.txt describe ah-update and engine.autoUpdate; the repo-pipelines and contributing docs say Copilot is off by default.

Repo (no effect on the installed plugin)

  • Repo root tidied: benchmark scripts moved from eval/ to tools/eval/, docs site requirements to docs/requirements.*, the engine lock example into ah-engine/; stray helper scripts removed (#149, #152).
  • Issue-triage workflow retired (community workflow handles it) and the community gate crash fixed (#146).
  • Model steps use structured output; GitHub Copilot is off by default in the model chain (copilot_fallback: false) (#150).

0.300.1 (2026-10-10)

Docs, site and repo patch. No engine or hook behavior changes.

  • Docs site: neutral near-black dark background with a faint teal hero glow (#125); new "How it works" section linked to the settings page (#127).
  • Docs: the Rust ah-engine is presented as the core component everywhere (README, GUIDE, HOOK-LATENCY, docs index, skills in both ports); the Node hooks are a temporary fallback removed in v1.0, never an option or mode (#127, #129).

0.300.0 (2026-10-09)

Highlights

  • Optional ah-engine (Rust) answers hook calls without starting Node per call. One thin trigger per event; the engine decides natively what it can prove identical to the Node hook and defers the rest to Node, never weaker than Node. Claude and Codex ports alike. macOS and Linux; Windows is not supported yet.
  • Everything tunable lives in plugin files (plugins/anti-hall/engine/), read at run time with hot reload, layered failover and self-heal. Nothing is compiled into the binary.
  • Local-only telemetry: ah-engine telemetry summary.
  • Released separately. The engine has its own version and GitHub Release (ah-engine-v0.1.0, six targets, build-provenance attested); the plugin pins it by sha256 in ah-engine.lock. A plugin without the binary, offline or on an unsupported platform runs the Node hooks exactly as before.

Added

  • Engine bootstrap with sha256 verification. The plugin ships ah-engine.lock (engine version and the sha256 of each release asset). On SessionStart hooks/ah-engine-bootstrap.sh (POSIX sh, detached, never fails a session) downloads the archive for the detected target (macOS arm64/x86_64, Linux x86_64/arm64 glibc or musl, WSL as Linux) from the GitHub Release over HTTPS and installs ~/.anti-hall/ah-engine/bin/ah-engine only if its sha256 equals the lock's entry (no trust on first use; a mismatch installs nothing). Atomic, idempotent, retried at most every 6 hours after a failure, previous binary kept as ah-engine.prev, a binary it did not install is never overwritten, log in ~/.anti-hall/ah-engine/bootstrap.log. Opt out: AH_ENGINE_BOOTSTRAP=0. A plugin tree without ah-engine.lock installs nothing and stays on Node.
  • One thin trigger per event. hooks/hooks.json and codex/hooks/hooks.json are generated from the dispatch table (engine/defaults/dispatch.toml) and call hooks/ah-hook.sh <Event> (Codex: --host codex; ten Codex events, all Claude Code events except WorktreeCreate/WorktreeRemove). hooks.registry.json lists the individual hooks. The wrapper runs the Node hooks when the engine is absent, times out, crashes or defers (exit 75).
  • Layered failover and self-heal for the engine files. Per file and per setting: the edited engine/defaults/, then the last-known-good copy (~/.anti-hall/ah-engine/defaults.lkg/), then the read-only engine/defaults.pristine/, then Node. Every fallback is logged and marks the engine degraded. A setting the edited files lack is taken from the pristine copy and appended to the edited file (existing text kept, backed up once, idempotent; skipped in a version-controlled checkout). ah-engine config heal does it on request.
  • Shadow and rollback. Per-entry mode = "shadow" runs an engine check beside its Node hook (Node decides) and logs dispatch_shadow; mode = "off" skips the engine's check. Whole rollback: ah-engine stop, delete the binary, AH_ENGINE_BOOTSTRAP=0.
  • Telemetry (ah-engine telemetry summary|events|rollup): per hook and check, the event, outcome, latency histogram and bytes injected into context, identifiers only, stored in ~/.anti-hall/ah-engine/, never uploaded. telemetry.enabled, telemetry.retention_days.
  • Docs: new "Install, go-live and rollback", "What still runs on Node" and "Measured results (pre-release)" sections in docs/AH-ENGINE.md; engine notes in the READMEs (both ports), GUIDE, CONTRACT-1.0, DEVELOPMENT, llms.txt and the system-briefing and settings skills (both ports); PRIVACY.md now lists the binary download, the local telemetry and the background process.
  • DevSwarm engine paths. The engine runs the DevSwarm gates and guards natively (parent and child gates, comms guard, reply tracker, child drain, codex nudge, task tracker, verify-first), with a new engine-devswarm-supervisor skill and its own settings area.
  • Reconcile and sweep behind a switch. The supervisor reconcile and sweep tail run natively only when devswarm_sup.reconcile_mode (engine file) and devswarm.sweepTailMode (setting) are set; both default to node, so nothing changes until you opt in.
  • Faster startup. The engine no longer links the macOS IOKit, CoreFoundation and CoreServices frameworks , and transcript-scanning Stop/prompt checks have their own script time limit so they no longer defer on real transcripts.
  • Skills call the engine with a Node fallback. Skills run through scripts/ah-run.sh, which uses the engine when it is installed and healthy and falls back to the Node script otherwise.
  • Command guard parity. The engine command check ports the read-only chain units and background-script safe flags from the Node guard (a gcloud read inside a read-only chain and python3 -I scratch scripts are not treated as heavy).

Changed

  • PRIVACY and README wording: "no telemetry" became "no analytics, nothing reported to anyone", because the engine keeps local-only counters; two requests are on by default (update check, one-time engine download).
  • Documented model default: jev.judgeModel is the alias haiku (the docs still said a pinned version).

Repo automation

These changes affect the GitHub repository only, not the installed plugin.

  • Copilot CLI install fix: the Copilot slot now always checks out its lockfile folder, so npm ci works when tools are none.
  • Claude and Copilot error diagnostics: each failing slot writes its own error line to the job summary.
  • Force-model dispatch: community.yml takes a force-model input that bypasses the daily cap for a test run.
  • AI_DAILY_CAP now counts actual model calls per day (a marker artifact is recorded when a slot answers), not workflow runs.
  • PR alerts in pr-check, and triage.yml folded into community and pr-check.
  • Docs-drift check plus a release docs review.
  • Discussions participation: announcements, Q&A follow-up, idea to issue conversion, and the /triage and /explain commands.
  • The weekly digest now posts privately to the project board.
  • Full board reconcile with Last update and Progress fields.
  • Dependabot auto-merge for patch and minor updates into dev.

Still on Node

DevSwarm mesh writes and daemons (ingest, supervisor, reaper), every call that consults a Jev integration, the semantic judge's model call, the statusline, and the blocking branch of several guards (the engine answers the quiet cases and defers any case that could block).

Measurements before release

From one replay of 2113 recorded payloads against the exact go-live bundle and the same-version Node hooks (not a field result): 0 of 68 blocks weaker than Node, 2059 identical outputs, 87.0 percent of hook rows answered natively, about 35.5 ms CPU per call for the engine against 158.2 ms for the Node hooks. Known gaps are listed in docs/AH-ENGINE.md.

0.203.3 (2026-10-10)

Repo

These changes affect the GitHub repository only, not the installed plugin (#87).

  • Model steps use two Claude credentials with immediate failover from the first to the second, then Copilot, then rules-only. Each job's summary names the slot used.
  • Per-job model routing: each automation job picks its own model from the config.
  • privacy-scan checks commit identity (author and committer) on pull requests.
  • The moderation sanitizer fix: model output is cleaned correctly before it is posted.
  • Leaner pull-request CI: Node 24 only on pull requests; the full matrix runs on tags, main and weekly. A nightly job runs the parity and bench tests.
  • Realtime board sync: the roadmap board and stale check update on pushes to dev/main, pull-request close and issue events (rules only, no model); the digest stays weekly.
  • Dev skills: every agent and lane brief names its issue and keeps it updated.

0.203.2 (2026-10-09)

Security

  • command-guard: the timeout prefix regex (TIMEOUT_PREFIX_RE) had two overlapping flag alternatives, so a timeout followed by many -k -x pairs made the PreToolUse classifier backtrack exponentially (CodeQL js/redos). The alternatives are now disjoint; every real timeout form strips to the same body. A regression test runs the CodeQL witness string (#56).

Repo

These changes affect the GitHub repository only, not the installed plugin.

  • Issue forms with priority, area and estimate fields; discussion category forms for ideas, Q&A, show-and-tell and general; saved replies and a release notes template.
  • Rules-first automation for triage, moderation and research briefs on issues and discussions, PR checks, and a weekly roadmap digest, each with an optional Claude or Copilot model step that falls back to rules only.
  • Privacy scan on pull requests (gitleaks plus repository rules for private paths and identifiers).
  • Roadmap board automation, PR path labels, PR title check, stale handling, dependency review, job timeouts, PR concurrency and SHA-pinned actions.
  • CodeQL via default setup (the advanced workflow is manual-only), OpenSSF Scorecard, and release drafter.
  • Docs site built with MkDocs and published on GitHub Pages at https://talas9.github.io/anti-hall/, plus the wiki; docs/REPO-PIPELINES.md lists every workflow and policy, and SUPPORT.md says where to ask.
  • Dev-only skills under .claude/skills/ (dogfood, gh-work, release, engine-lane, repo-hygiene); they are not part of the shipped plugin.
  • Dependabot bumps for GitHub Actions (checkout, setup-node, configure-pages, deploy-pages, claude-code-action), now targeting dev.

0.203.1 (2026-10-09)

Fixed

  • api-guard: the Python probe now runs with -I (isolated mode) instead of -s. With -s the probe's current directory (the temp dir) led sys.path, so Python listed the whole temp dir on its first import (40-73 s on a machine with about 490k temp entries) and a json.py placed in the temp dir would run inside the probe.

Security

  • On Linux the temp dir is the shared /tmp, so another local user could plant a json.py there and have it execute inside the api-guard probe. -I removes the current directory and user site from sys.path and ignores PYTHON* environment variables, closing that path.

0.203.0 (2026-10-09)

Highlights

  • Leaner injected context. The SubagentStart worker core, the task-tracker reminder and the DevSwarm workspace table all re-send less (measured per-turn and per-spawn reductions below).
  • Accurate limit advisory. The limit-conservation text no longer claims a Claude model has its own weekly bucket.
  • Language-agnostic. The Flutter-specific flutter-debug skill and agent are removed.
  • Alias-based model routing. The Jev judge and all plugin code route by model alias, never a pinned model version.
  • Guard and DevSwarm fixes. git-guard heredoc handling and self-credit messages, command-guard read-only chains, one canonical DevSwarm workspace-id resolver, spawn-time submodule fetch, and auto-archive records that no longer stay pending.

Changed

  • git-guard: a data heredoc is no longer vetoed by read-only neighbouring commands or a literal $VAR target, and the self-credit block names the command that carries the credit.
  • command-guard: a gcloud read inside a read-only chain and python3 -I scratch scripts are no longer treated as heavy.
  • edit-guard: a handover written outside the project gets a specific redirect to the project path.
  • Jev judge and plugin code route models by alias (haiku/sonnet/opus), with a guard test against pinned slugs.
  • DevSwarm: one canonical workspace-id resolver (meshId/uuid/prefix) across unarchive, send, gate and the other id-taking verbs; spawn fetches a missing pinned submodule commit before create and surfaces every submodule failure in warnings; an auto-archive record without doneHead no longer stays pending forever (already-written entries are repaired).
  • limit-conserve: reset time renders at minute precision so emit-dedupe matches across jitter; turn-gate pruning uses the correct tg family; jev-report keeps triage-derived rows out of Jev totals; eval trace extraction counts assistant usage once per message id.

  • Fixed the limit-conservation advisory (limit-conserve-inject.js, Codex anti-hall-context-conserve skill) claiming Sonnet draws on a "SEPARATE weekly bucket" or that a downshift preserves a "flagship weekly bucket". Per-model weekly buckets are not documented, and the usage screen shows only "All models" and "Fable only". The text now says only that Codex has its own limit and that cheaper models or fewer agents use less of the shared Claude pool, and makes no claim about Fable's relation to "All models". A test fails if the advisory ever again calls Sonnet, Opus or Haiku a separate bucket.

  • Removed the Flutter-specific flutter-debug skill and agent; anti-hall is language-agnostic. The removed code remains in git history.
  • The DevSwarm workspace table (parent inbox) is no longer re-sent when only a row's unread count changes; unread already arrives per turn in the inbox segments. Status/finish/risk changes still re-send it, and the keepalive now follows guards.injectionRepeatEvery instead of a fixed 10. Measured on a 3-workspace fixture over 16 delivered turns: 286 -> 198 injected chars per turn on average (the COMMS OVERRIDE line was already once-per-session plus keepalive).
  • SubagentStart (verify-first-subagent.js, default context.protocolLevel=compact) now emits a worker-only short core: iron law, stop-and-verify triggers, done/scope/autonomy/skip rules, the PROTOCOL.md pointer and WORKER. Measured per spawn: 3,564 -> 2,518 chars compact (7,701 under protocolLevel=full, unchanged). Main sessions keep the existing session core. The subagent size caps in verify-first-subagent-compact.test.js dropped to 2,800/3,100 to pin it.
  • The task-tracker SHORT reminder line is no longer injected on every prompt: it follows guards.injectionRepeatEvery (first turn after the FULL primer or a compaction, then every N delivered turns; 0 = every turn). The per-turn open-tasks / DISPATCH NOW note is unchanged, and the hook no longer emits an empty context block when nothing applies. Measured over 12 delivered turns with no open tasks: 1,896 -> 546 chars (158 -> 45 per prompt).

0.202.0 (2026-10-04)

Highlights

  • Bash guards return decisions instead of exiting (no process.exit; CLI output byte-identical).
  • Guards honour the evaluate() env argument, and guard output writes retry EAGAIN.
  • docs/HOOK-LATENCY.md hook counts corrected and scripts/hook-latency.js --grouped added.
  • Codex skill descriptions trimmed.

Changed

  • A guard's evaluate(payload, env) decision depends only on its env argument. The settings readers, the skip marker (skip-guard.js isSkipped(name, env)), the home directory, the coordinator entrypoint check and the Jev consult now take the env given to evaluate() instead of reading process.env; the CLI passes process.env, so command-line behaviour is unchanged. Guard output writes (guard-io writeAll) now retry EAGAIN for at most 3 seconds of real elapsed time (measured with a monotonic clock, so an oversleeping host cannot stretch it) instead of silently truncating.
  • The 15 Bash pre/post guard invocations return decisions instead of exiting. compact-declaration-guard, git-guard (pre and --audit), command-guard, coordinator-work-guard (pre and --post), merge-side-pick (pre and --post), merge-gate, scan-throttle, api-guard, ship-it-guard, output-verify-guard, devswarm-parent-reply-tracker and devswarm-child-drain each export evaluate(payload, env, { argv }) returning { exitCode, stdout, stderr } and never call process.exit; a thin require.main === module wrapper (hooks/lib/guard-io.js) reads stdin, writes both streams and exits with the code. A block can no longer turn into an allow when guards run in one process inside fail-open try/catch blocks. CLI behaviour is byte-identical. tests/hooks/guard-evaluate-parity.test.js runs every guard through the spawned CLI and in-process and compares exit code, stdout and stderr.
  • docs/HOOK-LATENCY.md per-event hook counts are corrected (Bash 9 PreToolUse and 6 PostToolUse on 0.201.0, plus the Codex port's counts), and scripts/hook-latency.js --grouped starts each event's whole hook set in parallel and reports group wall time and summed CPU.
  • Six Codex skill descriptions are shorter (anti-hall-doctor, -install-statusline, -omc, -omx, -orchestration, -update).

0.201.0 (2026-10-04)

Highlights

  • Tradeoff fixes A-F. git-guard skips heredoc bodies that are data and checks xargs, resolves git aliases and reused commit messages, and holds git/gh beside a data heredoc to a flag allowlist; api-guard, ship-it-guard and Bash edit parity see shell writes (no 64 KB skip); an opt-in unsupported-inference check and a keyless CLI judge backend; noisy advisories (root-cause nudge, output-verify) cut at the source; DevSwarm hooks exit early where they cannot act.
  • README tradeoffs list moved into GUIDE "Limits and escape hatches".
  • Replay false-positive fixes from 5,828 real PostToolUse events.
  • Benchmark harness (evals/anti-hall/) runs in the container sandbox, with --reps, a per-run cost reserve and --cases-dir.
  • New opt-in/advisory features: merge-side-pick advisory, Codex spawn-time orchestration rules, reduced tasklist nag, per-prompt Stop-nag budget.

Added

  • Idle-agent sweep (hooks/idle-agent-sweep.js, UserPromptSubmit, advisory; guards.idleAgentSweep default on, guards.idleAgentSweepCount 3, guards.idleAgentSweepMin 15; Claude and Codex). Field case: 34 named teammates had sent their final report and sat idle for up to ~1.5 h because nothing reminded the coordinator to stop them. Once per user prompt, the hook lists agents that finished but were never stopped, oldest first (up to 10 names, then "and N more"), with the exact call to end them. Claude: a named teammate whose newest event is an idle_notification with idleReason available or failed, with no later SendMessage to it and no TaskStop (by name or <name>@<team>); an idle with no idleReason (waiting on its own work) and background agents (already ended by their completed notification) are not listed. Codex: a multi_agent_v1 agent whose wait_agent result is completed/errored and that was never closed (close_agent) or re-tasked (send_input/resume_agent); the newer collaboration tool set has no close tool, so nothing is listed there. Fires when 3 are idle or one has been idle 15 minutes; a <task-notification> turn is skipped and a queued burst gets one copy. Replayed on the field session it fires at the first user prompt after the teammates finished (20 names, oldest idle 44 min) and at every prompt until the stop (34 names, the same 34 the coordinator stopped). Detection lives in hooks/lib/agent-scan.js (finishedTeammates) and hooks/lib/idle-agents.js.
  • api-guard and ship-it-guard see shell writes (guards.shellWriteChecks, default on; Claude and Codex). Both are now registered on Bash as well as Write/Edit/MultiEdit (Claude) and apply_patch (Codex). hooks/lib/shell-writes.js lists the files a command writes (>/>>/>|/&> incl. > f truncation, heredoc into cat/tee, echo/printf, tee [-a], sed -i, perl -i/-pi, cp/mv, literal python -c/node -e open-for-write paths; also inside sh -c, eval, $(…)) using command-guard's own parsers. ship-it-guard's existence gate runs on those targets (scratchpad and tmp-outside-a-repo writes excluded); api-guard checks the text a heredoc/echo/printf puts in a .py/.js/.ts file. A target the parser cannot know (a variable, a glob, dd/install/rsync, a script run) is allowed. Cost: two more hook processes per Bash call (api-guard exits before parsing unless the command names a .py/.js/.ts file; ship-it-guard is off by default).
  • Shell-write checks no longer skip commands over 64 KB, and cover more inline writes. The classify cap (a guard against catastrophic regex cost) made a ~96 KB cat > src/auth/big.py <<'EOF' invisible to api-guard, ship-it-guard and Bash edit parity. Those now scan the heredoc header (target and redirect) with an empty body, take the real body as the content, and parse text after the heredoc unless the header line has a cd; a big command with no leading heredoc is judged on its first 16 KB. 1 MB command: 34-60 ms per hook, wall time with node startup. api-guard still skips a body over its own 600 KB chunk cap, same as the Write tool. Inline-code literals now include perl 2- and 3-arg open with a write mode and node createWriteStream (ruby File.open(...,'w')/File.write and node writeFileSync/appendFileSync were already covered).
  • Bash edit parity covers inline code (guards.bashEditParity). A literal python -c "open('src/x.py','w')…" (or node -e writeFileSync, ruby -e File.write) into a repo file now gets edit-guard's delegation block on the main thread, like sed -i and redirects already did.
  • git-guard sees through git aliases (guards.gitAliasResolve, safety, default on; Claude and Codex). git <name> for a non-builtin name is resolved through the repo/global git config at hook time (one git config --get-regexp ^alias\. per repo per call; builtins never spawn git), following alias chains (loops resolve to nothing) and !shell aliases, and the expansion is scanned as the command it really runs. Before, git pf with alias.pf = push --force, or git ci -m "...Co-Authored-By: <AI>..." with alias.ci = commit, was allowed. Defining an alias whose body is a blocked git command is blocked too: git config alias.x, git -c alias.x=, GIT_CONFIG_VALUE_<n>=, and a shell alias x= (an AI-credited commit body used to pass; a force-push body was already blocked). A call to a shell alias or function defined in the same command is scanned as the git command it forwards to (g(){ git "$@"; }; g push --force used to pass). The PostToolUse audit follows aliases as well. Logic lives in hooks/lib/git-alias-scan.js.
  • git-guard checks reused commit messages before the commit runs (guards.gitReusedMessageCheck, safety, default on; Claude and Codex). For a git commit with no -m/-F, the message it would take from -C/-c/--reuse-message/--reedit-message <rev>, from HEAD (--amend) or from -t/commit.template is read and blocked when it carries an AI self-credit trailer: always when reused verbatim (-C, --no-edit, a no-op editor such as GIT_EDITOR=true), and on the editor path when the command sets no real editor. A real editor override and anything a commit-msg hook writes are left to the existing PostToolUse audit, which tells the agent to amend.
  • Unsupported confident inferences (opt-in). guards.inferenceCheck (default off) makes speculation-guard also block, once per reply, a causal claim with no hedge word ("The crash is caused by the cache race.") when no tool output, observation-tool input, pasted fenced block or task notification in the last 1 MB of the transcript mentions the stated cause (new hooks/lib/inference-check.js; Claude transcripts and Codex rollouts). Off by default because field precision is low: 1.00 precision / 0.95 recall on the new 84-case synthetic corpus (eval/inference-bench.js), but it flagged 3.8% of 3,276 real final replies and at most 18 of a 40-flag sample were causal claims about project state (precision ≤ 0.45).
  • Semantic judge without an API key. jev.judgeBackend (api default, cli, auto): cli runs the judge on the local claude -p CLI with your own Claude login, isolated (no tools, no MCP servers, no settings files, disableAllHooks) and marked ANTIHALL_JUDGE_CHILD=1 so it cannot recurse; measured about 5–6 s per turn end. Fail-open when the CLI is missing or not logged in.
  • eval/inference-bench.js + eval/inference-cases.json: offline precision/recall for both components (--hook, --codex, live --judge-cli).
  • merge-side-pick advisory (hooks/merge-side-pick.js, PostToolUse + PreToolUse Bash, Claude and Codex): after a conflict is resolved by taking one side wholesale (git checkout|restore --ours|--theirs, git merge|pull|rebase -X ours|theirs, git merge -s ours), a push with no test run since (npm/pnpm/yarn test, node --test, pytest, go/cargo/flutter/dart test, mvn/gradle test, make test, ...) gets one advisory line in the shared message shape. Never blocks; per-session state in ~/.anti-hall/merge-side-pick-<session>.json, pruned after 7 days. Setting guards.mergeSidePickAdvisory (default on, env ANTIHALL_MERGE_SIDE_PICK_ADVISORY). On Codex the advisory is shown only by Codex builds that support PreToolUse additionalContext (rust-v0.129.0 and later, docs/KB-claude-codex.md section 5.2); older builds ignore it and the recorder stays harmless. Designed from the owner-decisions description; the original report text was not found.
  • Experimental, opt-in: Codex spawn-time delivery of the orchestration rules (context.codexOrchFullOn=spawn, env ANTIHALL_CODEX_ORCH_FULL_ON; default session = today, byte for byte). Needs Codex >= 0.129 with hooks trusted. On a positively identified Codex session SessionStart sends the compact core plus the compact orchestration lines, and the full Codex-worded rules arrive once per context epoch on the first spawn_agent call through orch-on-spawn.js (same marker + O_EXCL claim as Claude; the retry scan also recognises Codex rollout developer messages). orch-on-spawn.js is now registered in the Codex port on the ^(?:collaboration)?spawn_agent$ matcher only (silent unless the marker is pending). Codex 0.160 surfaces PreToolUse additionalContext to the model (live probe 2026-10-04: tool_name collaborationspawn_agent, hook fires, context reaches the parent). Codex SessionStart chars in spawn mode are measured with evals/anti-hall/injection-profile.js.

Performance

  • Hook latency: DevSwarm hooks exit before loading their libraries in a session they cannot act on. devswarm-parent-gate, devswarm-parent-inbox, devswarm-child-turn, devswarm-child-gate and devswarm-child-drain each load a dozen companion libs at module scope (70 to 187 KB of source) and only then reach the check that returns early for a non-DevSwarm session or the other role. New hooks/lib/devswarm-primary-gate.js repeats those same payload-independent checks (the hook's own settings switch, the user skip marker, DevSwarm active, Primary vs child) before the heavy requires, drains stdin and exits 0. A session the hook can act on goes through the full hook unchanged. Measured, paired and interleaved against the previous build, DevSwarm env stripped: devswarm-parent-gate (Stop) 36.8 to 21.3 ms CPU p50, devswarm-parent-inbox (UserPromptSubmit) 38.2 to 21.5 ms, a bare node -e 0 being about 16.5 ms. One side effect, not a decision: devswarm-parent-gate and devswarm-child-gate (Stop) and devswarm-child-drain (PostToolUse) no longer rewrite the stable launcher under ~/.anti-hall/bin/ in a session they cannot act on. The launcher is still refreshed at SessionStart by devswarm-child-role.js and on every real DevSwarm Stop.
  • crypto and child_process load on first use in the hooks that rarely need them (hooks/lib/lazy-node.js; task-guard, tasklist-guard, speculation-guard/judge, claim-ledger, codex-nudge, compact-advice-guard, task-tracker, verify-first, phase-tracker, devswarm-child-turn and the shared libs jev-assist, emit-dedupe, command-allow, stop-ack, identity, ...). Call sites are unchanged. Measured gain is 0 to 4 ms CPU per hook, within noise on several, so treat it as small.
  • Measurement and the numbers that were checked and not changed are in docs/HOOK-LATENCY.md (addendum 2026-10-04, second).

Fixed

  • api-guard no longer treats the result of a chained require (const src = require('fs').readFileSync(f)) as the module itself, and no longer flags a property assignment on a module (fs.x = 1) as a fabricated API read.
  • git-guard shell-function wrapper expansion no longer splices call arguments into a name="$1" / local name="$1" assignment when the variable is only echoed or passed as data (test harnesses piping a command to a hook). It still blocks when the variable, or one derived from it, is later run ($c, eval "$c", bash -c, piped or here-string to a shell, source <(...)).
  • Judge hardening. The secret scrub now also covers sk_live_/sk_test_/rk_live_, glpat-, npm_, Authorization: Basic|Bearer values and quoted multi-word SECRET_KEY = "a b c" values, and the judge input is scrubbed before it is cut to length (a cut could leave a token fragment below every pattern). The claude -p judge child runs in a private empty temp dir (removed afterwards) instead of the shared temp dir. Every Stop, SessionStart and UserPromptSubmit hook exits 0 silently when ANTIHALL_JUDGE_CHILD=1 (new hooks/lib/judge-child-exit.js), as a second stop against recursion.
  • Benchmark harness (evals/anti-hall/, tooling only). Bash failed in every container run: the runtime's masked /proc subpaths make the kernel refuse bwrap's /proc mount (bwrap: Can't mount proc on /newroot/proc), so models could not run git/tests. run-in-container.sh now unmounts the masks as root with CAP_SYS_ADMIN and drops to node with no capabilities (EVAL_KEEP_PROC_MASKS=1 opts out). Also: --reps now reaches a single-arm run as --runs; new --run-reserve-usd per-run cost pre-check; --cases-dir reads cases from outside the repo with a generated manifest; -- --keep-temp copies each run's temp dir into <results>/kept/ with a relative tracePath, so traces outlive a --rm container.
  • git-guard: heredoc bodies that are data are no longer scanned as commands (setting guards.gitGuardHeredocData, default on, env ANTIHALL_GIT_GUARD_HEREDOC_DATA; Claude and Codex run the same hook). A note, commit message or PR body that mentions git push --force is no longer blocked when its heredoc feeds a prose/data file or a message: cat <<EOF > notes.md, tee notes.txt <<EOF, git commit -F - <<EOF, git commit -m "$(cat <<EOF ...)", gh pr create --body-file - <<EOF. All-or-nothing and fail-closed: every other command in the line must be on a short allowlist (no shell, interpreter, xargs, eval, source, or path/variable used as a command anywhere), targets must be literal prose/data files outside .git/.husky/hooks/.ssh/.config/.claude/.codex/.anti-hall/bin (no script, extensionless file, dotfile, git hook name or existing symlink), and the opener parse must be exact. A body fed to bash/sh/eval/source/xargs/python/..., piped into a shell, teed into >( ), or written to a file the same line runs is still scanned, and so is a script written by a heredoc. Credit trailers and git-config lines in a body are checked whatever the consumer. The pinned bypass suite (git-guard-heredoc-bypass.test.js) stays blocked; five tests that pinned prose into .md/.txt (and a handover file) as BLOCK now pin ALLOW.
  • git-guard: xargs runs get the checks a direct run gets. xargs options are parsed like getopt (GNU and BSD): GNU -i/-l/-e/--replace/--eof/--max-lines no longer swallow the command word, a required option value is taken even when quoted or dash-led (-d '\n'), clusters (-0n1), abbreviated long options and BSD -J/-R/-S are handled, and a bare {} is a literal word (xargs -I{} git ... is no longer split at the braces). An xargs-run sh -c, eval or nested xargs is re-scanned as a command. Before this, xargs -i git push --force ..., xargs -l, xargs -d '\n', xargs -J % and xargs sh -c '<force push>' were allowed.
  • task-guard IDLE NEGLECT no longer blocks while unmapped agents could be on the listed tasks. A running agent whose description names no #<id> was assumed to be on an in_progress task first, so the pending task it was actually on (or tasks waiting on its result) were reported as having "no in-flight agent". The Stop block now fires only when more dispatchable tasks remain than running unmapped agents, so at least one is uncovered however those agents are placed. The per-turn DISPATCH NOW line is unchanged. Replay of the 164 IDLE NEGLECT block entries in 14 days of transcripts: 58 had a live agent of the same session; of those, 34 no longer block and 24 still do. The 80 with no live agent are unaffected. Setting: guards.idleNeglectProvenOnly (default on).
  • task-guard IDLE NEGLECT: a hung or unrelated agent no longer hides a pending task. In the proven count, a running agent that names no #<id> stops counting as cover when its newest sign of life (launch, SendMessage resume, pending teammate message, output-file write) is older than guards.idleNeglectAgentMaxAgeMin (default 30; 0 = never), or when it was launched before the earliest uncovered task was created or last set pending/in_progress. Unknown age or task time = the agent still counts. The per-turn DISPATCH NOW line is unchanged.
  • Noisy advisories: root-cause nudge and output-verify cut at the source. Replayed 5,828 real PostToolUseFailure nudges from 30 days of transcripts through the hook: before, all 5,828 fired; after, 3,438 fire. The 2,390 removed were 362 expected exit-1 predicates (grep no match, test, diff, git diff --quiet, command -v, pgrep; only when the final statement or every && link is a predicate and the exit code is exactly 1), 87 harness refusals (the command never ran), and 1,941 repeats inside one human turn (the text never changes between calls). New lib/expected-failure.js and lib/turn-gate.js; setting guards.failureNudgeFilter (default on, env ANTIHALL_FAILURE_NUDGE_FILTER, off = nudge on every failure again). output-verify-guard gets the same once-per-turn gate per distinct signal set (guards.outputVerifyOncePerTurn, default on): 3,183 historical advisories, 69 still fire on the staged hook (the zero-count and runner-detection fixes already removed 3,114), 57 with the gate. No blocking guard changed. Claude-only hooks (the Codex port has neither PostToolUseFailure nor this PostToolUse shape), so no Codex mirror.
  • git-guard: git and gh beside a data heredoc are held to a per-subcommand flag allowlist. A masked body written to a.md could be run on the same line by git fetch --upload-pack='sh a.md' . (also --upload-p, --up, -u), gh repo clone o/r -- --upload-pack=... or git commit -F - --upload-pack=.... Beside a data heredoc, git may now use only commit/tag/notes/merge (message and common flags), status/log/diff/show/add (read-only flags) and rev-parse, plus --no-pager and -C <literal dir>; fetch, push, pull, clone, remote, submodule and config are no longer allowed there, and -c/--git-dir are not either. gh may use only pr/issue/release create/edit/comment with listed flags and no -- passthrough. Any other flag (abbreviations included) keeps every body scanned. A git push on the same command as a data heredoc now gets the body scanned again.
  • git-guard: wrappers, positional-arg forwarding and find -exec. stdbuf, caffeinate, ionice, flock (flock FILE -c 'cmd' is scanned as sh -c), setsid, chrt, taskset and doas are unwrapped with their own option grammars, directly and under xargs. A sh -c/bash -c script that forwards its positional args (sh -c '$0 "$@"' git ..., bash -c '"$@"' _ git ...) is scanned with the args spliced in (unquoted references word-split), and under xargs with a --force standing in for the stdin words. find -exec/-execdir/-ok/-okdir commands (\; or +) get the direct-run checks, including -exec sh -c '...'; a find -exec git push ... {} is blocked because a file name can be a +ref refspec. All of these forms were allowed before.
  • git-guard: replacement strings, flock -c anywhere, GNU parallel, and git's real value grammar beside a data heredoc. When an xargs -I/-i/-J/--replace string, find's {} or a parallel placeholder is the command word, the git subcommand or an argument before it (xargs -I{} git {} --force, xargs -I{} sh -c 'git {} --force', find . -exec git {} --force \;), the command is unknown and is blocked when a force or remote-delete flag is visible; xargs -I{} git push origin {} is now blocked like any xargs-run push. flock -c/--command[=] before FILE (or --command= after it) is scanned as sh -c. GNU parallel gets the xargs checks (the command before :::/:::: with the inputs appended, the joined line scanned as a shell line, parallel ::: 'cmd' inputs scanned as commands). Beside a data heredoc, --format/--pretty/--unified/--abbrev/-U take only an attached value (as in git), so git log --abbrev --output=x or git diff --unified --ext-diff no longer hides the unlisted flag; -U<n> is accepted on log, and tag's --sort/--points-at/--format take the next word. All of these forms were allowed before.
  • git-guard: custom placeholders, stdin-fed subcommands and short-option clusters. { splits a command only as a standalone word (blanks on both sides) and } only after ;, & or a newline, so parallel git {1} --force ::: push, xargs -I{x} git {x} --force and parallel's {.}/{/}/{#} are no longer cut apart before the runner checks; { ...; } groups, function f { ...; }, coproc [NAME] { ...; } and time -p { ...; } bodies still split. When xargs or parallel runs git with no subcommand word (redirections ignored), the subcommand comes from stdin: it is blocked if a force or remote-delete token (--force, -f, +ref, --delete, :ref) appears anywhere on the line, quoted text included and after quote removal (echo 'push --force o m' | xargs git, echo push '-'f o m | xargs git). When the script comes from stdin (xargs -I{} sh -c '{}', xargs sh -c '$0 "$@"', | parallel with no command, | sh, | bash -s), the line's quoted strings and echo/printf words are scanned as commands. Beside a data heredoc, log/diff/show accept only modelled short options (a lone listed letter, -n N, -nN, -N, -M/-C/-U with a number); clusters such as -pn or -wn 3, which git rejects only after --output=FILE has truncated FILE, keep the bodies scanned. A $HOME/ write target is treated like ~/. git branch --merged | xargs git branch -d, ls | xargs git add, xargs -I{} git add {}, find . -exec git add {} + and parallel gzip ::: *.log stay allowed.
  • Jev triage: the per-hash claim and the arrival drain lock now use the single lock primitive (companion/lib/lock.js) instead of hand-rolled O_EXCL markers; the lock gained adopt(path, token) so the detached drain worker takes over the lock its spawning hook acquired. Behaviour unchanged; the hygiene allowlist entry is gone.

Changed

  • Manifest: plugin.json now sets termsOfServiceUrl (the MIT LICENSE), so the plugin directory listing shows a terms link next to support and privacy. The Codex manifest already had it.

  • speculation-judge now sends the judge the latest user request (up to 2000 characters) and the newest tool evidence from the transcript (up to about 6000 characters, secret-scrubbed) besides the reply, so "verified with a tool" is visible to it; the prompt says what counts as supported. Its block reason now uses the shared message shape. Measured precision 0.78–0.81, recall 1.0 (Haiku ×2, Sonnet ×1), so it stays opt-in. PRIVACY.md and the plugin README privacy table list the added text.

  • tasklist-guard reduced nag (guards.tasklistNoTaskTools, default reduced). A session that is positively known to lack task tools (today: a Codex session whose transcript shows no task-tool evidence) gets a short nag (no TaskCreate demand: list the tasks in the reply, with the progress and history paths) that blocks at most once per session. A Claude session with no evidence keeps today's full TaskCreate demand: under claude -p the task tools are listed but deferred, so the tool list proves nothing. Evidence is read structurally from the transcript (a TaskCreate/TaskUpdate/TodoWrite tool_use, a deferred_tools* attachment naming TaskCreate, or a task_reminder attachment), never by substring, so the guard's own text cannot confirm itself. Values full (today's nag) and skip. Unset, context.protocolLevel=full makes the default full.
  • Per-prompt Stop-nag budget (guards.stopNagBudgetPerPrompt, default 0 = off = today's behaviour). task-guard and tasklist-guard consult hooks/lib/stop-policy.js just before they block; with N > 0 each blocks at most N times per user prompt (key: Stop payload prompt_id, else the last user entry uuid; no key = not applied), inside their existing session caps. Other Stop hooks and the PreToolUse guards are unchanged.
  • Jev recommend notice is not sent in non-interactive runs (jev.recommendNoticeHeadless, default false). claude -p exports CLAUDE_CODE_ENTRYPOINT=sdk-cli (verified live; interactive is cli); the notice is skipped for sdk-* entrypoints and the 30-day dedupe stamp is not written, so a headless run does not use up the interactive user's slot. JEV REVIEW DUE and the doctor line are unchanged; Codex is unchanged. Unset, context.protocolLevel=full makes the default true.

0.200.0 (2026-10-04)

Highlights

  • Less injected text. In the deterministic injection profile (synthetic payloads, chars not tokens, not a per-run saving) the Claude main-session text per context epoch goes from 26,370 to 13,068 chars (49.6%), SessionStart hooks (Claude and Codex) from 16,562 to 10,740 (64.8%), and the SubagentStart text from 7,701 to 3,371 chars in a normal spawn (43.8%). One setting rolls it back byte for byte: context.protocolLevel=full.
  • DevSwarm child workspaces no longer get the COMMS OVERRIDE / SELF_CONTINUE / REMINDER block on every turn (keepalive re-send on change, after compaction and every N turns).
  • Coordinator drift guard (coordinator-work-guard) and Bash edit parity: the main thread is nudged, then blocked, when it keeps doing state-changing work inline; Bash writes are judged like the Edit tool.
  • Codex guard parity: edit-guard, api-guard and ship-it-guard run on apply_patch; command-guard's heavy-command gate and its block paths work on the Codex main thread (stderr reason).
  • Codex wording: Codex sessions get Codex vocabulary in task-tracker, verify-first (orchestration rules and model routing: spawn_agent plus gpt-5.6 frontier/workhorse/fast tiers) and codex-availability instead of TaskCreate, run_in_background and Haiku/Sonnet/Opus; Claude text is byte-identical. The compact core now says "re-sent at session start and after compaction"; codex/README.md documents the context.protocolLevel=full rollback.
  • Jev fixes: speculationFramed is switchable, modelRouting records disagreements, triage no longer poisons its cache, arrival labelling.
  • Fewer false blocks: tasklist-guard, command-guard (read-only forms, pipe sinks), git-guard-adjacent push forms, AGENT-ROUTING and the shared-tree note.
  • Friendlier messages: block messages lead with the path that works and give an absolute, shell-quoted skip command.
  • Settings: the plugin options screen is down to 14 options; everything else lives in /anti-hall:settings, existing values migrate.
  • Benchmark tooling: injection profile with goldens, strength-study harness, scripts/hook-latency.js.

Fixed

  • Jev triage cache repair + prune. Entries cached before this release as a permanent no-label verdict (bare {_seq}, no label) are dropped by update and doctor --repair (repair-jev-triage-cache; real no-label verdicts now carry nl:true and are kept) so those messages are re-triaged. Stale triage claim files, a dead arrival worker's lock and an abandoned arrival queue are pruned (throttled, nothing else touched).
  • Guard block messages are readable: one shared shape (hooks/lib/block-message.js: a leading icon, anti-hall · <guard>: <what>, then Why: / Do instead: / Allowed here: / Override (only if the user explicitly asked):) replaces the ALL-CAPS walls of text in edit-guard, command-guard, git-guard, coordinator-work-guard, model-routing-guard, tasklist-guard, task-guard, speculation-guard, silent-agent-nudge, inbox-read-guard and the DevSwarm parent/child gates. Every fact the model needs (override command, TTL, exempt paths, guard name) is kept; Codex still gets exit 2 + stderr and no Claude-only tool names. Fixed icon set: ⛔ blocked, ⚠️ warning, 💡 tip, ✅ done, ⬆️ update, ❌ error.
  • The rest of the messages use the same shape: installer and CLI error lines (install-codex, statusline installers, ingest/reaper/supervisor forced-dry-run notices, jev-setup, defect, auto-handover-config, devswarm warnings), the handover-resume, fable- and codex-availability notes, and every DevSwarm injection segment (DEVSWARM <KIND>: banners are now <icon> anti-hall · devswarm-<kind>:). One icon, no ALL-CAPS banners; every fact, command and override variable is kept. Statusline rendering is unchanged (it emits no warnings).
  • Codex: command-guard and edit-guard block text no longer tells the model Claude-only things (a scratchpad script, run_in_background, a Haiku subagent). A Codex payload now gets spawn_agent / gpt-5.6-luna wording and no scratchpad advice; Claude text is byte-identical (golden test). Exit 2 + stderr is unchanged (#94).
  • DevSwarm child workspaces no longer get the COMMS OVERRIDE / SELF_CONTINUE / REMINDER block on every delivered turn: devswarm-child-turn.js now uses the guards.injectionRepeatEvery keepalive (re-sent on change, after compaction, and every N turns), matching the parent hook. The Codex port runs the same hook.

Added

  • Coordinator work window (coordinator-work-guard). The main thread keeps doing state-changing work inline instead of delegating it. A new hook counts successful state-changing Bash calls (WORK) over a 10-minute window. It adds one advisory note when the count reaches 4, and blocks the 7th WORK call in the window. Nudge delivery: PostToolUse additionalContext was live-observed delivering on Claude Code CLI 2.1.238 and is not doc-confirmed (docs/KB-claude-codex.md §1.4); re-verify after a CLI upgrade; blocks are the enforcement.
  • WORK is: state-changing git, a gh mutation, a Bash write into a repo file that is not a notes file, a script-file run, and inline -c/-e code that writes files or runs state-changing git/gh.
  • Recovery commands (git am|rebase|cherry-pick|revert --abort|--quit, git merge --abort, git stash pop|apply) and loosely matched inline code are counted but never blocked. Precise inline git/gh/repo-write code is blockable.
  • Script runs: scripts in the session scratchpad, a tmp dir outside any git work tree, or .anti-hall always count. A script written or modified this session counts, unless it sits in a package-manager or system location. A tracked, clean project script never counts. In a non-git project only a fresh or coordinator-writable script counts; an old one does not.
  • Observe-only mode: set guards.coordinatorWorkNudgeAt and guards.coordinatorWorkBlockAt to 0 and the guard only records state. guards.coordinatorWorkWindowMinutes 0 turns it off. Four new settings: guards.coordinatorWorkWindowMinutes, guards.coordinatorWorkNudgeAt, guards.coordinatorWorkBlockAt, guards.coordinatorWorkMaxEntries.
  • Skip key coordinator-work-guard (covered by the all skip). The guard is also off when safety.commandGuard is off or command-guard is skipped. Claude host only: Codex PostToolUse behaviour and Codex coordinator detection are unverified, so the hook is not registered on Codex.
  • Metrics: per plugin version, the posted share work / calls and the attempted share (work + blocks) / (calls + blocks), blocks, and a separate skipped-would-block counter. Counts survive pruning of old session files through a locked fold. dispatch-report and doctor show them.
  • scripts/coordinator-work-baseline.js <transcript.jsonl> [--from-line N] [--cwd DIR] [--json] replays a session transcript through the classifier and window, to get a "before" number.
  • Known gaps, all documented in docs/GUIDE.md:
    • obfuscated or computed inline bodies are not detected, and loose inline matches are count-only;
    • a binary compiled during the session (go build -o /tmp/p) is never counted;
    • a script run counts when it is a tracked script you modified this session. That is by design;
    • freshly generated, gitignored build launchers (./build/install/app/bin/app) count;
    • a fresh script inside a git submodule counts, because git status from the parent fails on that path;
    • deliberate mtime back-dating (touch -t, touch -d, cp -p) hides a fresh script;
    • a fresh script dropped into a package-manager or system location is not counted, and neither is one in a fake managed directory made outside the repo (for example ~/x/node_modules/.bin/p.sh);
    • an old, untracked, non-coordinator-writable script inside a repo is not counted;
    • text scripts under ~/Library or ~/.claude/plugins count if they are updated mid-session (an SDK install or plugin update);
    • just and task runs, git stash pop|apply landing edits, and cd "$X" && ... with an unknown cwd are not counted;
    • a trusted ./gen.sh > docs/api.md stays blocked by the Bash edit parity check;
    • the baseline resolves $VAR paths from its own environment, cannot classify direct-exec scripts that no longer exist, and judges freshness against current mtimes;
    • scripts fed on stdin are not counted: python3 - <<EOF, node <<EOF, sh <<EOF, echo … | sh, bash -s <<<;
    • wrapper forms hide the git/gh verb: time -p git, env -C d git, gh -R o/r pr merge, gh api -XDELETE;
    • writes via >|, a clustered cp -rt (wrong target read), install, dd of=, truncate, ln -sf and rm are not judged as repo writes;
    • git checkout . and git checkout <file> are not counted;
    • with a plugin path containing ', the printed skip command itself counts as WORK.
  • Bash edit parity (guards.bashEditParity, default on). In the main thread, command-guard applies edit-guard's verdict to Bash writes (sed -i, perl -i, tee, cp, mv, >/>> redirects) into repo files. Notes files the coordinator may edit stay allowed. Bash writes are judged like the Edit tool: a repo file the Edit tool may not write (including gitignored outputs like build/ or .env) is blocked; write under .anti-hall/ or the scratchpad, or delegate. Claude host only: Codex coordinator detection is unverified.
  • Codex: edit-guard, api-guard and ship-it-guard run on apply_patch edits (Codex 0.134 or later). Codex reports a file edit as tool_name: "apply_patch" with the raw patch in tool_input.command; the new hooks/lib/codex-apply-patch.js ports Codex's own patch parser (it agreed with the real apply_patch binary on 74 of 74 valid and malformed fixtures) and the guards check every added, updated, deleted and moved-to path. On Codex the main thread is a payload without agent_id/agent_type (confirmed from captured codex-cli 0.160.0 payloads); a Codex process started from inside a Claude Code session is treated as a worker. edit-guard blocks a patch it cannot parse on the main thread; api-guard and ship-it-guard let it through. A Codex block also writes the reason to stderr: Codex reads stdout JSON only on exit 0 and ignores an exit 2 with empty stderr, which a live codex exec run showed (the stdout-only block let the edit through; with the stderr reason it was refused, and a subagent's edit still went through). ship-it-guard's plan-conformance advisory stays Claude-only. Shell writes (cat >, tee, sed -i) still bypass all three on Codex. Claude behavior is unchanged: 353 Claude-shaped test spawns of the three guards gave byte-identical exit code, stdout and stderr before and after.
  • Docs: docs/KB-claude-codex.md no longer says Codex PreToolUse rejects additionalContext (supported since rust-v0.129.0).
  • settings.js judge on|off|status (and the /anti-hall:settings skills on both platforms): one verb to switch the opt-in semantic speculation-judge (jev.semanticJudge) on or off. on checks whether an Anthropic key is visible to the CLI process (never prints it; a key stored as a Claude Code plugin option is visible to hooks only, so the CLI says "not visible", not "absent") and prints how to add one plus the cost estimate (about $0.0001–0.001 and 1–3 s per turn end, estimated, not measured; no precision eval yet). status and on also report which backend is the semantic judge right now: Jev (when jev.enabled and the speculation integration is on, speculation-guard already asks Jev and the paid API judge exits early, as before), the Anthropic API judge, or lexical-only. doctor prints one info line with the active backend, or the enable command while the judge is off. No new userConfig option and no default changed.
  • scripts/hook-latency.js: a pure-Node hook-latency benchmark. It runs every command in hooks.json the way Claude Code does with a realistic payload per scenario, and reports wall p50/p95, CPU (process.resourceUsage() via a probe) and the per-tool-call total per event (--json or a markdown table). First measured numbers are in docs/HOOK-LATENCY.md, and the README no longer says latency is unmeasured.
  • Docs: a GitHub Pages site built from README.md and docs/*.md by the dependency-free tools/build-site.js; .github/workflows/pages.yml deploys it on pushes to main only (needs Settings → Pages → Source: GitHub Actions once). The site root also serves llms.txt, sitemap.xml and robots.txt.
  • Benchmark tooling (evals/anti-hall/). injection-profile.js is a deterministic injection-size profile (synthetic payloads, fixed root) that compares two plugin trees per channel (session start, skill listing, subagent start, Stop nags) and gates the cost-trim ratios; frozen normalised hook-output goldens (tests/fixtures/cost-trim/goldens/) with a privacy gate for the fixtures. A strength-study harness (arms, interleaving, global spend cap, trace metrics, study analysis, compaction ladder, probes) for the B0-B4 studies. Tooling only; nothing ships in the hook path.

Changed

  • The SessionStart verify-first core is now compact by default (context.protocolLevel=compact). It keeps every load-bearing clause inline (Iron Law, rationalization table, the seven rules incl. the done/hedge/merge block, scope and autonomy, the skip-guard clause) and points at the new generated PROTOCOL.md for the full wording. The orchestration rules A-N still arrive in full at SessionStart. Method and figures (deterministic injection profile, evals/anti-hall/injection-profile.js, synthetic payloads, 64-char root; a size measure, not a per-run saving): the SessionStart hooks (Claude and Codex alike) send 10,740 chars instead of 16,562 (64.8%); the core hook alone goes from 8,947 to 3,125 chars. The skill-listing trim (previous entry) adds to the per-epoch total. The SubagentStart text is compact too: the same core (subagent header, SKILLS: root-cause, deadly-loop) plus a short WORKER block (no re-delegation, the assignment is the authorization, tight summary, SendMessage before finishing) instead of the DISCIPLINES block and teammate note: 7,701 to 3,371 chars in a normal spawn (43.8%), 7,947 to 3,617 in a DevSwarm child workspace (child mailbox note unchanged). Same method as above (a size measure, not a per-run saving). context.protocolLevel=full restores today's subagent text byte for byte. Codex has no SubagentStart hook and is unaffected.
  • Experimental, opt-in: spawn-time delivery of the orchestration rules (context.orchFullOn=spawn, env ANTIHALL_ORCH_FULL_ON; values auto (default, = session), spawn, session, off). With spawn, verify-first-orch.js sends ORCH_COMPACT and a per-session marker (~/.anti-hall/orch-full/) and the new orch-on-spawn.js (PreToolUse Agent|Task|Workflow) is meant to send the full text once per context epoch (O_EXCL claim, silent for subagents, re-armed by compaction/clear/resume; needs positive Claude evidence: --host=claude in the Claude hooks.json command, a session_id, a transcript canonically inside <config>/projects/). It is NOT the default because it is unverified live: that a PreToolUse context reaches the coordinator, the real transcript line shape and Workflow-spawn behaviour have not been probed. Codex, unconfirmed platforms and a DevSwarm Primary always get the full text at SessionStart.
  • One-key rollback: context.protocolLevel=full (env ANTIHALL_PROTOCOL_LEVEL=full) restores today's text on every channel byte for byte (pinned by the frozen goldens in tests/fixtures/cost-trim/goldens/). docs/CONTRACT-1.0.md's "never silently shortened" sentence now names this deliberate exception. The legacy eval/run.js harness is frozen on full.
  • Tests: hooks.json command parsers in manifest-drift, codex-jev-hooks-parity and node-hook-flags accept trailing args (a non-match is now a failure, not a skip); node-hook-flags takes a per-script arg allowlist (--audit, and --host=claude only for verify-first-orch.js in the Claude hooks.json).

  • Third-party plugin-scanner (plugin-scanner 3.18.0) readiness: test fixtures no longer look like hardcoded secrets or dynamic code execution, the companion's ingest-daemon marker constant is renamed REAPER_DAEMON_MARKER (same value), and a comment no longer trips the eval heuristic. Added plugins/anti-hall/SECURITY.md, plugins/anti-hall/.codexignore, .github/dependabot.yml (github-actions, weekly) and a root package-lock.json.

  • Codex manifest: interface.composerIcon and interface.logo now point at assets/icon.png (the AH icon), and interface.screenshots lists two real captures (assets/screenshot-claude-code-session.png, assets/screenshot-git-guard-blocks.png), so plugin-scanner 3.18.0 scores the repo 100/100. The Claude manifest still has no icon key.
  • command-guard's anti-hall-plugin check reads only .claude-plugin/plugin.json when it walks up from a script (every anti-hall copy, the Codex install included, ships that manifest), and the Codex skills check the plugin root with test -d .codex-plugin instead of naming the manifest file. No hook or script names the Codex manifest file any more, so the image paths in it are not reachable from the hooks.

  • Background scratch-script runs allowed by guards.allowBackgroundScratchScripts, scripts written during the session, modified tracked scripts, other text-script runs outside package-manager/system locations, and inline -c/-e code that writes files or runs state-changing git/gh now count toward the main-thread work window. Running a tracked script you modified this session counts by design (the verify loop belongs to the subagent). The scratch hint now reserves scripts for read-only output capture. Patch application is for integration agents: the coordinator gets no exemption for git am or git apply, apart from the recovery commands above.

  • Contributing: day-to-day work lands on dev; main changes only through a pull request from dev.
  • Docs: docs/CONTRACT-1.0.md drafts the 1.0 contract, the settings, CLI verbs, hooks and state paths that semver will freeze.
  • The plugin options screen is down to 14 options: the manifest userConfig goes from 129 to 14: the 10 headline switches (the four safety.* guards, auto-handover on and threshold, jev.enabled, devswarm.supervisorMode, guards.modelRouting, limitConserve.mode) and the four API keys. Every other setting is reached through /anti-hall:settings, grouped by category (settings.js show, then show --section <category>).
  • Values you already set keep working: a non-default stored value is copied into ~/.anti-hall/settings.json by the update or doctor --repair run of the first release that runs the migration (this one, for an install coming from 0.121.x), and every removed option stays a read-only source below the settings file (a stored plugin option or a still-exported CLAUDE_PLUGIN_OPTION_* resolves exactly as before; a value equal to the schema default counts as unset). Defaults are unchanged. An older plugin version resolves the same values, because the settings-file value wins.
  • Safety options you set in /config survive their row leaving it: the four locked options that lose their row (guards.stashGuard, guards.editGuardAllow, guards.allowSubagentMailbox, devswarm.maintainerNotice.post) are now copied too, as human-confirmed values (the same write settings.js set <key> <value> --confirmed makes), because a value stored in Claude Code's plugin options can only have come from the person's own /config choice. Before, they were skipped and stayed readable only from Claude Code's stored options, so an armed stash guard would have fallen back to off if those were ever dropped. The copy still happens only when the stored value is the current effective one, so no guard changes behaviour.
  • The settings skills, the system-briefing skills, the guide, llms.txt, AGENTS.md, CONTRIBUTING.md and the Codex README describe the grouped show flow and say that /config only has the headline switches, the safety guards and the keys.
  • Tests: 320 test files that read home-dir state (settings, caches, registry) without their own HOME now require tests/helpers/isolate-home.js, which points HOME/USERPROFILE at an empty temp dir, so they can no longer read the real ~/.anti-hall or ~/.claude.
  • Skill descriptions are trigger-first and at most 200 characters on Claude and Codex; the long text moved into a When to use section of each skill. Skill-listing injection in the profile: 9,808 to 2,328 chars (synthetic fixture; a size measure, not a per-run saving).
  • Settings added this cycle: context.protocolLevel (compact default, full = rollback), context.orchFullOn (auto default = session), guards.coordinatorWorkWindowMinutes, guards.coordinatorWorkNudgeAt, guards.coordinatorWorkBlockAt, guards.coordinatorWorkMaxEntries, guards.bashEditParity, guards.sharedTreeAgentNote. New hook: orch-on-spawn.js (PreToolUse Agent|Task|Workflow, inert unless context.orchFullOn=spawn). New generated file: PROTOCOL.md (from tools/gen-protocol.js).

Fixed

  • coordinator-work-guard no longer loads the 237 KB command-guard.js for plain read-only Bash commands (git status, ls, cat | head, ...): a closed-vocabulary pre-filter (provablyNotWork in lib/coordinator-work.js, differentially tested against classifyBashWork) answers "not work" first. PreToolUse CPU p50 for that hook dropped from 53.4 to 45.0 ms in an interleaved before/after run (load 220); other commands classify exactly as before. The Node compile cache was measured (0 to 6 percent CPU change) and not adopted; see docs/HOOK-LATENCY.md.
  • Jev audit snippets for outputVerifyGuard now keep the first 200 + last 400 characters (joined with …, redacted as a whole, total <= 700) so the test-runner pass/fail summary at the end of the output is judgeable; other integrations' snippets are unchanged.

  • Jev: speculationFramed is now registered in jev-setup.js KNOWN_INTEGRATIONS, so jev status / jev mode can see and switch it (a registry test derives every call-site id and asserts it is in KNOWN_INTEGRATIONS and the settings schema). modelRouting decisions now record wouldChange plus an audit snippet whenever Jev's tier differs from the rule-based verdict (also in on mode, regardless of confidence), so jev report can count and label them.

  • command-guard: a pipe sink that decides an allow (the gcloud read carve-out and the bounded verification pipeline) now has to match the closed stdin-only grammar. ... | head /etc/passwd and ... | tail -n +1 --pid=123 no longer pass as bounded sinks; head|tail -c N and grep -m N -E PAT stay accepted.

  • A read-only gh api graphql call is no longer treated as a mutation. A graphql call is heavy unless it is proven to be a read: one query= field with no $, backtick, leading @ or mutation, and only read flags. The attached forms (-fquery=..., --raw-field=query=..., -F<x>, --field=) and --input now count as body flags on every endpoint; before, -fquery=..., --raw-field=query=... and --input b.json were not blocked.

  • Block messages give an absolute, shell-quoted skip command. The skip hint in edit-guard and the new guard is built from the plugin path, single-quoted, so a path with a space or $ works. A relative path broke when the cwd was not the plugin root.
  • Codex: command-guard's heavy-command gate now runs on the Codex main thread. Coordinator detection keyed on CLAUDE_CODE_ENTRYPOINT, which Codex does not set, so on Codex the heavy-command gate never fired (a live main-thread npm test ran). hooks/coordinator-detect.js now recognises any Codex payload (a non-empty turn_id plus a non-empty model, or an apply_patch call) and treats it as the main thread unless it carries agent_id/agent_type (a spawned subagent) or the process inherited CLAUDE_CODE_ENTRYPOINT (Codex started from Claude Code as a worker). Partial or mistyped payloads fail open. Verified live with codex-cli 0.160.0: a main-thread npm test is refused and does not run, and a subagent's npm test runs. Claude Code behaviour is unchanged: 2,140 command-guard/edit-guard spawns from the existing tests gave identical exit codes, stdout and stderr before and after.
  • Codex: command-guard and compact-declaration-guard now actually block. On exit 2 Codex takes the block reason from stderr and ignores stdout JSON; with stderr empty it records a failed hook and runs the tool call (codex-rs/hooks/src/events/pre_tool_use.rs, rust-v0.160.0). Both guards blocked with stdout JSON only, so on Codex they blocked nothing: a live codex exec main thread ran an armed git stash push. Their block paths now go through the new hooks/lib/emit-block.js, which writes the same reason to stderr; the same run is now refused. A hygiene test fails on any exit-2 path in a Codex-registered hook that leaves stderr empty. Claude shows the JSON reason when there is one (Claude Code hooks reference, "Exit code 2"), and the stderr text is identical, so Claude output is unchanged.
  • Docs: docs/CONTRACT-1.0.md and the Codex parity notes no longer say Codex has no SubagentStart event. A captured codex-cli 0.160.0 payload has hook_event_name: "SubagentStart" with agent_id/agent_type. verify-first-subagent is still not registered on Codex.
  • The root package.json / package-lock.json no longer read as 0.0.0. In the marketplace clone the root test-harness package was the first version file an agent found, so peers reported the installed plugin as "0.0.0 local dev build". Both files now carry the plugin version and a description pointing at plugins/anti-hall/.claude-plugin/plugin.json; tests/hygiene/manifest-drift.test.js pins package.json, both package-lock.json version fields and both plugin manifests to one version, and RELEASING.md lists them in the bump step. Nothing in the plugin reads the root package.json.
  • The Codex installer now registers the PostToolUse git-guard.js --audit hook. install-codex.js omitted it while the shipped codex/hooks/hooks.json template had it since the self-credit audit landed, so project/global installs never ran the audit; tests/codex/codex-hook-parity.test.js missed it because it compared hook basenames only, and now compares full normalized commands (flags and arguments kept) per event.
  • tasklist-guard no longer says "NO background agent is live" while agents run: it now reads running agents from the session transcript (as task-guard does) instead of the global 20-minute heartbeat, and an unknown count is treated as not stale.
  • speculation-guard: lock adversarial regression corpus. tests/hooks/speculation-guard-regression.test.js pins 67 adversarial and 19 evasion diagnoses (labelled, bounded, plan-worded, requirement-prefixed) plus 20 plain ones as must-block; test-only, no behaviour change.
  • command-guard false blocks on read-only forms (approved only when the WHOLE command is the read form, optionally piped into tail/head/wc/grep -c/-m; no substitution, ;/&&/||/newline or other redirection): firebase --version/-V (and the same bare version flag on the other deploy/infra CLIs) passes; gcloud describe/list/get/read passes with a separated value for --project/--region/--zone/--location/--limit/--freshness/--page-size/--sort-by (any other separated flag value, --format json included, stays blocked); a plain push chain may end with git ls-remote <configured-remote> [ref] or git rev-parse <ref>, and a plain push may be piped through | tail/head -N mid-chain. Deploy/delete/push-of-anything-else verdicts are unchanged.
  • AGENT-ROUTING no longer suggests Explore for prompts that tell the agent to write or mutate. The model-routing-guard write-signal list missed instruction forms such as save ... to <path>, clone <repo> into <path>, git format-patch, create new files and run the generator/tests/build, while its ambiguous words (build, release, tag, patch, fix, install) also suppressed read-only prompts that only mention them as nouns. Those words now count only in instruction position, an explicit read-only statement ("report only", "do not edit", "read-only", "no edits") overrides the bare stems, and the instruction forms above always suppress.
  • SHARED-TREE note no longer fires for spawns that work in a scratch dir. guards.sharedTreeAgentNote now stays silent when the new spawn, or the running agent, establishes a scratch work location ("work in a scratch clone", or work in / cd into / cwd = a /tmp, /private/tmp, /var/folders or scratchpad path), since those spawns share no git tree. A bare mention of a scratch directory or tmp path, an in-place edit statement ("in the repo", "in place", "working copy", "checked-out files", "on main") or a negation ("not a scratch clone") still warns.
  • The per-prompt task line no longer counts owner/external-blocked tasks. task-tracker now reuses the Stop task-guard's isOwnerBlocked predicate: open tasks: 2 (+1 blocked: external), and when only blocked tasks remain the "update or close them" instruction is dropped. Codex shares the same hook.
  • "Running-agent count unknown" no longer fires for agents that launched and finished in the same session. Root cause: an empty running list on a transcript larger than the 1.5MB tail was always treated as unprovable. agent-scan now makes one widened (12MB) read to prove the count, and when it still cannot, the line names the agent ids seen and the window it could check.
  • Canonical ANTIHALL_SCAN_THROTTLE and ANTIHALL_SESSION_END_REAPER env names (old ANTI_HALL_* names still read as deprecated aliases via schema envAliases; canonical wins); devswarm.js help now covers inbox tick, primary, ready-check and describes ensure as idempotent; settings-schema wording for reset confirmation and fable-availability/codex-availability; KB-claude-code-hooks.md and GUIDE now match the official hooks doc (over-cap hook output spills to a file with a 2,000-char preview, it is not truncated).
  • A failed unread-summary refresh after an ack is now recorded. applyReadAckOps swallowed a summaries/<hash>.json refresh failure; it stays fail-open but now logs op: 'summary-refresh' with the path and error to the bounded cursor-log.
  • limit-conserve threshold now resolves at call time. isConserving() read the module-load value when called without home, so a settings.json or env change after require() was ignored; it now always reads the unified settings store (the Codex status script shares this module).
  • update/doctor no longer tell an idle Primary to arm a wake watcher that would exit at once. update.js (wakeMonitorPostUpdate) and doctor's wake-monitor check advised "NOT live — arm it" while the watcher itself idle-skipped (no live non-held/non-ignored child), so the agent armed it, it exited, and every inbox tick kept saying idle-skip. All three now use one shared decision (idleSkipApplies in companion/lib/devswarm-live-children.js): when it applies they report "wake watcher not needed now (no live child workspaces; the mailbox tick covers you)" and add idleSkip: true to the status (existing fields unchanged); unknown liveness or devswarm.wakeWatchIdleSkip off keeps the arm advice.
  • Jev: triage no longer caches budget-skipped messages as a permanent no-label verdict (a bulk read used to fill the 500-entry cache with empty entries and evict real labels; attempted items are now an explicit cacheable null, unattempted ones retry). Concurrent triage of one message no longer calls Jev twice (per-hash O_EXCL claim, post-claim cache re-check, merge-on-write).
  • Jev: a direct message is now labelled at arrival, so the parent gate's cache-only question lookup finds a label while the message is still unread. Arrivals go through one bounded drain worker per HOME (capped queue file, O_EXCL lock stale after 30 s); a burst of 50 sends spawns one worker, not 50.

0.122.2 (2026-10-03)

Fixed

  • Two tests that failed only on Linux CI are now platform-independent (the command-guard message-order differential and the wake-watch lock-stale-bound test). Test-only; no plugin behaviour change.

0.122.1 (2026-10-03)

Added

  • Shared-tree agent note: the new setting guards.sharedTreeAgentNote (default on; env ANTIHALL_SHARED_TREE_AGENT_NOTE; settings file and env only) makes swarm-guard add one advisory sentence when a write-capable subagent is spawned without isolation: "worktree" while another write-capable agent is still running in the same working tree (two such agents can commit each other's uncommitted changes). Read-only agent types, isolated spawns and unknown agent state stay silent. Never blocks.

Changed

  • command-guard "allow plain push" accepts -u / --set-upstream. A session pushing its own branch with git push -u origin <branch> [2>&1 | tail -N] was blocked only because of the upstream flag. The flag is the single allowed flag slot (never combined with -q or any other flag) and requires an explicit remote AND ref; remote/ref vetting, the output-sink rules and the chain allow-list are unchanged (rev-parse/ls-remote/node/echo after a push stay blocked). git-guard.js is untouched.

Fixed

  • doctor: the live guard self-tests now run against an isolated temp home instead of the caller's ~/.anti-hall, so an unexpired user skip or an exhausted Codex quota no longer makes a working guard's self-test report FAILED. The temp home is removed when the run ends; the context-footprint measurements keep the real home on purpose.
  • Transcript-heavy hooks no longer hang at exit on Node 24+: the 19 hooks that may parse 1.5 MB or more of the session transcript before exiting (ask-guard, auto-handover, auto-handover-pause-nag, claim-ledger, codex-nudge, compact-advice-guard, compact-declaration-guard, devswarm-child-turn, devswarm-parent-inbox, dispatch-tier, edit-guard, limit-conserve-inject, precompact-snapshot, silent-agent-nudge, stale-agent-stop-note, task-guard, task-tracker, tasklist-guard, verify-first) now run node --no-concurrent-recompilation --no-concurrent-sparkplug. This includes the four that use emit-dedupe, whose scan widens to a 4 MB tail. That is 19 of 64 Claude entries and the 14 matching Codex entries, including the commands codex/install-codex.js writes. Every other hook command is unchanged. On Node 24 and 26 such a hook could deadlock inside process.exit(): a V8 background compile (Maglev or Sparkplug) waited for a main-thread garbage collection while the main thread waited to join that thread (nodejs/node#54918, #64274). The hook then sat until its timeout (30 s for a Stop hook). Replaying a real 14.8 MB transcript through silent-agent-nudge on Node 24.14.0: 4/600 hangs without the flags, 0/600 with them. Node 22 never hung (Maglev is off there). The flags and the hook list live in hooks/lib/node-hook-flags.js. doctor reports whether the running Node accepts the flags, and a hygiene test checks that exactly the listed entries carry them. Existing Codex installs pick up the flags on the next install-codex.js run.
  • task-guard idle-neglect message: it now says how to mark a task an agent already covers (set the task's owner to the running agent).
  • Heavy-command block message leads with the path that works: every variant of the command-guard.js block text (plain, DevSwarm child, DevSwarm Primary with and without the workspace advice) now starts with one line, to run or re-check the command yourself write it to a scratchpad script and run <interpreter> <script> with run_in_background, then the rest, tightened (about 15% shorter for the plain text). Message only: what is blocked and what is allowed are unchanged, proven by a before/after exit-code table over 49 commands.
  • The compact roster unread column summed direct and broadcast unread, so a Primary with no direct mail showed e.g. 34 while inbox tick / read-primary said 0. It now shows direct unread only (the tick's definition) and appends (+N bcast) only when broadcasts are unseen. --json is unchanged (directUnread stays direct-only; broadcastUnread stays separate).
  • Dispatch-tier hints are calibrated toward subagent: a workspace or workflow recommendation with Jev confidence below 0.6 is now shown as subagent (logged as low-confidence-subagent), and the classification question now says that a bug fix, a UI text change, any single-file or single-component task, and a priority label (P0/P1/P2) do not make a task a workspace or workflow task. Confident multi-step feature, release and sweep verdicts are unchanged. Verdicts already cached keep their stored answer, but the confidence floor applies to them too.

0.122.0 (2026-10-03)

Added

  • Optional "no blocking questions" guard (off by default): the new setting guards.noBlockingQuestions (off, advise or block; env ANTIHALL_NO_BLOCKING_QUESTIONS) watches the AskUserQuestion tool. advise lets the question through and reminds the agent to take the recommended option, say which one, and carry on; block refuses the call. A question whose first entry starts with DESTRUCTIVE: or CREDENTIAL: (in the header or at the start of the question text) is always allowed and noted in ~/.anti-hall/logs/ask-guard.ndjson. In a DevSwarm child workspace the message also says to send the question to the parent instead. Claude only (Codex has no ask tool). In block mode the skills that deliberately ask the user (settings, deadly-loop, ship-it, flutter-debug, devswarm) need that marker, or the setting turned off.
  • Agents-in-flight note on questions: the new setting guards.questionAgentsNote (default on; env ANTIHALL_QUESTION_AGENTS_NOTE) makes the same AskUserQuestion hook add one advisory line when background agents are provably in flight, naming how many and their short descriptions and saying they may act on an option before the answer arrives. Silent when none are running or the count cannot be proven. Never blocks and works with guards.noBlockingQuestions off. Claude only (Codex has no ask tool).
  • Dispatch-tier text off switch and no-workspace-repo suppression: the DevSwarm Primary dispatch-tier text (the workspace is the top fan-out tier) now goes through one gate used by task-tracker.js, verify-first.js and verify-first-orch.js. It is no longer injected in a repo whose CLAUDE.md/AGENTS.md forbids workspaces for real work, and the new setting devswarm.dispatchTierText (default on; env ANTIHALL_DEVSWARM_DISPATCH_TIER_TEXT) turns it off everywhere.
  • Inline-work nudge for Primaries: the new setting devswarm.inlineWorkNudge (default on; env ANTIHALL_DEVSWARM_INLINE_WORK_NUDGE; threshold devswarm.inlineWorkNudgeThreshold, default 5) adds one note per session, via the existing edit-guard.js hook, when a Primary has made more than the threshold of direct edits while actionable tasks are pending and it has proven zero live child workspaces. Silent when liveness is unknown, in a child workspace, and in a no-workspace repo. Never blocks.
  • Interval roster on the cron tick (off by default): the new setting devswarm.tickRosterEvery (default 0; env ANTIHALL_DEVSWARM_TICK_ROSTER_EVERY) makes every Nth inbox tick --quiet of a Primary append the compact roster table after the unchanged first line, when unread is 0 and a live child is proven. The JSON form of tick and --child ticks are unchanged.
  • Pending child questions are harder to miss: when a Primary has unanswered child questions, the per-prompt "own inbox" notice now starts with one line, QUESTIONS AWAITING YOUR REPLY: N (oldest Xm) — <workspace title>: <first 80 characters of the question>. The preview is cleaned of control characters and has secrets redacted; it is left out when the stored summary has no question text yet.
  • devswarm.js spawn note: when --source names a branch other than the default branch and the DevSwarm app already has a workspace (active or archived) on it, the result carries one line in warnings saying the app may show the new workspace nested under it. Output only; nothing else changes, and a failed lookup stays silent.
  • Stale-agent stop note: the new hook stale-agent-stop-note (PreToolUse TaskStop, advisory, never blocks) adds one line when TaskStop names an agent that was messaged or resumed after its last report and has not reported since. Setting guards.staleAgentStopNote (default on; env ANTIHALL_STALE_AGENT_STOP_NOTE; settings file and env only, no plugin option).

Changed

  • Advanced settings now live in /anti-hall:settings: the 25 advanced settings no longer have a /config row (the manifest userConfig goes from 154 to 129 keys). Non-advanced settings keep their /config rows; /anti-hall:settings shows and changes every setting. The new settings in this release (guards.noBlockingQuestions, guards.questionAgentsNote, guards.staleAgentStopNote, devswarm.dispatchTierText, devswarm.inlineWorkNudge, devswarm.inlineWorkNudgeThreshold, devswarm.tickRosterEvery) are advanced: set them via /anti-hall:settings (settings file or env); they have no /config row. dispatchTierText and inlineWorkNudge are independent switches.
  • Stored plugin-option values are copied into ~/.anti-hall/settings.json on update: plugin-option values are now migrated into ~/.anti-hall/settings.json for every plugin-option setting except the 10 headline switches, locked/home-only keys and the credential options, and only when the stored value is already the effective one (the migration never changes what resolves; it re-checks "still unset" under the settings lock). Stored options are read under both pluginConfigs keys (anti-hall@anti-hall, anti-hall) and both shapes. The 10 headline keys are flagged headline in the schema. New permanent default-equivalence test.
  • Both READMEs end with a "Links" section naming Documentation, Support and Privacy; documentationUrl now points at the documentation start page (docs/README.md).
  • Eight over-200-character tokens (seven regex literals in model-routing-guard.js, command-guard.js, claim-ledger.js, plus the command list in devswarm.js's unknown-command error) are rebuilt from short joined parts so the directory scanner can read them. Behaviour is unchanged; tests/hygiene/regex-source-equality.test.js pins each rebuilt RegExp.source and .flags to the original.
  • The icon key is removed from the Claude manifest. The plugin icon ships at the directory's default path, plugins/anti-hall/.claude-plugin/icon.png (512x512 PNG, found without an icon key), and a copy is kept at assets/anti-hall-icon.png. The listing icon is whatever was uploaded in the directory portal; the repository copy is read only at first submission. tests/hygiene/plugin-icon.test.js checks the icon and that no other image or font file is tracked under the plugin folder.
  • Skill and model-policy documents refer to this repository's own docs by plain repository path (docs/KB-….md) instead of a full URL. Affected: the devswarm, jev, ship-it and update skills, their Codex counterparts (devswarm, doctor, jev, update) and both MODEL-POLICY.md copies. Vendor documentation URLs in the Jev skills are unchanged.
  • The devswarm skills point at the archived orchestration design and plan under docs/archive/superpowers/ (the old path no longer existed).
  • The Jev skill leads with the plugin option for storing a key, marks the key file as an opt-in legacy path, and matches the real set-key output.
  • CI installs a pinned networkx in a virtual environment so the api-guard third-party tests run; a new hygiene test checks the paths declared in hooks.json, monitors.json, skills and plugin-root references.

Fixed

  • A DevSwarm Primary that has live workspaces but no way to be woken by their messages (its session cron died with a restart and its watcher exited when no child was live) is now told so. Each prompt shows a NO MAILBOX WAKE PATH line (or a shorter one naming just the missing watcher or tick), spawn adds the same instruction to its warnings, and the Stop gate blocks once per cap when both are missing. Session crons do not survive a restart; the line says what to re-arm.
  • The wake watcher's idle-skip line no longer says the cron covers you. It reports how long ago the mailbox tick last ran, or says no recent tick was seen and to check CronList.
  • A transient lock-heartbeat failure (the lock's reclaim sidecar briefly busy) no longer makes a healthy wake watcher exit. The watcher skips that restamp and retries next tick; it exits as lost only on a genuine lost lock, or when transient failures outlast the lock's 2-minute stale threshold (with a log line).
  • Releasing a lock now checks ownership and unlinks under the reclaim sidecar (waiting at most ~40 ms for a busy one, then the old behaviour), so a stealer that publishes at that moment cannot have its lock deleted.
  • Unanswered-question nag and Stop block skip archived-and-dead, held and ignored askers: the Primary's "N unanswered QUESTIONS" notice and "QUESTIONS AWAITING YOUR REPLY" line (devswarm-parent-inbox.js) and the Stop block on unanswered questions (devswarm-parent-gate.js) used to count questions from workspaces that are archived, held or archive-ignored. All three now share one filter and one set of inputs (filterLiveAskers, built on row-eligibility.js; active descriptors count as rows in both hooks): held and ignored askers are dropped; an archived asker is dropped only when its session is provably dead, so an archived-but-live child still blocks and nags (R18). A sender with no row stays counted when the summary is a legacy one without archivedRegistryRows (it may be archived-but-live), and is dropped as retired otherwise. The block's loop signature follows the filtered set.
  • "QUESTIONS AWAITING YOUR REPLY" no longer names archived, held or ignored children: the line now lists only questions whose sender row is not archived, held or archive-ignored (the parent gate's policy, via row-eligibility.js); a sender whose eligibility cannot be determined is left out. The rest of the per-prompt notice is unchanged.
  • Roster counter survives non-quiet ticks: a non-quiet inbox tick rewrote the wake-tick marker without its seq, resetting the devswarm.tickRosterEvery counter; the previous seq is now carried over.
  • Cron tick prompt wording: the prompt now says to decide on the FIRST printed line only and treats any lines after it (the optional devswarm.tickRosterEvery roster) as informational.
  • Roster note only in the Primary's tick prompt: the FIRST-printed-line wording is in the Primary's cron-tick prompt only. A --child tick never prints the roster, so the child prompt keeps its 0.121.8 wording and the injected child payload stays under its size budget.
  • Edit-guard and command-guard advice in no-workspace repos: the block message shown to a DevSwarm Primary no longer recommends spawning a child workspace when the repo forbids workspaces for real work (or devswarm.dispatchTierText is off); it goes through the same primary-tier.js gate as the other Primary tier text. Advice text only: the same edits and commands are blocked.
  • agent-scan: terminal evidence (a stop notification) stamped before an agent's latest resume no longer re-marks the resumed, running agent as finished; the resume check allows 2 s of clock skew. A named in-process teammate that was sent a message and has not reported since now counts as running (rows carry a pendingMessage marker).
  • agent-scan: a teammate report is recognised only as the harness-injected record. Typed or pasted text, a peer's cross-session message, a tool result, an assistant message, a report for a teammate this session never spawned, prose before the block, and an inner timestamp in the future of the record's own all no longer count, so a forged report cannot hide a running teammate. A teammate spawned outside the scanned window is "unknown", never "finished". A teammate named like a background agent id no longer replaces that agent's row or clears its terminal state.
  • silent-agent-nudge: a teammate whose only running evidence is an unanswered message never causes a Stop block, at any guards.silentAgentNudgeMin.
  • Codex: documented that stale-agent-stop-note is Claude-only (no agent-stop tool matcher is documented for Codex).
  • A "Repository not found" error from hivecontrol no longer silences a live workspace's mail for hours: live rows are always pulled and keep reporting the failure; archived, held or ignored rows (and a repo with no live row) are suppressed only after 3 identical sweeps in a row, and the match is limited to hivecontrol's own error line.
  • The workspace stop gate re-opens the stop budget for the kinds it just satisfied when nothing is owed any more.
  • The Primary gate says so (one non-blocking line) when it cannot determine its own unread count, instead of skipping its own row silently.
  • Key notices (hooks, jev-setup, migration) say "the plugin's options screen" instead of /plugin config.
  • Held and archive-ignored children no longer keep the wake watcher armed: the wake-watch and inbox tick idle-skip now use the parent gate's policy (live = not archived, not held, not ignored).
  • Archived or pruned workspace rows are no longer reported as errors by reconcile, and the supervisor log line lists an unchanged failure set once instead of repeating every row each sweep. Nothing is moved or deleted.
  • The workspace stop gate no longer re-demands a heartbeat beside an inbox pull after a report, and the Primary gate no longer nags a Primary whose own descriptor has no inbox path.
  • The one-time generic Jev key binding notice follows the effective transport (plugin option, env or settings), not only the home settings file.

0.121.8 (2026-10-03)

Fixed

  • Directory listing: the plugin icon moved to the plugin's top level (./icon.png) and is now an original "AH" monogram; the display name is "Anti-Hall"; the description is shorter; the marketplace entry has a homepage and an author URL.
  • Skill front matter is valid strict YAML (three skill descriptions were not).
  • Literal NUL bytes removed from two source files; they now use the \0 escape.
  • The monitor command quotes ${CLAUDE_PLUGIN_ROOT}, so it works when the install path contains a space.

Added

  • README and PRIVACY have a "What it runs and writes" table.
  • Hygiene tests: strict skill front matter, no NUL bytes in shipped files, icon at the plugin's top level, quoted monitor command.

0.121.7 (2026-10-03)

Changed

  • devswarm.js roster prints a compact table of live workspaces with a "+N archived" line; --all lists archived ones and --json gives the full data as before.

Fixed

  • The stale-handover reminder no longer makes a reply end with "refresh the handover first" after the handover was refreshed in that same turn.
  • The workspace stop gate's cap message no longer says there will be no further stop blocks; it now says the limit starts over after a report.

Added

  • The stalled-agent check records its phase timings when a run is slow, to diagnose a rare 30-second timeout whose cause is not yet known.
  • Orchestration guidance: Claude Code rejects subagent files named report, summary, findings or analysis; name them differently.

0.121.6 (2026-10-02)

Fixed

  • Stop-time task check: it no longer guesses the state of a task whose records are far back in a long session (it wrongly nagged tasks marked as waiting on the owner and re-opened finished ones). It now recovers the real state, and says so in one line when it cannot.
  • Stalled-agent check: an agent that was stopped is no longer reported as silent. A resume, a launch and a finish now count only from real system records, not from text that merely quotes them.
  • Quoted text is no longer mistaken for a real signal in four other checks: the merge gate (only something the user types can sign off a hedge; the assistant can no longer clear its own), the "safe to compact" declaration (table rows are ignored), DevSwarm idle tracking, and the task-list reset check.
  • Judge setting: the jev.semanticJudge setting now actually enables the optional judge. Previously only the environment variable did.
  • Docs corrected: CI matrix, release guide, Codex hook-event count, Windows wording, activate skill.
  • A docs link check no longer fails on local, git-ignored files.

Added

  • Handovers are never committed: git-guard blocks a commit that includes a handover (setting guards.handoverCommitGuard, on by default). Doctor lists tracked or out-of-format handovers, and the archived-workspace message names the exact handover path.
  • The manifest declares the plugin icon. A code-owners file.

Changed

  • New files named CONTINUE-HERE.md are blocked (existing ones stay editable). The deadly-loop state record moves under .anti-hall/.
  • Historical plans and audit reports moved to docs/archive/. The test workflow pins its actions and declares read-only permissions.

0.121.5 (2026-10-02)

Fixed

  • Claim ledger and second-opinion judge: both now check the reply actually being sent rather than the previous message, and a reply is never counted as evidence for itself.
  • Marketplace listing text: the listing now shows the current short description instead of an outdated long one.
  • speculation-judge reads the reply being stopped: the opt-in judge now sends the Stop payload's last_assistant_message to the model, not the previous transcript message. The transcript is only the fallback.
  • claim-ledger judges the right message: when the transcript is one message behind the Stop payload, the payload text is the reply, the transcript's last message counts as evidence, and tools_this_turn is the running tool count for this turn. The last message of a session is now judged too. When the transcript is up to date the record and its hash are unchanged.
  • Jev turn pointer: speculation-guard and claim-ledger no longer log a turnRef when the judged text came from the Stop payload, because the transcript's last line may be the previous turn. Nothing reads turnRef programmatically.

Changed

  • License file: the plugin folder now ships its own LICENSE file.
  • The privacy policy now states that the optional statusline reads your account email for display only.

0.121.4 (2026-10-02)

Fixed

  • Doctor twin-record repair: the repair stage no longer reports "failed" on every run. Records it leaves untouched on purpose are now listed once each as "needs review", with the record id and the reason.
  • Archived DevSwarm child: an archived child workspace is no longer woken by its own mailbox watcher, and its tick no longer asks it to re-arm. It is told to delete its own mailbox cron. A live workspace is never silenced.
  • Archived-workspace handover check: it now recognises a flat handover file.
  • Speculation check: it now judges the reply actually being sent rather than the previous message. That mix-up caused false blocks and let the real reply go unchecked. If the helper that reads the reply cannot load, the check falls back to reading the transcript instead of allowing silently.
  • Mailbox watcher handoff line: the "handed off to " line now says when that version is only cached and not yet registered.

Changed

  • Plugin icon: it is now 512x512 (under 256 KiB), so every shipped file is under the 256 KiB per-file limit.

0.121.3 (2026-10-02)

Fixed

  • DevSwarm ingest warning: the "DEVSWARM INGEST FAILING" banner now appears once per episode instead of on every prompt. When the DevSwarm app's hivecontrol is timing out, it says so (the DevSwarm app is not answering, so check or restart it) instead of telling you to run doctor, which cannot fix a timeout. The doctor advice stays only for a real configuration fault.

  • Stop-time open-task message: it now names tasks that were created long before, which previously showed as "(subject unknown)". The name is recovered only from a real, paired task-creation record; otherwise it still shows "(subject unknown)".

0.121.2 (2026-10-02)

Changed

  • Docs: the README is cut to about 70 lines. All the detail moved to docs/GUIDE.md; docs/README.md is now the single documentation start page, and the README links to it in one place. The plugin README and the Codex README are slimmed in the same way.
  • Demo: the README GIF is now a real Claude Code session. A force-push attempt is blocked by git-guard, and the agent falls back to a normal push. assets/demo/README.md documents how it was recorded.

Tests

  • Every doc must be reachable from the docs start page; the Jev recommendation's full text is pinned in the KB doc.

0.121.1 (2026-10-02)

Added

  • Jev: an optional backup vendor, jev.fallbackTransport (none|vercel|typesafe). It retries once when the primary times out, returns a server error, or is rate-limited or out of balance. It never retries on a rejected key. A per-vendor circuit breaker skips a failing vendor; when both vendors are failing, Jev is skipped. The docs note that the two vendors likely share a backend, so this is not full redundancy.
  • jev-setup enable --fallback.
  • Jev reporting: every decision records the vendor that served it, and whether the backup was used. The report and the weekly scorecard break down per vendor. The Vercel per-request cost is read.

Changed

  • Jev keys are vendor-bound: jev_vercel_api_key and jev_typesafe_api_key. The legacy jev_api_key is bound to the vendor in the home-only setting jev.genericKeyVendor (default vercel), which is changed with jev-setup bind-generic-key. A key is never sent to a vendor it was not entered for. Behaviour change: a generic key used with the typesafe transport is no longer sent until it is bound, or re-entered as the vendor key.
  • Jev: requests refuse redirects. Secret redaction now happens once, inside the client. Error bodies are never logged.
  • jev-setup reads and writes the effective configuration. enable previously wrote a file that a newer settings file overrode. status warns when the two config files disagree, and doctor checks it. The balance is shown per vendor ("n/a" where a vendor has no balance endpoint). set-key no longer prints the key length.

Fixed

  • Silent-agent warning: agents launched far back in a long transcript are now seen. A queued follow-up message no longer counts as the agent having finished. A resumed agent is tracked from the time of the resume.

0.121.0 (2026-10-02)

Changed

  • Internal: scripts/devswarm.js is split. It is now a dispatcher (about 80 KB) plus 15 modules under scripts/devswarm-lib/ (the largest is about 156 KB). It is a pure move: the export surface (198 names), the CLI behaviour and the help output are unchanged. Every file is now under 256 KiB, and a hygiene test enforces that.
  • Note for anyone patching the plugin locally: the verb implementations moved to devswarm-lib/.

Tests

  • The mutation kit and the source-text checks treat the dispatcher plus devswarm-lib as one source unit. An export-surface snapshot test was added.

0.120.15 (2026-10-02)

Fixed

  • DevSwarm roster and plan display. Shows "N/total done", plus "k doing" or "doing #i", instead of a last-touched step index that jumped backwards for plans worked in parallel. A re-opened or replaced plan shows the true count with "(plan changed)". Out-of-order completion no longer reads as all-done, which had silenced the stall, idle and burn signals.
  • archive. Completes the app-side archive for workspaces whose builder is closed but still listed. The branch-name fallback runs only when exactly one non-archived builder has that branch, and it is the one being archived. An empty target is never sent. Other builders on the same branch or worktree are snapshotted before and after, and any side effect is reported. doctor --repair matches by exact id only.

Added

  • Jev recommendation. A one-time recommendation notice when Jev is off (first session, then at most every 30 days; setting jev.recommendNotice), a doctor section, and a README "Enable Jev" block. The wording states only what is measured.

Security

  • git-guard. No verdict changes. 181 heredoc-scanning bypass forms are pinned as must-stay-blocked. The substitution block message now explains that heredoc bodies are scanned as shell, and suggests the Write tool. The GUIDE documents the limitation.

0.120.14 (2026-10-02)

Changed

  • BEHAVIOUR CHANGE: API keys come from sensitive plugin options. Jev and Anthropic keys are read from the plugin's sensitive options. Reading a key from the machine environment or a key file is now off by default and needs the home-only opt-in jev.allowLegacyKeyRead (Jev) or guards.allowAnthropicEnvKey (Anthropic env key), set in your home settings. An update migration enables the Jev opt-in once for existing installs that already have a key file, so those keep working. If you use an env-var key, or an Anthropic env key, set the matching opt-in or move the key into the plugin option. Key files must live under ~/.config or ~/.anti-hall and hold a single token.
  • Previous release's Jev endpoint change (loopback-only test override) was hardening; it is not a feature change.

Added

  • .anti-hall/ gitignore check. Doctor warns when .anti-hall/ is not git-ignored; --repair writes .git/info/exclude; a one-time reminder; setting guards.gitignoreHint.
  • Community files. Issue forms, PR template, code of conduct, expanded CONTRIBUTING, and an issue-triage workflow (inactive until a repo secret is set).

Security

  • Known secret shapes are redacted from text sent to the classifiers (best-effort, not a guarantee).

Docs

  • Privacy. PRIVACY.md, the README "Network and data" section, and manifest privacy, documentation and support URLs.
  • Directory and validator. No shipped text names input-rewrite or permission-decision fields; the Codex manifest no longer references the icon file; skills link out-of-plugin docs by URL; settings descriptions spell no shell pipeline; the Jev skills are reworded.
  • Demo GIF re-recorded. AGENTS.md compacted to about 27 KB, with a 30 KB soft-budget test.

Tests and internal

  • devswarm source-unit test helper (split stage 1); the migration test pins the old build by commit SHA.
  • CI runs once per release, on rc-v* tags, sharded; main and PRs run a reduced matrix.

0.120.13 (2026-10-02)

Security

  • Jev test-endpoint override is loopback-only. The Jev classifier's test-endpoint override (ANTIHALL_JEV_TEST_ENDPOINT) is now honoured only for loopback addresses. Previously, an environment value could redirect the request carrying the gateway API key to an arbitrary host. Users with a Jev key configured should update.

Fixed

  • spawn submodule repair. Reports a bare 40-hex sha; byte-identical set-aside files are no longer reported as conflicts (the duplicate copy is removed); differing files state which copy is in place; a leftover set-aside dir holding env or credential files produces a warning.

Docs

  • Release policy. RELEASING.md now states that rc tags are pruned after the release is published.

Tests

  • Defect history. Added a single-tag test for the defect-history backfill.

0.120.12 (2026-10-02)

Added

  • Plugin icon (plugins/anti-hall/.claude-plugin/icon.png), also referenced from the Codex manifest interface (composerIcon, logo).
  • Short manifest description plus keywords in both manifests.
  • CONTRIBUTING, SECURITY and issue templates.

Changed

  • BEHAVIOUR CHANGE: scan-throttle is now advisory-only. It injects additionalContext and never rewrites the command or tool input.
  • Codex anti-hall-activate and anti-hall-context-conserve skills run shipped scripts (codex/scripts/) in place of inline node -e one-liners.

Fixed

  • The NUL dedup-key separator in doctor-runtime.js is written as an escape instead of a literal control byte.

Docs

  • README overhaul: demo, what-it-blocks table, verify and uninstall, glossary.
  • llms.txt uses absolute links.
  • The Windows claim is corrected.
  • Settings descriptions for allowGcloudReads and allowBackgroundScratchScripts reworded (text only, no detection change); GUIDE matches.

0.120.11 (2026-10-02)

Fixed

  • silent-agent-nudge is no longer disarmed by an agent-list "running" row, and it survives compaction by adopting task_status attachments.
  • killed and cancelled agents are treated as terminal (notification and task_status shapes). An adopted agent with no timestamp and no output file is never nudged.
  • Orchestration rule I wording: no phantom heartbeat file, and "running" in an agent list is not progress.
  • limit-conserve honours the caller's home in inbox tick, so it no longer reads or writes the real home from in-process tests.
  • Hooks emit block JSON via a synchronous write.

Docs

  • Docs alignment: Jev mode count, hook table, monitors and --auto, inbox tick in llms.txt, the locked safety key, deadly-loop-multi marked Claude-only, the Windows snippet removed, private project names genericised, and a scratchpad volatility note in the orchestration skill.

Tests

  • LOCK LOST watcher test no longer races the watcher's own lock refresh (rare CI-only failure).

0.120.10 (2026-10-02)

Fixed

  • The wake-watch update-available line is emitted once per new version, persisted across re-armed watchers, and says /reload-plugins.
  • The task-guard IDLE NEGLECT text tells you to set blockedBy for tasks waiting on an in-flight task.
  • command-guard allows chains of individually light segments with an allowed bounded form. Shell-wrapped and heavy segments still block.
  • Read-only filters before a bounded last stage (... | grep -E x | head) are treated as bounded.
  • Codex rate limits that surface only in background job logs are detected and recorded, so the Critic seat falls back and the Codex nudge stays quiet.
  • The stale-version notice says /reload-plugins.
  • A pipeline filter stage with any output redirect is never treated as read-only (fixes a Linux-only allowance).
  • Dormant roster rows no longer show STRAYING, and supervision skips stall/burn nudges for them.

0.120.9 (2026-10-01)

Fixed

  • tasklist-guard demanded a progress header started: taken from its own clock, so it changed every Stop and could never be satisfied. It is now derived once from the transcript and persisted.
  • The running-agent count reads "unknown" instead of 0 when a launch is outside the transcript window, and never fires DISPATCH NOW on an unknown count.
  • parent-gate re-checks the live own-unread count before blocking on a cached one.
  • git-guard's command-substitution block suggests the Write tool for message text.

Added

  • mesh read also returns messages.
  • mesh history re-reads consumed broadcasts without moving cursors.
  • inbox messages --with-broadcasts.

Changed

  • After an update, /reload-plugins is the step to take. It was verified on 2026-10-01 to load new hooks and skills; restart only if a path still shows the old version.

0.120.8 (2026-10-01)

Fixed

  • devswarm.js spawn repairs a submodule worktree DevSwarm failed to add. Root cause: DevSwarm's background, un-awaited worktreeInclude copy pre-creates the submodule dir before git worktree add -b. The repair sets the files aside, adds the worktree (attaching to an existing branch if needed) and restores the files without overwriting; conflicts are reported. The output includes a pinned-commit hint.
  • send/mesh read/heartbeat from a non-worktree cwd (e.g. the session scratchpad) fall back to CLAUDE_PROJECT_DIR for a Primary.
  • model-routing-guard no longer blocks a flagship on reasoning-heavy briefs (analyze/synthesize/read a PDF page by page), and never suggests haiku for them.
  • command-guard allows a background direct-exec of an executable script in the session's own scratchpad, and names the allowed shapes. git pull and mutating fetches are labelled as state-changing remote operations.

0.120.7 (2026-10-01)

Changed

  • /anti-hall:update runs update.js (and the DevSwarm supervisor/ingest installer steps) IN the main session instead of delegating to a Sonnet subagent. Reason: it runs migrations, and the main session must judge the result, per the owner's rule that deploys and migrations are never handed to a delegated model by default.
  • command-guard exempts the anchored update.js and installer invocations.
  • model-routing-guard blocks a subagent spawn that would run the update (setting guards.updateInSession, default on).

0.120.6 (2026-10-01)

Fixed

  • devswarm.js archive reported appArchive.ok:true from hivecontrol's exit code alone, without checking the DevSwarm app, so a workspace could stay live in the app while anti-hall treated it as archived. The app archive is now VERIFIED against the app-DB snapshot, or against hivecontrol's archived:true response when the DB isn't readable. Note in passing: workspace list all keeps archived rows, so it isn't an archived signal. An unverified archive returns partial:true with manualCommand.
  • archive <branch|meshId|uuid> archives the app side alone when the anti-hall descriptor is already archived.
  • roster rows carry appArchived; appStillLive and a one-time Primary note surface workspaces the app still shows after an anti-hall archive.
  • doctor detects that mismatch, and only an explicit doctor --repair runs the verified app archive.
  • Verified live against DevSwarm on throwaway workspaces.

0.120.5 (2026-10-01)

Fixed

  • devswarm.js roster no longer lists an archived workspace twice. A native branch-name row whose worktree maps to an archived descriptor folds into the archived row (hints:['archived']). The output adds liveCount/archivedCount.
  • command-guard names state-changing remote operations (gh pr create/merge, git push) as such in its block reason, instead of "heavy". A ;-joined cd that would qualify with && gets a "use cd <dir> &&" hint. Behaviour is unchanged.

0.120.4 (2026-10-01)

Security

  • git-guard now blocks launcher-dir writes via glob redirect targets and same-command variable targets (pre-existing).

Fixed

  • git-guard reads a redirect target as its first word (fewer false blocks after an odd quote).
  • command-guard allows git push ... > <own scratchpad file> (symlinks refused; newline-safe).
  • Test fixtures use neutral ids.

0.120.3 (2026-10-01)

Fixed

  • The update's reconcile step: a DevSwarm native message-count timeout with nothing imported or lost is now "skipped (native unavailable: timeout)", reported in one aggregated summary line, with one retry after a backoff, instead of "failed" per worktree.
  • devswarm.js archive/unarchive/archive-request accept mesh ids, resolved like send --to. An ambiguous id lists the candidate UUIDs and archives nothing.
  • parent-gate: unread broadcast rows to a done (archive_ready) child no longer count as neglect, and the escalation suggests archiving when every flagged child is done or idle.
  • New regression tests for edit-guard handover writes and devswarm.js send loops.

0.120.2 (2026-10-01)

Fixed

  • Compact advice/declaration guards no longer fire on a MENTION of the phrase: straight double-quoted text and backticked /compact mid-sentence are treated as quotes. Instruction forms ("Run /compact …", a standalone command line, a bare SAFE TO COMPACT line) still count.

0.120.1 (2026-10-01)

Fixed

  • Compact-declaration guard deadlock: the auto-handover nag's mandated footer (GOOD POINT TO /compact NOW) and its pasteable /compact focus: … line were read as a SAFE TO COMPACT declaration, blocking every tool call in later turns; refreshing the handover (which the block itself demanded) was also blocked. Now only an explicit SAFE declaration counts, .anti-hall/handovers/** edits are always allowed, and leading hook reminders no longer stop a new user prompt from resetting the turn.

0.120.0 (2026-10-01)

Fixes from the second cross-project peer bug sweep (Primaries plus their workspaces), three rounds of adversarial security review of the command/git/edit guards, and DevSwarm identity, inbox and mesh fixes.

Fixed

  • DevSwarm store id-collision guard no longer rejects a submodule path of the registered workspace. The persisted toplevel of a child running in a submodule (e.g. <ws>/sub) was compared against the workspace-keyed id as if it were a different worktree. The guard now resolves both directions to the outermost superproject root and treats them as the same worktree.
  • Child summaries and sends refresh the plan progress clock; identical repeats do not. A child that was working (heartbeat --summary, mesh sends) but had not changed plan step was judged stalled. noteActivity now records activity with new text; text already among the recent signatures does not refresh, so a looping child still stalls.
  • child-gate: one-time inbound-only warning when the native queue is unreachable and the outbound half is already satisfied. The gate re-demanded a heartbeat right after one was sent and prescribed inbox pull, which hits the same dead native channel. It now warns once (own kind, cap 1) with an exit that does not need the native channel.
  • Queued-prompt bursts no longer repeat COMMS OVERRIDE and the primary TASK-LIST block; task-guard accepts workspace branch/title owners. N cron ticks delivered together showed the static block N times; it now collapses via emit-dedupe (keepalive = guards.injectionRepeatEvery). task-guard also accepts an owner spelled as the branch name or title of a live, non-archived workspace.
  • No NOT DRAINING claim on graced fresh sends to stuck children; own-inbox escalation is count-free. The Primary's own-inbox-only escalation is held while usage limits are at conservation, and the forced-ack state is cleared on a fully clean pass so new mail starts a fresh budget instead of escalating immediately.
  • mesh read/inbox row shape is normalized; mesh read --last/--since added (peek-only); project resolves from DEVSWARM_BUILDER_ID outside a worktree. The filters require --peek so they cannot consume unread rows.
  • Guard housekeeping. tasklist-guard ledger wording, heredoc/comms hook housekeeping, edit-guard matches .anti-hall/** allowlist names against the project root when cwd is a subdirectory, submodule or symlinked spelling (additive, never narrower), and a line starting with RETRACT/RETRACTED/RETRACTING is a free-form retraction. Docs: one git writer per worktree.
  • DevSwarm roster step counter is monotonic; a done-reported child no longer shows STRAYING stall/burn. The displayed plan step could move backwards, and a child that had reported done was flagged as stalled/burning while it was simply awaiting its parent.
  • The devswarm-fold-mesh journal test no longer flakes on same-millisecond registry timestamps. The 2+ LIVE fold guard now seeds the second live entry on a later millisecond.
  • command-guard judges shell-loop bodies per command and tracks cd. A for ... do node .../devswarm.js send ...; done loop was blocked while the single send passed, because the loop keyword stayed glued to the body and defeated start-anchored exceptions; the keyword is now stripped and each body segment is classified on its own. Bounded checks follow cd; path-qualified interpreters that exist as regular files are accepted; [npx] vitest run / jest on 1-2 explicit test files and a push of a HEAD sha are allowed (full suites, watch, --coverage stay heavy). The devswarm.js exemption now requires the file name to end at devswarm.js.
  • git-guard: odd-quote heredoc prose no longer blocks the executed launcher. Launcher-dir redirect targets are attributed per line, so an apostrophe in a heredoc body does not make the real command look like a launcher-dir write. ditto is now a copy verb for launcher-dir writes.
  • codex-quota parses "try again at " and codex-nudge stays quiet while limited. A Codex usage-limit message giving a reset as an ordinal date ("try again at Oct 3rd, 2026 ...") was not parsed, so no quota record was kept and the nudge kept firing. The date is now parsed, an unparseable limit message falls back to a 6 h cooldown, and codex-nudge is skipped while a quota record is live.
  • DevSwarm Primary registration and workspace listing no longer key a submodule cwd as its own worktree. register-primary, workspaces list, archive-prefix and the ingest installer workdir used the submodule toplevel, minting phantom Primary rows. They now use the superproject-folding resolver. doctor detects existing phantom submodule Primary rows; only doctor --repair archives them (nothing is deleted).
  • command-guard read-only-verify and push carve-outs tightened (review round 2). cd is honoured only as an &&-chained step (not after ||, in a pipe, or after ;) and the cd target and script path are realpath'd, so symlinks cannot steer an anti-hall script past the refusal; the vitest/jest carve-out requires each operand to be an existing regular *.test.*/*.spec.* file (a bare .test.ts is a filter pattern); a short-sha push source must resolve through git to exactly HEAD (a tag named like the sha no longer matches); a leading ! negation (while ! npm test; ...) no longer hides the heavy verb.
  • work-detect: a << after an unquoted # is a comment, not a heredoc. The lines below it were dropped as heredoc body and hid real work from the housekeeping-only verdict.
  • git-guard: the quote-blind launcher backstop is skipped unless the text can contain an anti-hall path; a backslash-newline collapses only after an odd backslash run. The backstop made the 160 KB cd-chain PERF test about 10x slower; an O(n) precheck restores it. A launcher-dir-looking path with an escaped backslash before the newline no longer false-blocks.
  • work-detect: an escaped \<< is not a heredoc, and a partly quoted delimiter (<<E"OF") ends at its unquoted word. Both dropped real commands from the housekeeping-only verdict.
  • Defect channel: two concurrent defect report calls for the same fingerprint could lose one line. The creating process opened the file without O_APPEND and overwrote the other's append. The file is now created with 'ax' (O_CREAT|O_EXCL|O_APPEND).

Security

  • git-guard: a leading ! negation was not treated as a wrapper word, so if ! <push with a force flag>; then ..., while ! <force push>; ... and if ! <push of an empty-source refspec>; ... slipped past the force-push and ref-deletion checks. ! is now resolved like if/then.
  • git-guard: git accepts unique prefixes of long options, so abbreviated push deletion options (--del, --dele, --pru) and force/mirror options (--force-w, --for, --mir) were not recognised. git push long options are now resolved by unique-prefix match; an ambiguous prefix counts as every candidate (fail-closed, --f blocks as force).
  • git-guard: a launcher-dir write (cp/tee/mv/sed -i/install/ln/cd+write) placed after a heredoc body with an unbalanced apostrophe was not detected; a quote-blind per-line launcher backstop now blocks it.
  • git-guard: a quoted redirect target spanning a newline (>"$HOME/<NL>/../<launcher dir>/x", also >>, 1>, &>, x>) escaped the launcher-dir write block because the target was cut at the newline. Redirect words are now read from the raw text (a quoted span runs to its matching quote) in both the quote-aware pass and the backstop.
  • git-guard: a backslash-newline continuation in the middle of a launcher-dir path (> ~/<dir>/\<NL>bin/y) bypassed the launcher-dir write block; continuations are now removed (as bash does) before launcher analysis.
  • git-guard: git push deleting a remote branch/tag (--delete, -d, origin :<ref>, --prune) is now blocked with a reason naming the skip.json override; deleting published refs needs explicit owner confirmation.

0.119.0 (2026-09-30)

Fixes from a cross-project peer bug sweep plus Jev measurement fixes: the wake watcher, STALE DATA banner, reload/restart wording, guard hints and home isolation, and jev report gains rare-trigger counters and --exclude-project.

Fixed

  • Wake watcher no longer wakes non-DevSwarm sessions. monitors.json ("when": "always") starts the watcher at every session start, and every stdout line is a transcript event, so a plain terminal session got REFUSED TO ARM: not-a-devswarm-session and spent a model turn on it. The harness-started entry now passes --auto; with it the expected refusals (not-a-devswarm-session, disabled-by-settings) go to stderr only. A model-armed watcher still prints its one refusal line.
  • Wake watcher and companions no longer emit the node:sqlite ExperimentalWarning, and the watcher can hand off on every later release. On Node 22.5+ the first require('node:sqlite') wrote an ExperimentalWarning to stderr, and under the harness Monitor every stderr line is a wake event, so each one cost the model a turn for noise. All companion and devswarm.js loads now go through companion/lib/sqlite-quiet.js, which swallows only that one warning during the load. Separately, the handoff loop guard allowed one handoff per process chain (stamp 1), so every second release in a long-lived chain ended in STALE BUILD plus a manual re-arm. The stamp is now the target version, and a handoff is allowed whenever the newest version is strictly newer than the stamp; same-or-older targets and legacy or non-semver stamps are still refused.
  • The ingest STALE DATA banner is emitted once per stale episode and says what is at risk. The banner repeated on every turn and only told the user to run /anti-hall:doctor. It now goes through staleBannerOnce (a recovery clears the record; re-emits only when staleness crosses a coarser tier, with a 50-turn keepalive). The text names what may be stale (roster and app-state freshness) and notes that mesh sends are written directly to the store and are not affected. When launchctl print-disabled lists the ingest job as disabled, the banner says self-heal is skipped and gives the one-line launchctl enable command; anti-hall never runs it.
  • The Primary no longer gets told a dead, older session "was active until" a recent time. Session recency compared another session's start and its handover mtime against the current session's start, so a long-running live Primary was warned about an older dead session, and unconditionally when its own transcript was not matched. Both checks (primary-seat stale-resume and anchorSessionDrift) now compare transcript last-activity by stat only, and stay silent when the current session's activity is unknown. handover-resume also adds a WRITER KEPT RUNNING line when the handover's writer transcript was written more than 5 minutes after the handover.
  • tasklist-guard no longer counts .anti-hall/handovers writes as work. Writes to .anti-hall/handovers/ (which the handover flow itself directs) and cwd-relative .anti-hall/progress|history|handovers/ paths counted toward the work threshold. The exemption pattern now covers handovers and accepts a path with no leading slash.
  • PARENT INBOX hints and update wording name the stable launcher and the right reload step. Hints in devswarm-parent-inbox embedded the version-pinned plugin-cache path, which goes stale after /reload-plugins or once an update prunes that version directory. A shared preferStableLauncher (in hooks/lib/stable-launcher.js) now returns the ~/.anti-hall/bin launcher when it already exists and devswarm.stableLauncher is on, and is used by the parent-inbox and child-turn hooks and the primary-seat conflict text in devswarm.js. The update skill, version-alert, and the registry-lag note now say /reload-plugins only when the plugin cache alone changed, and a full restart of Claude Code when the harness registry names a newer build than the running session. The update skill also respawns a Sonnet delegation with model:"opus" when Sonnet is unavailable or usage-exhausted (never Haiku).
  • edit-guard's DevSwarm block text points at the exempt notes locations. The block message now says session notes and reports can go in .anti-hall/history/** or the scratchpad (exempt), and that repo docs need a subagent or a trusted .anti-hall/edit-allow.json.
  • command-guard blocked the background scratch-script shape its own block text recommends, and the timeout form of anti-hall CLI calls. A run_in_background scratch script chained with a read sink (script > out; wc -l out) was refused because the allowance covered a single unbroken segment only. The check now accepts segments joined by ;, && or |, each either a scratch-script segment or a bounded sink (wc, head, tail, grep -c / grep -m N), with at least one script segment required. Separately, timeout N node <dir>/devswarm.js roster | head was blocked while the same line without timeout passed, because the anchored anti-hall CLI exemptions did not see through the wrapper; isHeavySegment now also matches the command with a leading timeout [-flags] <duration> removed. The heavy-pattern block text now names the working shape: use the literal absolute scratchpad path (not $VAR) and chain only wc/head/tail/grep -c/grep -m N.
  • askDetached (Jev) defaults to the 3000ms ceiling instead of the 1500ms interactive budget. A detached Jev call no longer inherits the shorter interactive timeout.
  • Spawned hook children carrying the test-isolation marker can no longer resolve the real home. resolveHome in companion/lib/test-home-guard.js now refuses a child whose HOME is missing or real instead of falling back to the developer's ~/.anti-hall; covered by new cases in no-real-home-entrypoints and git-guard tests. The commit also bounds Jev row volume per git-guard invocation.
  • Rare Jev triggers are counted and shown in jev report. speculationFramed and devswarmOnBrief / devswarmExtraSanctioned occurrences are now recorded as counters and listed in the report.
  • jev-report --exclude-project <name> drops the raw rows of a leaked or synthetic project from the report.
  • Test hygiene. New os.homedir() call sites in handover-resume.js and emit-dedupe.js (forget) now go through resolveHome (the homedir-call-site-ratchet baseline stays 315), and the handover-resume test's git spawn isolates HOME/USERPROFILE.

Notes

  • git-guard still treats a launcher-dir write command quoted inside a heredoc body as a real write (fail-closed); a narrower exemption was tried and reverted after review found real bypasses — write such text with the Write tool.
  • workflow-drift-guard false positives on conditional test.skip(cond, …) come from oh-my-claudecode, not anti-hall.
  • the 2026-09-28 10:57–11:01Z jev-assist rows from project v117-int are synthetic; exclude them with --exclude-project v117-int.

0.118.0 (2026-09-30)

DevSwarm wake-watch/limit-skip and nudge/roster fixes, a codex-nudge scratchpad-exclusion fix, and further git-guard hardening — five field-reported fixes, each root-caused against a reproducing test before the fix.

Added

  • devswarm.wakeWatchIdleSkip (default true, #39) — a DevSwarm Primary with 0 LIVE (non-archived) child workspaces gains nothing from an armed wake-watch Monitor (nothing will ever message it), yet re-arming it every cron tick still costs a turn. When on and the caller is a Primary with 0 live children, inbox tick reports watcherArmed idle-skip (never the boolean false), so the cron prompt's "re-arm if watcherArmed false" rule stays inert; if devswarm-wake-watch.js is started anyway it prints [wake-watch] idle-skip: no live child workspaces — not arming (cron fallback covers) and exits 0 without ever acquiring the watch lock. A live child (companion/lib/ devswarm-live-children.js hasLiveChild) leaves both paths unchanged; an archived-only child still counts as 0 live. The cron fallback itself is never removed or altered. Metric: ~/.anti-hall/devswarm/rearm-cues.jsonl (trigger idle-skip), surfaced by doctor as "wake-watch idle-skips: N". Codex mirror: inbox tick's watcherArmed documentation updated in plugins/anti-hall/codex/skills/ anti-hall-devswarm/SKILL.md (Codex has no Monitor tool, so a Codex workspace never produces idle-skip itself).

Fixed

  • nudge <id> escalated immediately ("poke-exhausted") on the very first call after a child had been idle over an hour. computeLiveness's resolved-to-alive paths (heartbeat-fresh, nudged→alive on advance, and the general recompute) carried a stale nudgeAttempts/nudgedAt forward from an EARLIER episode instead of resetting them — only the heartbeat-fresh branch reset the poke budget. A LATER, unrelated stale spell then inherited an already-exhausted budget and escalated on its first pokeOrEscalate call. companion/lib/liveness.js's two remaining resolved-to-alive returns now reset nudgeAttempts/nudgedAt on that episode boundary too. The escalation notice into the parent's store now also tells the owner to click/continue the workspace in the DevSwarm app (still allows reassign/archive), and lastNudgeError is preserved unchanged.
  • nudge <child meshId> failed {"ok":false,"error":"no descriptor for workspace ..."} while send --to <that same meshId> worked. cmdNudge looked up only readDescriptorFile(home, id) (an exact builder-id match), with no meshId resolution at all. It now falls back to the SAME resolution send uses (resolveSendTarget: meshId match, then exact registry-id, then a one-hop retired-redirect) before failing closed.
  • roster from a cwd outside any project (e.g. a session's scratchpad) returned bare {"ok":false,"reason":"no-project"} with no guidance. repoKey is always git-derived (there is no repoId→repoKey registry to fall back through — DEVSWARM_REPO_ID is a caller-declared label, not a store key, and trusting it here could silently read the wrong project's roster), so roster now returns an actionable error string (same wording as send's own no-project error) instead of inventing a new resolution path.
  • task-guard's IDLE NEGLECT block never mentioned delegating a task to a live DevSwarm child workspace. When at least one live (not-archived) DevSwarm child workspace exists, the IDLE NEGLECT reason now adds one line: delegate the task by setting its owner to the workspace id (TaskUpdate owner), which already counts as attended via the existing devswarmChildAttended check — detection is unchanged, this only documents the option.
  • The "CHILD NOT DRAINING" nudge repeated every single turn for an idle child. devswarm-parent-inbox.js's normal-tier unread segment (buildUnreadSegment) passed no keepaliveTurns to its on-change dedupe call — unlike the urgent tier just above it — so an UNCHANGED segment re-fired on every delivered turn once the prior copy was consumed. It now passes keepaliveTurns: 2 (same pattern the urgent tier already uses), capping repeats to 2 before going quiet until the content actually changes (new mail, a drained inbox, a status flip) or the keepalive window resurfaces it. Suppressions are counted in a new not-draining-suppressed.jsonl, reported by doctor --check alongside the existing cron-found-mail.jsonl line.
  • codex-nudge.js counted scratchpad-only and out-of-worktree edits as "substantial code edits", nudging for a Codex review over disposable repro scripts. Seen twice for *.py files written to the session's own scratchpad (/private/tmp/claude-<uid>/<encoded-cwd>/ <session>/scratchpad/), and once for a scratchpad .js repro script. The edit counter now excludes any edited path strictly inside the session's own scratchpad (reusing lib/scratchpad.js's ownScratchpadDirs/isInsideDir, derived from transcript_path's encoded-cwd segment — the same mechanism edit-guard.js/command-guard.js already use) and any path outside the session's git worktree, when payload.cwd resolves to one (via companion/lib/identity.js's resolveContext, the single canonical worktree resolver — no new filesystem walk). Both checks fail open when the needed payload fields are absent.
  • git-guard: more force-push/wrapper bypasses closed (0.117.2 follow-up, A1-5/A1-6). None needed a quote desync. git push --mirror force-updates and deletes every remote ref and is now treated as a force push. time -p/command -p used to resolve the effective command verb to -p itself, skipping the wrapped command entirely - both now skip their own -p flag. coproc git push --force ... is now recognized as a wrapper, like exec. An xargs-run git push is now treated as a force push whether or not the visible argv already carries --force: xargs appends words it reads from stdin as trailing arguments, so a hidden -f (e.g. echo -f | xargs git push origin main) cannot be ruled out statically. env -S 'STRING'/--split-string (which word-splits STRING and runs it as a new command) is now unwrapped and re-scanned, and a bash -c $'...' ANSI-C-quoted payload is now decoded before being re-parsed (it used to carry a literal leading $ into the recursed command, so the verb never resolved to git). A literal string piped into a bare shell (echo "git push --force origin main" | bash) or fed via a here-string (bash <<< "git push --force origin main") is now recognized as feeding that shell's stdin as its script, the same as bash -c "...".
  • git-guard: quadratic scan on many git commit -F - segments (0.117.2 follow-up, A1-6). extractQuotedLiterals(cmd) ran once per -F -/--file=- segment over the FULL, unchanged command string; 20000 such segments took about 13s, over the hook's 10s timeout (a fail-open miss of a trailing force push). It is now memoized per distinct command string, computed once instead of once per segment; the same input now decides in well under 100ms.
  • emit-dedupe: a fixed-turn-count keepalive block (parent-inbox-urgent, keepaliveTurns: 10; NOT DRAINING, keepaliveTurns: 2; limit-conserve; verify-first; task-tracker) could go silent forever on Codex. The fallback window used when the transcript cannot confirm consumption (always true on Codex, which has no hook-attachment transcript) doubled as BOTH the pending-suppression window AND the "new turn" detector once dedupeWindowMin (default 20 min) replaced the old fixed 15s constant for the former. A dedupeWindowMin of 20 minutes meant a turn only counted as "new" every 20 minutes, so a keepaliveTurns: 10 block needed over 3 hours of unchanged content before it re-surfaced — effectively never, on a normal cadence. The "new turn" check is now back on the original fixed 15s threshold regardless of dedupeWindowMin, which continues to govern only the initial pending-suppression window.
  • git-guard: an xargs-run git command merely MENTIONED inside a quoted commit message or redirect target was wrongly treated as a live command (A1-2, 0.117.2 follow-up). xargsGitVerdict ran from the quote-BLIND backstop split (cuts at every |/;, inside or outside a quoted string), so e.g. git commit -m 'docs: explain why ls | xargs git push is blocked' or echo '... | xargs git push origin' > notes.txt false-blocked. It now runs only from the quote-aware segment scan and the quote-aware per-line backstop recovery pass; a real echo -f | xargs git push origin main still blocks, and an xargs-run git command inside a heredoc BODY line still blocks too (heredoc bodies are deliberately scanned as shell, unrelated to this fix).
  • DevSwarm inbox tick's wake-watch idle-skip scoped "does this Primary have a live child?" from the CALLING process's cwd instead of the tick id's own registered worktree (A1-3). A tick invoked from a different clone/checkout of the same project than the one id is registered under resolves to a DIFFERENT repoKey (repo-keyed off --git-common-dir, which is per-checkout), so hasLiveChild's project scoping silently missed every real sibling child and idle-skip fired even with a live child present. The scope now comes from id's own registered descriptor's worktreePath; when that cannot be resolved, the idle-skip check is skipped entirely (fail open, never idle-skip on an unresolvable scope) rather than falling back to the calling cwd.
  • git-guard: the -F -/--file=- stdin-commit-message scan was still O(N^2) on many segments with quoted text (A1-4, 0.117.2 follow-up). A1-6's memoization cached only extractQuotedLiterals's result ARRAY; the candidate-text .join() and the self-credit regex test against that joined text still ran unmemoized on every -F - segment, so a command with quoted text on each segment (e.g. git commit -F - <<<'m'; repeated 20000 times) still paid an O(cmd.length) join + regex scan per segment. Both are now memoized per distinct command string alongside the literal array.

Changed

  • command-guard.js's "Heavy command detected" delegation text now surfaces the scratchpad-script path first. The "write it to the scratchpad and run it with run_in_background" option used to be the very last clause of the block message, easy to miss before reading it as a plain "spawn a subagent" instruction. A one-line version now appears right after the first sentence in both the DevSwarm-primary and non-DevSwarm variants. Text only — no change to which commands are allowed or blocked inline.

0.117.2 (2026-09-30)

Security fix for git-guard.

Fixed

  • git-guard could miss a force push that came after a heredoc or a stray quote (0.116.1 and 0.117.1). git-guard splits a command into its separate commands while tracking quotes. Some text threw that tracking off: an apostrophe in a heredoc body (It's done), a quote in a # comment, or a heredoc inside a quoted $( … ). When that happened, the rest of the command looked like one long quoted string, so a later force push was allowed. git-guard now also runs a quote-blind check. It cuts the whole command at every newline, ;, &, |, (, ), $( and backtick, ignoring quotes, heredocs and comments, and applies the git rules to each piece: force push in every form, push arguments built by command substitution, command-valued git config, and AI self-credit trailers on commit-creating commands. These pieces are also checked inside eval and bash -c payloads. The command is blocked if either check blocks it. Rules that are not about git, such as launcher-directory writes, still use only the quote-aware check. Known and deliberate: quoted text that has a separator right before a literal force push, such as git commit -m "don't; git push --force", is now blocked. A plain mention like git commit -m "never git push --force" is still allowed.
  • Slow arithmetic scan. Long inputs with many << inside $(( … )) took about 9 s to check, close to the hook's 10 s timeout. The scan now continues from where it stopped instead of starting again at the beginning, so a 96 KB input is checked in about 30 ms.
  • Follow-up hardening on the quote-blind check above. The quote-blind check cuts a command into pieces at every ; & | ( ) and backtick, so it could also cut INSIDE a git command's own quoted argument - git -C "$(pwd)" push --force-with-lease origin main, or a -m/--trailer message containing one of those characters (a conventional-commit subject like feat(x): y opens with (). That split the piece holding the git verb from the piece holding --force/the trailer, so the command was allowed. The check now also re-reads each physical line on its own with the normal quote-aware splitter and applies the git rules to it - a line is almost always quote-balanced by itself even when the whole command is not. Also, a force push inside the condition of an if/while/until/elif (e.g. if git push -f origin main; then …; fi) was never checked at all; those are now recognized the same way then/do/else already were. And a regex alternation grouped in parentheses (e.g. grep -E "(a|git push -f)" x) was wrongly blocked, because the quote-blind check's own (/) cuts split the still-open quoted pattern across more than one piece before its closing quote; the check now tracks the quote state across pieces on the same line instead of only the one piece right after a |.

Fixed

  • UserPromptSubmit injection burst deduplicated by a configurable window (field report) — a DevSwarm Primary idle through a multi-day usage-limit outage had every queued cron/wake prompt re-delivered together on resume, and each one re-triggered anti-hall's full LIMIT CONSERVATION, TASK-LIST, DEVSWARM COMMS OVERRIDE and DEVSWARM WORKSPACES injection blocks — one resumed turn carried ~60 identical copies (~150 KB). hooks/limit-conserve-inject.js, hooks/task-tracker.js and hooks/devswarm- child-turn.js already deduped their static blocks via hooks/lib/emit-dedupe.js (as does hooks/devswarm-parent-inbox.js's WORKSPACES table); the fallback window used when a burst's consumption cannot be read from the transcript is now driven by the new context.dedupeWindowMin setting (default 20 minutes, 0 = off/disables emit-dedupe entirely) instead of a fixed 15s constant. Suppression counts (per session and overall, with a last-suppressed age) now surface in /anti-hall:doctor via hooks/lib/emit-dedupe.js's new summary().
  • inbox tick reports watcherArmed limit-skip while LIMIT CONSERVATION is active — the mailbox-tick cron prompt tells the agent to re-arm a lapsed Monitor wake-watch when watcherArmed reads false, which directly contradicted LIMIT CONSERVATION's own "defer non-urgent work" instruction while both were active at once. inbox tick (scripts/devswarm.js cmdInboxTick) now checks hooks/limit-conserve.js isConserving() after the existing idle-skip check and, while conservation is active, reports the string watcherArmed limit-skip instead of false — sibling behavior to devswarm.wakeWatchIdleSkip's idle-skip, no separate setting. The cron prompt text (hooks/lib/devswarm-wake.js drainCmd's rearmClause) now says both idle-skip and limit-skip mean "do not re-arm"; the cron job itself is never altered. Metric: same ~/.anti-hall/devswarm/rearm-cues.jsonl ledger, trigger limit-skip, surfaced by doctor as "wake-watch limit-skips: N".

0.117.1 (2026-09-28)

Four field-reported fixes from DevSwarm child workspaces, all root-caused against a reproducing test before the fix.

Fixed

  • edit-guard's own-scratchpad exemption broke after an in-session cd. The session scratchpad directory is named after the session's ORIGINAL project cwd, but the exemption re-derived it from the LIVE cwd on every call — after a cd (e.g. into a DevSwarm child worktree), a legitimate Write/Edit into the session's own scratchpad was wrongly blocked with the edit-delegation message. lib/scratchpad.js's ownScratchpadDirs() now prefers the encoding read off transcript_path's parent directory (the harness's own record of the session's original cwd), falling back to cwd when the transcript path is absent, with no change to which session a scratchpad belongs to (still scoped strictly to session_id).
  • A child's first-ever heartbeat --summary (before any register) was silently dropped — then the first fix for that opened an impersonation hole. cmdHeartbeat's mesh-broadcast ownership check required a pre-existing registry row for the target id before considering any ownership leg; a workspace's very first interaction is often a direct heartbeat <id> --summary ... with no prior register, so the broadcast was refused caller-not-registered and never reached the shared store — devswarm-child-gate.js's Stop gate then kept re-prescribing the same heartbeat forever. The first cut of the fix gave broadcastFamilyOwns a first-claim leg that granted ownership of ANY never-registered id to ANY cwd-resolved caller — which let an unrelated caller impersonate a sibling's FUTURE id and lock the real owner out once claimed (review finding B-a0-impersonation). First-claim now additionally requires GROUND TRUTH: the DevSwarm app's own database (builderForWorktree, the same ground-truth reader the register-path app-archive guard already trusts) must name a builder for the caller's own resolved worktree whose id is exactly the id being claimed — a bare DEVSWARM_BUILDER_ID declaration is never sufficient proof by itself. No app-DB match (DB off/unreadable, no builder row for the worktree, or a different builder id) fails closed, same as any other refusal. Also: a dropped --summary now carries a plain top-level note: 'summary NOT recorded: <reason>' instead of only a nested dropped/dropReason pair that was easy to miss.
  • First-claim heartbeat ownership: follow-up hardening round (the app DB "ground truth" was itself env-forgeable). The fix above's own claim that the app-DB ground-truth reader "never consults env" was wrong: its DB file path honors ANTIHALL_DEVSWARM_APP_DB with no gating, and a real CLI invocation's ctx.env defaults to process.env — the same process env the untrusted caller controls — so a caller could point the "ground truth" DB at a throwaway sqlite file of its own and claim any never-registered id through the very leg meant to stop that. The override is now honored only when an in-process caller (tests, another in-process embedder) supplied its own env explicitly; a real CLI invocation always resolves the fixed per-OS app DB path. This leg also now matches an ACTIVE app-DB builder row only — an archived/hidden-only row at the worktree no longer grants first-claim.
  • A workspace could register itself with the exact id system. Mesh code reads a row's sender/from field to attribute a message; a workspace registered as system would make its own outbound rows indistinguishable from a genuine system-authored one. 'system' is now a reserved EXACT id (never a substring match — an id that merely contains "system", e.g. ecosystem-service, still registers normally) on a fresh register, refused the same way the existing reserved-token ids are.
  • devswarm.js nudge <id> gave no reason for "poke-exhausted." A failing poke (e.g. the session's channel is unreachable) was silently swallowed; pokeOrEscalate now persists the failure as lastNudgeError (carried across sweep ticks) and, once attempts are exhausted, the escalate result carries lastNudgeError plus a human line: "session appears dead/unreachable — needs a manual continue in the DevSwarm app."
  • Escalation notices into the parent's inbox showed a blank sender. The synthetic escalation row had no sender at all (rendered as an empty "from:" line) — root cause was two-fold: the row literal was missing sender: 'system', and the delivery path itself (via: 'message') only persists {workspaceId, ts, hash, body} and silently drops every other field including sender; delivery now goes through the mesh-aware via: 'row' insert (already used elsewhere) so sender is actually written.

0.117.0 (2026-09-28)

Measured. Every new judgment/automation surface in this release ships its own effectiveness report: DevSwarm dispatch demand (scripts/dispatch-report.js), Meeseeks supervision — plans, warnings, corrections, respawns, token burn (devswarm.js supervision-report), and Jev's own agreement-vs-deterministic-signal and follow/override counts (jev-report.js, folded into supervision-report per integration).

Added

  • DevSwarm Meeseeks supervision: step plans (P1). spawn -p turns a numbered list in the brief into the child's step plan (~/.anti-hall/devswarm/plans/<key>.json; a Scope: line becomes its file scope). A brief without a list is never refused. New verbs plan set|show; heartbeat --step N --status doing|done|blocked. The workspace table and roster show step 3/7 · 42m · progress 18m ago. Settings devswarm.planTracking (on) and devswarm.planRequired (off). Rows without a plan render exactly as before.
  • Straying detection + correction (P2). For planned children only, the supervisor sweep raises advisory signals: stall (busy, no step progress for devswarm.stepStallMin, 30 min), off-scope (ready-check --allow over scope ∪ extras), idle (stale liveness verdict) and burn (see below). Each episode warns once, capped at devswarm.strayWarnMax (2) per signal per step. The Primary gets one capped DEVSWARM STRAYING line per warning per session on Stop (stable alert kind devswarm-straying, never a block), STRAYING: … and +N extras on the table cell, and plan.straying on the roster. devswarm.js correct <id> [--dry-run] sends "step N '': . Return to step N or reply BLOCKED " and records warned_at only after the send succeeds. Nothing is automatic; respawn stays manual.
  • User-sanctioned extras. devswarm.js scope add <id> --glob G --note TEXT: the child tags extra work a user asked for; tagged globs stop counting as off-scope and the Primary sees the note. The child-turn hook tells a planned child to do this.
  • Token burn. The sweep reads each planned child's own session transcript incrementally (byte offset per workspace, complete lines only, one count per message id) and computes weighted tokens: input + output + cache writes + devswarm.burnCacheReadPct% (10) of cache reads. A burn warning fires past devswarm.burnTokensWarn (2M) since the step last moved ("used 2.1M tokens since step 3 last moved", also in the correction text). The table and roster show · 1.8M tok; the table's change-dedupe ignores the rising figure. Codex children have no Claude transcript and get no burn figure.
  • Five Jev supervision integrations, as recommendations. jevIntegrations.devswarmOnBrief, devswarmExtraSanctioned, devswarmWaitKind, devswarmLoop and devswarmStepMap default on: asked detached from the sweep only when a deterministic precondition fires (one ask per input per devswarm.supervisorBlockerLabelReaskSec), inputs capped and secret-scrubbed, answer read from the cache next sweep. The verdict and confidence ride on the warning ("Jev: waiting on CI/owner/peer, not stuck 0.92"); Jev never suppresses a warning, blocks or kills, and the Primary decides. shadow logs only. jev-setup.js mode knows all five.
  • Supervision effectiveness metrics. ~/.anti-hall/logs/devswarm-supervision.ndjson (1 MB × 5, daily rollups in devswarm-supervision-daily/) records plans, step changes, warnings, corrections, corrections followed by step progress, extras, done (time-to-done, steps done vs planned, tokens per completed step), token periods and Jev answers. devswarm.js supervision-report [--days N] [--json] reports warnings by signal and repeats, the correction follow rate, extras, tokens per workspace and per step, burn warnings and their corrected rate, and per Jev integration its agreement with the deterministic signal and the Primary's follow/override counts. doctor prints a one-line 7-day summary.
  • Primary-run respawn that keeps progress (P3). devswarm.js respawn <id> [--dry-run] replaces a straying child. It refuses unless the caller holds the Primary seat, the plan has warned_at (a correct was sent) and devswarm.respawnGraceMin (20) minutes have passed; nothing runs it automatically and it never kills. In order: it asks the child to commit and push its WIP and waits up to devswarm.respawnWipWaitSec (120 s); commits anything still dirty or unpushed through a private index onto a new park/<branch>-<ts> branch and pushes it (the child's worktree is untouched, nothing is stashed or discarded, and a failed park or push aborts the respawn with the local park branch kept); writes plans/<id>.handover.md; spawns <branch>-r<N> with -s <default branch> (not -s <old branch>, which would make the old branch the merge target) whose step 1 merges the old or park branch, followed by the remaining steps, scope and extras; archives the old id and asks the owner to close its app tab. --dry-run prints the plan with no side effects. Metrics: respawns, WIP parked or not, aborts, time to first step progress in the new workspace and whether it finished, in supervision-report.
  • Token savings. A 24h field measurement found idle DevSwarm wake turns re-paying ~620K-930K tokens each (mostly cache reads) against a ~90K fixed per-turn floor, with 48 Monitor-expiry turns and 45 mailbox-tick turns running on overlapping 30-minute cadences. Four owner-approved levers, each with a settings key:
  • Monitor re-arm only at the cron tick (devswarm.rearmOnTickOnly, default true). The SessionStart wake directive (hooks/lib/devswarm-wake.js monitorArmLine) no longer tells the agent to re-arm a lapsed Monitor watcher inline on the tool's own final/expired event — it replies in ≤1 line and stops; the cron tick's own watcherArmed:false check (drainCmd's rearmClause) is the one re-arm trigger. false restores the pre-0.117.0 wording. The cron fallback itself is never disarmed (unchanged, non-negotiable). Metric: ~/.anti-hall/devswarm/rearm-cues.jsonl records each re-arm cue by trigger (tick/expiry; scripts/devswarm.js recordRearmCue), written from inbox tick's existing watcherArmed computation.
  • Prune completed tasks (guards.pruneCompletedTasksAfter, default 10). Claude Code's own TaskCreate reminder re-prints the whole task list, completed tasks included, every few turns. Once completed/cancelled tasks exceed the threshold, task-guard.js emits a one-line advisory at Stop (never a block) to prune them via TaskUpdate status=deleted after recording them in the history ledger. Metric: the completed-task count at Stop (visible in the advisory itself). Advisory only — never gates the Stop hook.
  • Deadly-loop scales to risk. skills/deadly-loop, skills/ship-it (and their Codex mirrors codex/skills/anti-hall-deadly-loop, codex/skills/anti-hall-ship-it) now lead with a decision table: the full Reviewer+Auditor+Critic trio only for guard/security/parser/schema/CI/shell changes, one reviewer for normal code, none for text/doc-only changes. Guidance only, no metric.
  • Periodic peer/check crons default to a longer cadence. Verified (git grep/command grep, not the shell grep alias) against every CronCreate reference in plugins/anti-hall/skills and plugins/anti-hall/hooks: the only agent-created periodic cron in the plugin is the DevSwarm mailbox-wake tick, which this release deliberately leaves at its existing 30-minute cadence per the owner's mailbox-tick carve-out. No second periodic peer/check-in cron directive exists in the codebase to retune.
  • Jev: dispatchTier, speculationFramed, and a durable review-due reminder. dispatchTier is a shadow advisory on whether pending work belongs in a DevSwarm child workspace or a plain subagent. speculation-guard's framed-expectation calls (speculationFramed) are now judged by Jev in shadow instead of the bare heuristic alone. A JEV REVIEW DUE reminder (SessionStart + doctor) resurfaces periodically so labelled-but-unreviewed Jev calls don't go stale.
  • Orchestration: dispatch demand survives the machine-wide heartbeat, respects maxParallelDispatch, and honours addBlockedBy. The parallel-dispatch nudge used to be blanket-suppressed by a single machine-wide heartbeat, so it could silently stop firing even with pending unblocked tasks; it now fires per pending/unblocked task and is capped by the new maxParallelDispatch setting. SendMessage-resumed agents now count as running (not orphaned) against that cap. Tasks are scoped to the current task-store epoch so a restored/reset store doesn't resurrect stale nudges, and nudges now show the task subject. addBlockedBy dependencies are honoured with an order-independent fixpoint (see the 0.116.0 cyclic-blockedBy fix); a live DevSwarm child workspace now counts as "attended" for IDLE NEGLECT purposes, and IDLE NEGLECT uses one consistent priority filter.
  • DevSwarm misc: a maintainer notice lets one post reach every project's Primary; doctor/wake now warn when the mailbox cron has stopped ticking (e.g. after a restore) instead of going silently dark; hivecontrol startup-state samples are captured to detect a crash-paused workspace; an archived child that tries to re-register is told to hand over and stop instead of continuing; hook wake text now points at one consistent CLI path and wake instruction everywhere it's quoted.
  • Mesh read directive corrected. The SessionStart COMMUNICATION OVERRIDE text grouped mesh read with the genuinely read-only roster/inbox read-primary calls; mesh read with no flags actually advances the caller's broadcast cursor and returns no ackCommand. The directive now flags mesh read as consuming (use mesh read --peek to preview without consuming) and scopes the ackCommand claim to inbox read-primary, where it's actually true.

Fixed

  • Plan-file writes are locked (no lost step updates). The supervisor sweep's Jev step-map write and the child's heartbeat --step, scope add, plan set and done (and the Primary's correct) each did an unlocked read-modify-write of plans/<key>.json, so a write landing between another writer's read and its rename was dropped. Every writer now goes through one read-modify-write under the plan's lock (companion/lib/lock.js) on the fresh on-disk plan.
  • AGENTS.md regenerated (tools/gen-agents-catalog.js) to cover the two new settings keys; still fits the Codex 32 KiB cap.
  • work-detect: a housekeeping segment can no longer hide real work. isDevswarmHousekeepingOnly classed a whole segment as pure DevSwarm/crontab housekeeping once its leading verb matched, even when that same segment also redirected into a real file (crontab -l > src/app.js, node devswarm.js roster > README.md) or hid work in a command substitution (crontab -l $(sed -i s/a/b/ src/x.js)). Each housekeeping-looking segment is now also checked for an escaping write/substitution before being excluded from the tasklist-guard work count; the legitimate mailbox-wake crontab install (redirect target under /tmp or the session scratchpad) still stays housekeeping-only.
  • compact-declaration-guard/compact-advice-guard: the "SAFE TO COMPACT" matcher no longer fires on a quoted mention, a question, or a negated/conditional sentence. findAdvice matched the bare phrase anywhere in the text, so a reply merely describing the guard (`SAFE TO COMPACT` is declared, 'SAFE TO COMPACT' must be last), asking about it (Is it safe to compact now?), or ruling it out (far from safe to compact, once this lands it will be safe to compact; first I need to …) was misread as a live declaration. stripQuoted now blanks single-quoted and backtick-quoted spans (unless the quoted text itself names a /compact//clear//new invocation), matches inside a question-ending sentence are ignored, negation detection covers far from, not yet, and the once … it will be … ; first … conditional shape, and the bare safe to compact/clear wording only counts at a line/sentence start or as the unambiguous ALL-CAPS SAFE TO COMPACT form — the more specific forms (good point to /compact, safe for a context reset, run/then/now /compact, a standalone /compact line) are unaffected.
  • guards.compactDeclarationGuard default reset to ON. It shipped opt-in (default off) in 0.116.0 pending the false-positive fixes above; now that the shared matcher ignores quoted/question/negated/conditional mentions, the PreToolUse block is back on by default. Opt back out with /anti-hall:settings if needed.
  • git-guard: a round of self-credit and force-push hardening. Self-credit detection now covers pipes, variables, same-command files, merge/commit-tree/rebase -x, interpret-trailers, gh pr merge --body-file, hook-authored trailers, and git tag -m/-F (with a widened -F stdin trailer scan); force-push detection covers command-valued config/env values and quoted code-string literals passed to a call, and an alias-body force flag is now found however its body is quoted (splitting on quotes/parens/braces/shell separators, not whitespace alone). Writes into the ~/.anti-hall/bin launcher directory are blocked (best-effort; a launcher-dir bypass gap closed in a follow-up round). The >| clobber-redirect tokenizer bug (pre-existing) is fixed, a dangling-symlink target no longer false-blocks, and Write/Edit now gets a hint to use those tools instead of a Bash heredoc-to-file, with Write/Edit itself denied on a matching block. The heredoc mailbox-message exemption (added mid-release to let a legitimate mailbox heredoc through) was removed after it failed security review — no heredoc-shaped exemption ships in 0.117.0.
  • git-guard: follow-up hardening rounds (self-credit Jev consults, tokenizer, rm-through-symlink, project .anti-hall/ copies). Jev self-credit consults are now memoised, capped (8 distinct texts per hook call) and bounded to 4 s total consult time per invocation, so a pathological command can no longer exceed the PreToolUse hook timeout. The escaped-operator tokenizer now handles \>| and \>& correctly. rm through a symlinked parent directory into ~/.anti-hall/bin is blocked (a real file/dir reached via a symlinked parent component previously bypassed the launcher-dir guard). Ordinary project-local .anti-hall/ copies (a project's own .anti-hall/, not the launcher dir) are no longer false-blocked.
  • command-guard: plain-push and read-only carve-outs tightened. A colon-bearing plain-push ref (push origin :refs/...) now requires an explicit remote instead of being silently accepted; allow-plain-push now also accepts HEAD:refs/heads/<current>, a trailing 2>&1, and a final | tail/head filter without losing its narrow scope.
  • devswarm respawn: secret-shaped untracked files are excluded from the WIP park branch. .env, key/credential-shaped filenames and similar are never committed onto a park/* branch by the respawn WIP-park step, even when they're dirty/untracked in the straying child's worktree.
  • tasklist-guard: a reset task store after a session restore now gets an explanation instead of a bare nudge; the progress-freshness path now keys on the session's project root, not a submodule cwd, so progress recorded from inside a submodule is found correctly.
  • wake-watch: consistent primary role from repo root or subfolder — armed from a subdirectory no longer flips role detection relative to being armed from the worktree root.
  • work-detect: anti-hall's own bookkeeping and scratch writes are no longer counted as "work" for freshness purposes, and real work chained onto a housekeeping command (e.g. crontab -l && sed -i ... src/app.js) still counts — closing both directions of the false-classification gap the earlier housekeeping-segment fix opened.
  • limit-conserve: an expired reset window no longer injects a stale "conserve tokens" reminder.
  • mcp-reaper: recognises <name>-mcp-suffixed server process names; orphan-kill rules are unchanged.
  • doc coverage + regenerated catalogs. Doc-coverage gaps closed and ratchet baselines bumped for the v117 merge; AGENTS.md and docs/KB.md counts regenerated to match.

0.116.1 (2026-09-28)

Fixed

  • wake-watch: no more false "+N new mail" wake after a version handoff, or when an older hook rewrites the mailbox summary. A counter with no recorded history (an older seen-file predating it, or none at all) was diffed against a fabricated 0 baseline on the very next arm, producing an immediate false wake for history the watcher never actually had. That counter is now seeded from its own first live observation instead. A related flap: the broadcast channel's bucket is now chosen per row (the first bucket whose row actually carries a total), matching how the primary snapshot picks it, instead of falling through per field — an older build rewriting the summary without one field no longer flips the channel onto an unrelated, stale bucket and re-fires a false wake on the next new-build rewrite. The seen-file merge also now preserves any keys a newer build wrote that this build doesn't own, instead of dropping them on rewrite.
  • Handoff tests reap their child processes. The handoff-forwarding test suite was leaking orphaned watcher processes; it now reaps its own children.
  • The update skill runs update.js on a sonnet subagent, matching the model-routing floor instead of a heavier default.
  • The post-handover new-work gate nudge skips scheduled housekeeping prompts (a mailbox-wake tick, a DevSwarm peer-check tick, a broadcast bug-sweep tick) — these fire on every prompt and can't act on a "park this work" nudge, so gating them was pure noise. Real user prompts are unaffected. A new autoHandover.gateHousekeepingMarkers setting adds extra markers on top of the built-in defaults.

0.116.0 (2026-09-27)

Added

  • compact-advice-guard (Stop): no /compact recommendation at low context or right after a compact. Field defect: a reply declared "SAFE TO COMPACT NOW" and repeated a /compact focus: … line a few turns after a manual /compact, at low context, with "no background agents running" as its only reason. The guard blocks once per declaration when the final reply recommends compacting and either context % is below autoHandover.pct - guards.compactAdviceMarginPct, or a compact boundary (Claude compact_boundary; Codex compacted/context_compacted) is within guards.compactAdviceRecentTurns turns. The threshold-fired auto-handover path stays allowed; double-quoted spans, fenced code blocks, and phrasing immediately preceded by "not"/"n't"/"never"/"no need"/"no reason"/"retract(ed/ing)" are excluded from matching (single-quoted/backticked mentions and other negated or questioning phrasing are not — see the opt-in note below). Registered for Claude and Codex.
  • compact-declaration-guard (PreToolUse, opt-in — off by default): no new work in the turn after declaring SAFE TO COMPACT. Field defect: after the declaration the model started a background agent and a state-changing shell command in the same turn. Agent/Task spawns, Write/Edit/MultiEdit/NotebookEdit and state-changing Bash are blocked while the current turn holds an active declaration; read-only tools pass. The next user message or an explicit RETRACT SAFE TO COMPACT — <why> line clears it (task notifications do not). Codex: Bash only, matching the port's PreToolUse policy. Ships with guards.compactDeclarationGuard defaulting to false in 0.116.0: a reviewer probe found the shared quote-stripping/negation matcher it uses (lib/compact-advice.js) still hard-blocks a turn on harmless single-/backtick-quoted mentions, questions ("Is it safe to compact now? No"), and negated phrasing outside its narrow lookbehind ("far from safe to compact", "will be safe to compact") — false positives are too costly on a hard PreToolUse block. Opt in with /anti-hall:settings once you've read the caveat; the Stop-time compact-advice-guard above stays on by default (a soft, once-per-declaration, recoverable block).
  • The shared compact-advice matcher also recognizes the Codex handover wording ("GOOD POINT FOR /compact OR /new NOW", "SAFE for a context reset") and its retractions.
  • Handover skill (Claude + Codex): a quiet background queue is necessary, not sufficient, to declare SAFE TO COMPACT — declare it only on the auto-handover threshold or genuinely high context, never within guards.compactAdviceRecentTurns turns of a compact, and retract before doing more work.
  • Jev: daily rollups and longer log retention. jev-assist.ndjson kept one 1MB backup (about a day) and jev-triage.ndjson truncated itself at 1MB, so no audit could see more than a day. jev-assist.ndjson now rotates at 2MB and jev-triage.ndjson at 1MB into .1 … .N (jev.logRotatedFiles, default 10, about 20 days); lowering N never removes generations already on disk. Before each rotation every retained row is folded into ~/.anti-hall/logs/jev-daily/<YYYY-MM-DD>.json (per id × backend × mode: counts, fresh calls, changed-vs-baseline, timeouts, failures, cost, p50/p95 latency, outcomes). Rollups are never removed by default; jev.rollupRetentionDays is an opt-in.
  • jev-report reads history. It reads every rotated generation (not only live + .1), and summarises days the raw logs no longer cover from the daily rollups (rollupHistory in --json, an "Older history" text section). A day is never counted twice; rolled-up p50/p95 are labelled approximate.
  • settings show --section jev lists every Jev integration's effective mode (master switch and env kill switches folded in), the stored value, its source tier and the log it writes to (jevEffectiveIntegrations in --json).

Fixes

From a peer workspace sweep:

  • Scratchpad exemption encodes the cwd like the harness. The harness names the per-project scratchpad dir by replacing every non-alphanumeric cwd character with -; the shared helper replaced only /, so any cwd containing . or _ (e.g. dotted worktree paths) computed a dir that never exists and the own-scratchpad exemption never matched in edit-guard or command-guard. One shared encoder now mirrors the harness; the session id is still required, so sibling sessions stay blocked.
  • command-guard: 2>&1 / >&2 / &> are redirections, not background separators. The splitter split on every lone &, so x 2>&1 | tail never reached the bounded-check allowance. An & after >/< (fd dup) or before > now stays in the word; a real background & still splits. Only 2>&1 is ignored when judging a check's shape; any other >& routes output around the sink and disqualifies it, and &>file targets go through the write-redirect rule.
  • command-guard: the delegation hint names the exact inline-allowed shapes instead of promising any "bounded single-target check": python3 -m pytest -q <one file>, node --test <1-2 files>, ctest -R <name>, <cc> -fsyntax-only, git clone --depth 1 <https-url> <scratch/tmp dir>, a non-heavy command with --check/--dry-run/--list, or an interpreter on an existing script with --check, each piped to tail/head/wc/grep -c/grep -m N; scratchpad scripts are run_in_background only. The guards.allowReadOnlyVerify descriptions (GUIDE, /config, settings schema) now say the same.
  • tasklist-guard: progress freshness is relative to the last work, not wall-clock. An idle turn (cron tick, status check) more than 30 min after the last progress write was blocked although no work had happened since. Freshness is now judged against the newest file-changing action when known; new work after the progress write still blocks.
  • task-guard: a self-referential or cyclic blockedBy no longer counts as an honest block. A blocked by A, or A/B blocking each other, satisfied the old membership check and silenced the generic drain nudge for tasks that are not honestly blocked at all. an order-independent set fixpoint over the blockedBy graph now requires a blocker to name a different open task whose own chain reaches something that can actually progress (no honest blocker, or an owner/user/external blockedOn): self-references, pure cycles and dangling ids never count as honest blockers, and the verdict no longer depends on task order.
  • speculation-guard: quote/blockquote masking can no longer hide the session's own hedge. A > line is now masked only up to the earliest sentence-break after quoted text (em dash, "; so", ", so", …), so a hedge appended after a quote on the same line is still caught; a reply that is 100% blockquoted/fenced is no longer masked into nothing, and neither is a reply that is a single straight-quoted or inline-code hedge with nowhere else to hide (masking to nothing falls back to matching the unmasked text). Straight " pairing is now per-line with an even-count guard, so a stray unpaired quote no longer mispairs across the line.
  • devswarm send: a no-project failure explains the fix (run from inside the repo) instead of returning a bare reason code, like relay and reap-orphans.
  • command-guard: narrow read-only gcloud reads accept a 2>&1 stderr merge. A trailing, unquoted, space-separated 2>&1 is now the one redirection a gcloud read may carry (whole-command carve-out and per-segment inspect exemption); every other redirection, and a redirect glued to a flag value, is still refused.
  • devswarm: stale-main spawn refusal names the submodule it refers to. Run from inside a submodule, the refusal now leads with submodule <name>: (the path in the superproject, else the .git/modules dir) instead of pointing at the meta-repo.
  • devswarm: WAL-health tick performance. health() re-scanned the full os.tmpdir() per WAL (123 readdirSync calls on a 33k-entry tmpdir per tick, ~96% of inbox tick CPU); it now lists tmpdir once per call and threads the filtered list through, cutting a tick from ~5.5s to ~0.15s CPU with unchanged behavior.
  • devswarm-wake-watch: a false wake on role flip. A Primary whose descriptor sits at the main worktree resolves as child when armed from the worktree root and primary when armed from a subdirectory; both roles shared one lastTotal field but meant different counters (NDJSON count vs. mesh summary total), so a re-arm under the other role compared against a stale, wrong-shaped baseline and fired a false wake. The seen file now carries meshTotal/ndjsonTotal by name and each role reads/writes its own; legacy files migrate on read.
  • devswarm reconcile: archived/pruned worktrees are skipped, not reported as failures with no detail. An archived-but-present row reached the app-DB archive guard, which refuses with a reason and no .error field, so the fallback printed "unknown error"; an archived+pruned row hit the missing-worktree path with no concept of "archived". cmdReconcile now classifies archived rows (via the DevSwarm app DB or anti-hall's own archived marker) as skipped with a skipReason, separate from genuine failures, and the per-target error now falls back through the parsed reason and the hivecontrol exit code/stderr before ever saying "unknown error". update.js and doctor-repair.js report skipped counts separately.

Review-wave hardening (a second pass over the peer sweep above):

  • task-guard and speculation-guard above are all second-pass hardening from this same review wave (self/cyclic blockedBy, blockquote/quote masking).
  • compact-advice-guard/compact-declaration-guard: prefer last_assistant_message, block like sibling PreToolUse guards. Claude/Codex Stop payloads carry last_assistant_message — the exact text the model just produced — which is now preferred over the transcript-tail fallback; compact-declaration-guard now blocks via exit 2 like command-guard.js/edit-guard.js instead of a flat {decision:'block'} on exit 0.
  • compact-advice-guard: declaration matching requires explicit forms only. "safe to clear the cache" and a bullet merely mentioning /compact (or a fenced-code example of it) no longer count as a declaration; only explicit forms (SAFE TO COMPACT, safe to /compact, good point to /compact, /compact now, run /compact, the Codex skill's GOOD POINT FOR /compact OR /new) match, and fenced code blocks are blanked before matching alongside the existing blockquote/quoted-text stripping. A tokens-latch handover (firedVia:'tokens', or 'stop-tokens' from the Stop-time nag) may now declare SAFE below the pct threshold too, as long as no compact has happened since the latch fired.

Wave-2 fixes (a second peer sweep):

  • devswarm: inbox messages --since filters before the per-source row cap. A --since window on an inbox with more rows than the cap could come back empty even though matching rows existed, because the cap truncated the source before the since-filter ever ran; the filter now applies first, so a recent window is never emptied by an oversized old inbox. Behavior change: a plain (non-acking) inbox messages read that the cap truncates now returns the NEWEST rows per source instead of the oldest; read-primary and its non-mutating preview peek-primary keep the oldest prefix (cursor safety), so peek-primary still shows exactly what read-primary will deliver.
  • devswarm-wake-watch: a sender's own broadcast/direct send no longer wakes it. A send loop back to the sending session's own inbox counted as new mail and re-armed its own wake, producing a self-triggered wake cycle; a sender's own outgoing message is now excluded from what counts as new mail for it.
  • devswarm: emitted ackCommand uses the stable launcher instead of a raw path, matching every other devswarm-emitted command shape and staying correct across install layouts.
  • devswarm-child-gate (Stop): a --summary heartbeat sent earlier in the turn now satisfies the gate. The gate's activity window was a fixed 5-minute wall-clock interval, so a heartbeat sent earlier in a turn longer than 5 minutes read as stale and the gate re-demanded one. The window is now anchored to this session's previous Stop check (persisted lastCheckAt), with the 5-minute floor only for a session's first check. The continuation Stop right after a block also records its check time, so a report sent in that continuation does not keep satisfying later turns.
  • tasklist-guard: sees an early TaskCreate, and devswarm/cron/agent spawns no longer count as file changes. A TaskCreate issued before any file-changing tool call wasn't picked up by the freshness check; housekeeping tool uses (Agent, Task, CronCreate, CronDelete, devswarm stable-launcher chains) are now excluded from what counts as work, so they no longer mask genuinely stale progress or falsely satisfy it.
  • devswarm: one consistent inbox-drain instruction. read-primary and the emitted ackCommand previously implied two different drain sequences; both paths now point at the same instruction.
  • devswarm-comms-guard: no unresolved-advisory for in-process subagent targets. A send addressed to an in-process subagent id was flagged as an unresolved target even though it never needed mesh resolution; subagent ids are now recognized and skipped.

Deadly-loop round-2 fixes:

  • git-guard: a force flag inside a !shell alias body is found however it is quoted. git -c alias.x='!sh -c "git push --force"' x, '!f() { git push --force; }; f', '!eval "…"', the reversed-quote form and '!(git push --force)' were allowed because the body was split on whitespace only, leaving --force", --force; or --force). Alias bodies are now split on quotes, parens, braces and shell separators too, and an inner -- of another command in a shell body no longer disarms the check. This gap predates 0.116.
  • devswarm-comms-guard: the session-index lookup runs before the agent-id silent allow. The widened agent-id pattern (bare a<hex>) was tested first, so a workspace-backed peer session whose name happened to be agent-id-shaped was silently allowed. The target is now looked up first; a workspace-backed match is blocked whatever its shape, and only unresolved or non-workspace agent-id targets are silently allowed.

Jev logging:

  • parentGateQuestion decisions were never logged. The triage and assist caches evict the lowest _seq first, but _seq came from a per-process counter that restarted at 1 in every hook process; once a cache reached 500 entries each new entry sorted lowest and was evicted in the same write. The triage cache froze, the question-needs-answer lookup never found a label, and parentGateQuestion never reached its decision. _seq is now seeded from the stored max; an already-frozen cache heals in place.
  • jevIntegrations.triage: off now disables triage labelling (it only checked jev.triage); ANTIHALL_JEV_TRIAGE=0 does too.
  • modelRouting is consulted only on routing blocks, so a window with no blocks correctly logs zero rows (documented, fixture-tested).

New settings

Key Default /config
guards.compactAdviceGuard true guards_compact_advice_guard
guards.compactAdviceRecentTurns 10 guards_compact_advice_recent_turns
guards.compactAdviceMarginPct (advanced) 10 —
guards.compactDeclarationGuard false (opt-in) guards_compact_declaration_guard
jev.logRotatedFiles (advanced) 10 —
jev.rollupRetentionDays (advanced) 0 (keep all) —

0.115.2 (2026-09-27)

Fixes

  • command-guard: a shell comment containing ) no longer triggers a false heavy-command block (peer-reported). The segment splitter split on ), ;, |, & even inside comment text, so # 1) go to x produced a bogus go to x segment classified as the heavy verb go. An unquoted # that starts a word (start of input, or after unescaped whitespace, ;, &, |) at nesting depth 0 and outside backticks is now dropped to end of line, as the shell does. A # that is not a comment stays code (a#b, $#, ${#x}, ${x#y}, ${x:- #}, (( 2 #)), $[1 #], backticks, $(echo a)#b, escaped/quoted #, continuation lines); the newline still ends the comment; the existing # refusal for the allowlist/bounded-sink paths is unchanged.
  • command-guard: four splitter bypasses closed (latent, not new regressions). Each let a real chained command (e.g. ; go build) reach the shell unclassified:
  • ANSI-C $'…' quoting was read as plain '…', so an escaped \' closed the quote early and hid what followed — in both the segment splitter and the $( … ) substitution extractor.
  • A <<< here-string was parsed as a heredoc on its operand (<<< "#"), swallowing every following line as "body".
  • A heredoc delimiter was cut at the first non-identifier character (<<EOF#x read as EOF), so the body never terminated. The delimiter is now the full shell word with quote removal; words the parser cannot model exactly are not treated as heredocs (nothing is skipped).
  • command-guard / git-guard: << inside arithmetic ($(( … )), (( … )), $[ … ]) is a left shift, not a heredoc. A letter-led operand ($((1<<y))) opened a fake heredoc whose "body" hid every following line. The shared heredoc parser now rejects << whose innermost enclosing context is arithmetic, so every caller (both guards' heredoc scans and the substitution extractor) is covered.

0.115.1 (2026-09-26)

Fixes

  • devswarm-wake-watch: orphaned watchers never exited when their starting parent (Monitor shell, or the stable-launcher wrapper ~/.anti-hall/bin/wake-watch.js) died. Field evidence: a watcher started by a test's intermediate parent 19 hours earlier was still running with PPID 1, holding its per-child watch lock — the same shape can leave a live session's next watcher refusing to arm (REFUSED TO ARM: lock-held) or reporting watcherArmed from a dead process. devswarm-wake-watch.js now captures its startup process.ppid and checks every poll tick (parentGone()): it exits cleanly (releasing its lock) the moment it is reparented (ppid changed / became 1) or process.kill(startPpid, 0) throws ESRCH.
  • stable-launcher: the generated ~/.anti-hall/bin/ wrapper forwards SIGTERM/SIGINT/SIGHUP to its target and spawns asynchronously, instead of blocking synchronously with no way to react to a graceful shutdown signal; stdio and exit code passthrough are unchanged.
  • A fleet test (devswarm-fleet-8143ced316d3.test.js) no longer leaks a live watcher process on assertion failure — the spawned watcher child is now cleaned up unconditionally, not only on the success path.

0.115.0 (2026-09-26)

Changed

  • doctor --prune-cache warns that a cron or Monitor job naming a versioned cache path directly will break once that version is pruned, and points it at the version-independent ~/.anti-hall/bin/ launchers (devswarm.js, wake-watch.js) instead.

Fixes

  • doctor --prune-cache keeps versions referenced in recent session transcripts. A CronCreate/Monitor job created during a session (hooks cannot read CronList) could still literally name a versioned cache path; pruning that version left the job failing silently on every tick. Prune now does a cheap, bounded (<=2s, <=20 files, 512 KB tail each), read-only best-effort scan of the newest recent (<=7 day) session transcripts under ~/.claude/projects/*/ for any versioned cache path they mention, and keeps those versions. The scan is bounded and fail-safe: if the scan fails, nothing is removed.

0.114.1 (2026-09-26)

Fixes

  • command-guard: P0 — the version-independent stable-launcher form (~/.anti-hall/bin/devswarm.js / wake-watch.js) was wrongly blocked as a heavy command. Since 0.109 (devswarm.stableLauncher, default on), every hook-emitted directive — the mailbox wake cron prompt, the Monitor re-arm command, the DevSwarm comms-override role text, and Stop-gate drain/handover pointers — names the version-independent launcher under ~/.anti-hall/bin/ instead of the plugin-relative scripts/devswarm.js path. The command-guard allowlist only ever anchored scripts/devswarm.js, so every one of those emitted commands fell through to the generic node <file>.js HEAVY_PATTERN and was blocked in coordinator context — breaking every Primary's cron tick and inline mesh command that used the launcher form (peer-reported by the downstream Primary). Added anchoredAntiHallStableLauncher(), a home-anchored carve-out (accepts ~, $HOME, "${HOME}", and the resolved absolute home directory immediately followed by /.anti-hall/bin/devswarm.js or /.anti-hall/bin/wake-watch.js, no other prefix) with the same whole-invocation scope as the existing scripts/devswarm.js carve-out. A look-alike path (evil/.anti-hall/bin/devswarm.js, /tmp/x/.anti-hall/bin/devswarm.js) is still NOT exempt, and chaining a second heavy command after the launcher invocation still blocks on that segment.

0.114.0 (2026-09-26)

Fixes

  • command-guard: P1 — the leading-cd carve-out let a nested repo (submodule or untracked nested .git) hijack a push. cd realsub && git add z && git commit -m x && git push origin subbr qualified for the plain-push carve-out whenever realsub merely lived under the outer repo's directory tree, even when it was a git submodule or any independently-git init'd nested repo with its own .git/remote/branch — branch/remote resolution then ran against the WRONG repository. Path containment was never a repo-identity check. resolvedLeadingCdTarget now requires the cd target to share the payload cwd's git-common-dir (the real .git store) — linked worktrees of the same repo share it and still qualify, but a submodule or untracked nested repo never does. Fails closed on any git error.

  • command-guard: "allow plain push" missed three common shapes. Peer-reported: cd <repo> && git add a b && git commit -q -m "fix: x" && git push -q origin main && git log --oneline -1 was blocked as "heavy-pattern" even though it should have qualified for the carve-out. Root-caused to three independent gaps, all reproduced first with the exact command via a real PreToolUse payload: (a) -q/--quiet on push was not recognized by PLAIN_PUSH_SEGMENT_RE; (b) a leading cd <path> segment was not recognized at all; (c) a trailing read-only segment (git log/status/show) was not recognized at all. isAllowedPlainPushChain now accepts -q/--quiet on push (still refusing it combined with --force or any other flag), ONE optional leading cd <path> — only when it realpaths to the payload cwd's own repo toplevel or a directory inside it (fails closed on another repo, a $() argument, or an unresolvable path) — and optional trailing git log --oneline [-N] / git status [--short|-s] / git show --stat [-N|HEAD], only after a push segment has already appeared in the chain.

  • doctor: orphaned-workspace-process remedy now runs anti-hall's own archive verb, not raw hivecontrol. The re-archive remedy for an archived/app-archived hit printed hivecontrol workspace archive <id> directly. It now prints node <cli> archive <id> (the stable-launcher path when available, falling back to the plugin's own scripts/devswarm.js) — anti-hall's own verb carries the app-archive retry-once fix above plus the app-DB target-gate verification, neither of which a bare hivecontrol call gets.

  • doctor: orphaned-workspace-process report gave the wrong remedy for an ARCHIVED (not gone) workspace. Peer report: the claude process in an archived DevSwarm tab stays alive, and killing it just relaunches it from the pty shell — the only real remedy is re-archiving the workspace in the app. orphanedWorkspaceProcessCheck previously suggested kill <pid> for every hit regardless of reason. It now prints the exact safe remedy line (hivecontrol workspace archive <full id>) with an explicit "do not kill it — the pty shell relaunches it" warning for archived/app-archived hits, and keeps the plain kill suggestion for a genuinely gone workspace (worktree removed from under a still-registered row, no relaunch mechanism). Still report-only — never kills or archives anything itself.

  • devswarm: app-archive retry regex missed a reworded DevSwarm error text. attemptAppArchive's single retry only fired on the exact phrase "could not confirm terminal process boundary"; a later DevSwarm build rewords the same transient failure as "Could not confirm terminal stopped", so every first archive attempt failed and the built-in retry never triggered. APP_ARCHIVE_RETRYABLE_RE now matches the stable "could not confirm terminal" prefix, covering both wordings while still excluding unrelated hivecontrol errors from the retry.

0.113.0 (2026-09-26)

Features

  • command-guard: narrow read-only gcloud in the main thread (owner-approved). Three shapes now run inline instead of being delegated: gcloud auth print-access-token on its own; gcloud <group…> <describe|list|get-iam-policy|read> … --format=json|yaml|value(...), optionally piped into tail/head/wc/grep -c/grep -m N/jq; and T=$(gcloud auth print-access-token); curl -s|-sS [-H "Authorization: Bearer $T"] <https URL> (; or &&). The URL's parsed host must be googleapis.com or a subdomain of it (no userinfo, no IP literal, no other host), and -L/--location, --resolve, --connect-to, -x/--proxy, --url and -K/--config are refused, so the token never leaves Google. The curl must be a GET, and its output must be piped into a bounded sink or jq, or capped with --max-filesize. Everything else stays blocked: create/delete/deploy/set/update/add-iam-policy-binding/remove-*/patch/import/export/ rollback/start/stop/ssh/scp/submit/run/apply, curl -X other than GET, -d/--data*/-F/-T/--upload-file/-o/-O/--output, @file, any other curl flag, redirects, and any chained segment. $T may appear only in the Authorization header. Subagents are unaffected. New setting guards.allowGcloudReads (default true).
  • command-guard: background scratch scripts in the main thread (owner-approved). A Bash call with run_in_background: true may run ONE segment <python3|node|sh|bash> <script file> [args…] inline when the file is an existing regular file in the session scratchpad or a tmp root (os.tmpdir(), /tmp, /private/tmp). The path is checked on its realpath, so a symlink out is refused. Refused: an interpreter option before the file (-c/-e/-m), an env prefix or wrapper, chaining, pipes, $/backtick/backslash/process substitution, a stdin redirect, and a write redirect outside the scratchpad/tmp. Foreground runs keep today's verdict, and Monitor is unchanged. New setting guards.allowBackgroundScratchScripts (default true).
  • doctor --prune-cache (opt-in plugin cache prune, owner-approved). It never runs automatically: update.js, the supervisor, crons, SessionStart and hooks never call it. Without --confirmed it lists the old ~/.claude/plugins/cache/anti-hall/anti-hall/<semver>/ dirs it would remove and their total size. --confirmed removes them and logs each removal. It always keeps the newest 3, every dir holding an installPath registered in installed_plugins.json, every version a live process runs from (process cwd via the 0.111 doctor scan, or the cache path in its argv), the running version, and anything it cannot parse. Symlinks and paths outside that root are refused, both when listing and again just before each removal. If the live-process scan is unavailable, nothing is removed. New setting updates.allowCachePrune (default true), which only enables the verb.

Security fixes

  • command-guard gcloud reads: the read verb must be the last command-path word (P1). The new read-only gcloud carve-out, and the older read-only cloud-inspect exemption, accepted read/list/get-iam-policy in ANY position. That included a separated flag value, so gcloud compute instances reset vm1 --zone read --format=json, gcloud secrets versions access latest --secret read …, gcloud kms decrypt … read … and similar ran inline. Both paths now accept only gcloud <group…> <verb> [≤1 positional] [--k=v…]. The verb is the last path word. Every flag is --k=v or a known boolean (--quiet, --uri); a separated value is refused. No path word may be access/reset/suspend/resume/publish/call/execute/decrypt/ encrypt/sign/print-/attach-/detach-/add-/set-/remove-/delete/create/update/ deploy/ssh/scp/run, nor a hyphenated form of a mutating action (delete-access-config, reset-windows-password, …). run is allowed only as the product group, and read only as logging read. This also closes a pre-existing hole in the older read-only cloud-inspect list exemption, which took the same separated-flag-value bypass.

  • command-guard gcloud token curl: closed set of token variable names (P2). The token may only be assigned to T, TOKEN, ACCESS_TOKEN or GCLOUD_TOKEN. Any other name (HTTPS_PROXY, http_proxy, CURL_CA_BUNDLE, SSLKEYLOGFILE, …) could be an env var curl itself reads, and is refused.

  • doctor --prune-cache: registered installPath compared canonically (P2). Both sides go through fs.realpathSync.native, and the comparison is case-insensitive on darwin. A registered path spelled in a different case, or reached via /tmp, .. or a trailing slash, is kept.

  • doctor --prune-cache: live-process match on the cache suffix (P3). A process cwd or argv now counts as live when it names plugins/cache/anti-hall/anti-hall/<ver> followed by a path boundary, whatever the prefix (/tmp vs /private/tmp, a symlinked home). <ver>.bak never matches.

  • command-guard gcloud read sinks: no env/input/file access (P3). A jq filter after a gcloud read or the token curl may not use env, $ENV, input, inputs, input_filename, import, include or $__loc__. grep -f/--file, including in a short-flag cluster, is refused.

  • command-guard background scratch scripts: same F1 rule as the script check (P3). A script whose realpath is inside any anti-hall plugin root (this install, a cache copy, or a dev checkout), or any --confirmed argument, never qualifies.

  • command-guard gcloud token curl: canonical ASCII host only. The raw URL host must be plain ASCII and equal to the parsed host, with no curl URL globbing ({…}, […]). This refuses percent-encoded and full-width lookalike hosts.

Changed

  • command-guard gcloud reads: separated flag values are no longer read-only. gcloud … describe foo --region r (a value in its own argv word) is no longer treated as read-only inline; write it --region=r to keep running inline. This closes the separated-flag-value gap the P1 fix above found in both the new gcloud read carve-out and the older read-only cloud-inspect exemption.

0.112.0 (2026-09-26)

Features

  • edit-guard: a per-project doc-edit allowlist (owner-approved). A repo may list repo-relative globs in .anti-hall/edit-allow.json ({"paths":["docs/**","PLAN.md","*.md"]}). The main thread (a DevSwarm Primary or coordinator) may then Edit/Write matching files directly instead of delegating them. It uses the 0.111 command-allowlist trust model and the same lib/command-allow.js machinery. It applies only after settings.js trust-edit-allow <repo> --confirmed records the sha256 of the exact file bytes in ~/.anti-hall/trusted-edit-allow.json. Any edit to the file revokes trust, and a symlinked file is refused. Absolute, .. and match-everything globs are ignored. A match never covers a path outside the repo (checked on real paths), .git, .anti-hall, .claude, .codex, hook config, ~/.claude, or a symlinked or hard-linked target. Each path segment is folded with NFKC plus lowercase before these checks, so hookſ.json (long s) and .huſky are denied like the plain spellings the filesystem treats them as. The main thread can never edit .anti-hall/edit-allow.json itself. Subagents are unaffected. Doctor reports untrusted, changed or ignored entries. New setting guards.projectEditAllow (default true). Codex registers no Edit-family hooks, so this is Claude-only. The Codex settings skill documents the trust command.
  • inbox tick --quiet (peer ask, downstream Primaries 2026-09-26). inbox tick's JSON carries duplicate legacy+new field names (unread/unreadTotal, cursor/cursorNdjson, storeCursor/cursorStore), which made a cron-prompt directive eyeballing raw JSON error-prone. --quiet prints one line: tick <id>: unread N, known true|false, meshGap true|false, watcherArmed true|false on success, or a loud ok:false ... line + non-zero exit on failure. The JSON default (no --quiet) is unchanged. The DevSwarm wake-cron prompt now uses --quiet and words its stop condition against that line.
  • devswarm.js send takes several recipients. --to <id1>,<id2>[,…] or a repeated --to sends the same body to each recipient (duplicates dropped), so a caller no longer needs a shell for loop, which command-guard treats as heavy. Every recipient is attempted even after one fails. The result lists per-recipient ok/seq/bytes, and the exit is non-zero if any recipient failed. --quiet prints one line per recipient. A list does not combine with --broadcast, --to-primary or --cc-primary. Documented in the devswarm skill (Claude and Codex) and help. New setting devswarm.sendMultiRecipient (default true; false keeps "last --to wins").

Fixes

  • DevSwarm wake-cron default moved off :00/:30 (peer ask, downstream Primaries 2026-09-26). WAKE_CRON_DEFAULT was */30 * * * *, which fires exactly on :00/:30 — every other machine's */N cron piles onto the same wall-clock instant. Now 7,37 * * * *: same 30-minute cadence, off-minute offset. The MAILBOX WAKE CronCreate directive also now states plainly that ANY existing job already running inbox tick for the workspace — any schedule, id, or partition UUID — counts as present; never create a second one.
  • handover-resume.js GUIDED RESUME PATH is now adaptive. It used to unconditionally tell the agent to run a "section-10 resume-verification checklist" and read state.md/decisions.md/trials.md — a handover written without that section or those files left nothing to follow. It now detects (fail-open) what the referenced HANDOVER actually has and only mentions what exists, falling back to a generic 3-step check (git status --short --branch, pwd, CLAUDE.md/AGENTS.md re-read) when the checklist section is absent — still requiring the resume-verified: line either way.
  • model-routing-guard: deploys, migrations and secret work are never pushed to haiku. An opus spawn that ran a production deploy (wrangler … cors set, deploy_webui.sh prod) was blocked with "respawn with haiku", which is the wrong advice for deploys, migrations, rollbacks and secret/credential work. A spawn is deploy-shaped when it has one action signal: deploy, migrate, rollback, token rotation, or wrangler/terraform/kubectl apply/ firebase deploy/db migrate. Two distinct context words out of prod, secret and credential also count; one stray "no secrets" does not. At or above a floor, the rows that push toward haiku are skipped and every other row still runs. Below the floor, or with no explicit model, the spawn gets an advisory naming the floor instead. It is never blocked. New setting guards.modelRoutingDeployFloor (sonnet default, opus, or off for the old table).

  • command-guard: a read-only --check run of a real script file is allowed again. 0.111 refused every interpreter verb on the check-flag path, so python3 tools/gen_contract.py --check | tail -5 blocked. The main thread may now run <python*|node|ruby|perl|php> <existing script file> --check|--dry-run|--list piped to a bounded sink. The script must be an existing regular file named right after the interpreter. Inline code (-c/-e/-m/-p/--eval/--require), stdin or heredoc scripts, $/backtick/process substitution, env-assignment prefixes and wrapper verbs (sh/bash/eval/exec/xargs/env/nice) never qualify, and the remaining arguments must be non-heavy. An anti-hall script (this plugin, or any other anti-hall install found by its plugin.json) never qualifies, and neither does any command carrying --confirmed, so the main thread cannot flip a safety switch or trust an allowlist this way. Claude and Codex share the hook. New setting guards.allowReadOnlyVerifyScripts (default true).

  • devswarm.js spawn refuses a flag value that is really the next option. spawn b -s main -t -p "brief" made -p the workspace title and dropped the brief without a word. Spawn now refuses before fetching or creating anything, and the message names the flag. It checks -s/--source, -a/--agent, -t/--title or -p/--prompt with no value, or with a value that starts with -. For -p, only a single option-shaped token counts, so a brief that opens with a - bullet is still accepted. New setting devswarm.spawnStrictFlagValues (default true).

  • devswarm-parent-gate.js: the fresh-mail grace window (parentGateNeglectGraceMin, 0.111.0) now covers mesh-direct send --to messages, not just the native NDJSON inbox. A Primary that ran devswarm.js send --to <child> and ended its turn seconds later still got hard-blocked with DEVSWARM NEGLECT (peer field report), because the grace window's own scoping blanket-excluded every STORE-ONLY-union row (a mesh-direct send is store-only, no NDJSON line at all) regardless of who sent it. The exclusion now keys on the row's sender: a store-only row this Primary itself sent, still within the grace window, is downgraded to a stderr advisory — "N unread — awaiting child pickup (Ns)" — exactly like a fresh native-inbox send; a row from anyone else (or with no resolvable sender) still blocks unconditionally, unchanged. The pre-existing busy-child downgrade already applied to store rows and needed no change. New tests in tests/hooks/devswarm-parent-gate-neglect-grace.test.js cover a fresh own mesh-direct send (advisory), an old own send (still blocks), a fresh send plus an unanswered child question (still blocks), a fresh third-party mesh-direct row (still blocks), and an unresolvable sender (fail-open, still blocks).
  • devswarm-parent-inbox.js's "CHILD NOT DRAINING" per-turn nag: the 120s grace window (unreadIsGraced) was defeated by a busy child's own heartbeat. A field report showed the exact scenario the grace window exists to fix still reproducing verbatim: a 9-second-old own mesh-direct send to an actively-working (uuid-addressed) child, whose heartbeat was 1 second old, still produced "CHILD NOT DRAINING" and blocked the next Stop. Root cause: unreadIsGraced disqualified grace the instant the child recorded ANY heartbeat after the send — but a heartbeat file is rewritten on every inbox tick/turn cycle regardless of whether the specific message was ever read, so a busy, continuously heartbeating child (precisely who this window protects) almost always fails that check within its very first tick. The heartbeat check is removed entirely and replaced with the same sender-keyed predicate used in the Stop-gate fix above: grace applies while the message is within the window UNLESS its sender is positively known to be someone other than this Primary (an unresolvable/legacy sender still gets grace, matching the pre-fix lenient default). devswarm-store.js's summary projection now additively emits oldestDirectUnreadSender (the sender of the oldest unread row, computed alongside oldestDirectUnreadTs at zero extra cost) so the hook can make this call without a second store read. New/updated tests in tests/hooks/devswarm-parent-inbox-grace-window.test.js reproduce the exact field shape (fresh own send, fresh child heartbeat -> still graced) and cover a known third-party sender (never graced, however fresh).

Security

Hardening for the new 0.112.0 command-guard/edit-guard/model-routing surfaces above, found and fixed before release (F1–F4):

  • F1 — command-guard script --check: an anti-hall script, or any command carrying --confirmed, never qualifies for the read-only allowance, closing a path where the main thread could otherwise flip a safety switch or self-trust an allowlist through the new script-check exemption.
  • F2 — edit-guard project allowlist: path segments are folded with NFKC + lowercase before the deny-list check, so unicode-confusable spellings (hookſ.json long-s, .huſky) are denied exactly like their plain-ASCII equivalents.
  • F3 — command-guard read-only verify: a comment or an unbounded check no longer passes. x --check #| tail -5 was allowed because the comment hid the sink from the shell, and so was x --check && x --check | tail, where the first check is unbounded. An unquoted # now disqualifies the line, as does a background &; every pipeline that runs a check must end in a bounded sink.
  • F4 — model-routing deploy floor: weak single-word context no longer counts as deploy-shaped. One stray "no secrets" in an otherwise mechanical opus task used to count toward the floor; it now takes two distinct context words (out of prod, secret, credential) to qualify, keeping the original BLOCK for the weak case while still letting genuinely deploy-shaped spawns through.

Fixed

  • mcp-reaper: GRACE=0 honored via a deterministic parseGrace(), not a wall-clock e2e bound. The e2e test asserted elapsed time < 2500ms for a real subprocess run (spawnSync + two real ps scans) to prove GRACE=0 wasn't coerced to the default 3 by a 0 || 3 gotcha; that bound is wall-clock/scheduler-dependent and flaked under heavy local CPU load. The grace-parsing logic is now a pure, exported parseGrace() (used in production at mcp-reaper.js:569) covered by exact-equality unit tests (no subprocess, no timing); the e2e test keeps only a generous-deadline poll for the orphan's actual death.
  • "Restart" advice no longer implies harness registration that hasn't happened. Two sites could tell a user to just reload/restart to pick up a newer anti-hall build even when the Claude Code harness (installed_plugins.json) had not actually re-registered it yet — restarting alone does not load a version the harness has never seen (claude plugin update anti-hall@anti-hall is required first; see doctor.js's own harness-registration check). Field repro (2026-09-26): the marketplace clone had fast-forwarded to 0.111.0 while installed_plugins.json still reported 0.110.0 and no cache dir existed for 0.111.0 — devswarm-wake-watch.js's "update available" line said the version "is registered" when it was only known from the marketplace clone.
  • plugins/anti-hall/companion/lib/devswarm-wake-watch.js: checkStaleVersion() now returns a registered flag (true only when installed_plugins.json itself already names the newest version or newer); formatUpdateAvailableLine() branches wording accordingly — pointing at /anti-hall:update instead of claiming "is registered" when the harness has not caught up.
  • plugins/anti-hall/hooks/version-alert.js: CASE 2 ("already downloaded ... run /reload-plugins") now checks the harness's own installed_plugins.json (read-only, reusing update.js's resolver) before claiming a reload/restart is enough; when the harness lags the mirrored cache dir it instead tells the user to run /anti-hall:update first.

0.111.0 (2026-09-26)

Security

A security review of the three new main-thread command-guard allowances (read-only verify, per-project allowlist, plain push) reproduced nine bypasses; each is fixed with regression tests (tests/hooks/command-guard-security.test.js), and every reproduced case blocks again (or more strictly than 0.110.0).

  • Splitter: an unquoted backslash is an escape. git commit -m \" ; npm test ; echo \" hid npm test inside a fake quoted argument (bash reads \" as a literal quote), and every allowance inherited the gap. The shared splitter and its quote scanners now treat \x outside quotes as a literal pair. No shell-scan differential or command-guard corpus verdict changed.
  • Read-only verify: a check flag never launders a wrapped payload. sh -c "npm test" --check | tail, bash -lc '…' --dry-run, eval "…" --check and node -e "…execSync('npm test')" --check were allowed. The check-flag path refuses shell/interpreter/wrapper verbs and -c/-e/--eval/-lc, and the segment minus the flag must be non-heavy under the full unwrapping classifier.
  • Project allowlist: no shell expansion in a matched command. npm run deploy -- "$(npm${IFS}test|sh)" matched ^npm run deploy -- \S+$. Any $, backtick, backslash or <(/>( anywhere (quoted or not) disqualifies the command.
  • Project allowlist: ^.*$ is not an anchored rule. A pattern now needs a literal command word after ^, a closing $, no unbounded wildcard (.*, .+, [^;]*, [\s\S]+, quantified wide groups) and no top-level |; doctor reports each ignored pattern with its reason (its allowlist report also never printed before — a TDZ read of cwd, fixed).
  • Project allowlist requires per-user trust. A cloned repo's working-tree .anti-hall/command-allow.json could authorize itself. It now applies only while ~/.anti-hall/trusted-command-allow.json maps the repo's real path to the sha256 of the file bytes (any edit → untrusted); symlinked files/dirs are refused. New verb: node scripts/settings.js trust-command-allow [<repo>] --confirmed (prints the patterns; records nothing without --confirmed). doctor reports untrusted/changed allowlists with that command.
  • Plain push: the remote must be a configured remote name. git push ../other-repo main and git push host/evil main qualified; the remote must be one of git remote (fail closed).
  • Allowlist audit log: no symlink follow, secrets redacted. The log dirs are lstat-checked, the file is opened O_NOFOLLOW (mode 600), and the logged command is redacted (secret-named flag values, then the existing scrubSecrets).
  • Verify redirect/clone targets are resolved first. | tail > …/scratchpad/../../etc/x passed a raw substring test. Targets are resolved and realpath'd and must land inside the session's own scratchpad or a tmp root (helpers shared with edit-guard via hooks/lib/scratchpad.js).
  • Verify clone: only git clone --depth 1 https://… <tmp dest>. The local-path clone form and non-https --depth 1 sources no longer qualify.

Features

  • command-guard per-project command allowlist (owner-approved 2026-09-26). A repo may declare its own sanctioned exact commands — e.g. a deploy script the project's own rule says must never be delegated to a subagent — in <repo-toplevel>/.anti-hall/command-allow.json ({"patterns":["^anchored regex$", ...]}). Applies in the MAIN THREAD ONLY (a subagent never reaches this carve-out; it already passes through command-guard before this point). Every pattern must be literally anchored (^...$) or it is ignored, never matched. The WHOLE command must be exactly one unbroken segment (no chaining, pipes, subshells, backticks, or $( ) command substitution — reuses the guard's own splitSegmentsDetailed, no new parser) and carry no unquoted redirect anywhere, or it never qualifies regardless of the pattern. A qualifying match writes one audit line to ~/.anti-hall/logs/command-allow.ndjson. Default config is empty (no behavior change for a repo that never opted in). New setting guards.projectCommandAllow (default on). doctor reports an unanchored/invalid pattern in a repo's own config as a warning.
  • command-guard "allow plain push" (owner-approved 2026-09-26). In the MAIN THREAD ONLY, git add/git commit/a plain git push [remote] [ref], and &&/; chains made up only of those three, run inline instead of being delegated. ref must be omitted, HEAD, or the current branch (resolved fresh via git symbolic-ref --short HEAD, fail-closed if unresolvable). --force/-f/--force-with-lease/--force-if-includes/ --mirror/--delete/-d/--all/--tags/+refspec/src:dst to another branch, any other chained segment, and pipes/redirects/subshells stay exactly as blocked as before. git-guard.js keeps its own independent force-push/AI-credit checks, untouched. New setting guards.allowPlainPush (default on).

Changed

  • Roster wording: "no upstream" no longer reads as an error. A workspace spawn creates a branch with no upstream until it is pushed — normal, not a problem — but the Primary's per-turn workspace table (devswarm-parent-inbox.js's riskMarker()) rendered it with a warning glyph, ⚠ no upstream. Now renders local only (not pushed), no glyph. The underlying noUpstream field and its priority over a stale unpushed count are unchanged.

  • command-guard "narrow allow": bounded read-only verification for the coordinator. The coordinator may now run a short, single-target, read-only verification command inline instead of delegating it — e.g. re-running one delegated test file to verify a subagent's "done" claim — when it is a --check/--dry-run/--list flag, a -fsyntax-only compile check, one python3 -m pytest -q <file>, one or two explicit node --test <files>, ctest -R <name>, or a scratchpad-scoped git clone, AND its output is piped to tail/head/grep -c/grep -m N/wc, AND no write redirect targets a path outside the scratchpad/tmp, AND every other segment on the line is trivially safe (cd/pwd/true) or the piped sink itself — any other segment (including a chained second heavy command, or a --check hidden inside a still-heavy invocation) keeps the whole line blocked. Full suites, builds, installs, deploys and pushes stay gated. New setting guards.allowReadOnlyVerify (default on).

  • devswarm.js relay <seq|receipt> --to <id> [--note-file <path>]. Forwards a message the caller already received (its own inbox) to another workspace, verbatim, prefixed with a provenance header (relayed from X, seq N, M bytes). <receipt> (an inbox read-primary readReceiptId) resolves only when it covers exactly one message; otherwise it refuses ambiguous rather than guessing. Verifies the relayed byte length against the source and refuses (ok:false) on a mismatch or an empty source body — never a silent partial relay.
  • devswarm.js inbox read-primary --format text. Prints one from/seq/body block per message instead of the raw JSON. The default two-step read-then-ack-primary --receipt flow is unchanged; the new, opt-in --ack-after-print flag acks immediately after printing instead.
  • devswarm.js send --quiet. Prints one line (sent seq N -> X, B bytes, ok) instead of the full JSON; failure still prints a loud ok:false ... line and keeps the non-zero exit code.
  • devswarm.js send --to <id> --cc-primary. A direct --to send also copies the Primary with the identical message body, best-effort — reported under the result's ccPrimary, never flips the primary send's own ok/exit code.
  • spawn help text now documents the real hivecontrol args. devswarm.js help spawn names every hivecontrol workspace create flag (-s/--source, -a/--agent, -p/--prompt, -r/--remote, -t/--title) plus anti-hall's own --from-local, with an example and pointers to send/roster/inbox for following up with a spawned child.
  • devswarm.js help --short. Peer request: a DevSwarm Primary spent a day driving raw hivecontrol because it never discovered devswarm.js archive existed. Prints one line per verb (verb — purpose), generated from the SAME source of truth as the full help/help <verb> listing (the dispatcher's own verb list + VERB_HELP), so it can never drift out of sync. A hygiene test independently scans the dispatcher's case '<verb>': statements and asserts every one appears in the short list. The Primary's SessionStart directive now names it as the pointer to the full verb list.
  • devswarm.js ready-check <sha>. A generic, read-only readiness verdict for a child's "READY \<sha>" claim — works against any git repo, not DevSwarm-specific. Reports ff (is --base, default origin/main, an ancestor of sha), the base...sha file diff (files), submodule pointer bumps (gitlinks), deletions under --watch-deletions dirs (deletions_under), files outside --allow globs (outside_allowed), and a verdict:'ok'|'review'|'block' + reasons[]. Runs git read-only; no fetch unless --fetch is passed.
  • devswarm.js inbox tick reports watcherArmed. Whether a live Monitor wake-watch currently covers this workspace (fresh lock, pid alive) — the Claude-branch CronCreate prompt body now checks this field and tells the agent to re-arm Monitor when it reads false (the harness caps a Monitor at 30 minutes; this cron's own 30-minute fallback cadence is exactly when it would have lapsed).
  • devswarm.js roster/the parent gate now name what a child is waiting on. A row (or block reason) whose transcript is paused on an unresolved AskUserQuestion/ExitPlanMode now carries a truncated (~120 char) preview of the actual question/plan text, not just "waiting on a human" — reusing the SAME childBusyState detector both surfaces already relied on.
  • doctor reports live processes leaked into archived/gone DevSwarm workspaces. A bounded (≤2s), report-only scan cross-references every live process's cwd against archived (anti-hall's own marker, or the DevSwarm app DB's) or gone (worktree removed) workspace paths and prints pid, command name, cwd, and a suggested manual kill — never kills anything itself.

Fixes

  • DevSwarm store migration's "SOME COUNTS UNVERIFIED" no longer false-flags every workspace that has ANY native mesh traffic. migrateOne's count-verify used to require the store's TOTAL message count for a workspace to EQUAL today's legacy-inbox line count — but a workspace's store also accumulates rows from native mesh writes that have nothing to do with the legacy inbox, so that equality permanently breaks the instant any such row exists (not a data-integrity signal; confirmed on a live store where 100% of an actively-used project's workspaces failed the old check despite complete data). Verification now checks legacy-line COVERAGE (every legacy line is represented in the store, freshly imported or already covered by another path's row) via the same cross-path identity migrateLegacyInbox already used correctly. migrate-state.js's CLI output now names each still-unverified workspace (title/id + the specific field that failed), capped at 10 with a "...and K more", and states plainly that sources are never deleted so no data can be lost — re-running the (idempotent) migration re-verifies.
  • silent-agent-nudge.js could block Stop with no subagent involved at all. ~/.anti-hall/agents/ is a single home-scoped directory shared by every project/session on the machine, but heartbeatCandidates() globbed EVERY *.json in it and treated each as a subagent heartbeat, falling back to the filename as id when one was missing. That misread phase-tracker.js's rolling recent-spawn.json ({ts} only, a 20-minute "orchestration live" marker, not an agent) and other projects'/workspaces' devswarm-<branch>.json files as silently-dead agents. A heartbeat candidate now requires its OWN id and status fields (never the filename) plus a session field that matches the session about to Stop; recent-spawn.json and devswarm-*.json are also excluded by name as defense-in-depth. The heartbeat convention (skills/orchestration/SKILL.md) gained a required session field (session_id of the spawning session); a heartbeat written before this field existed has no owner to verify and is now treated as not-ours (fail-open toward no nudge, never toward blocking on an unverified file).
  • The parent Stop gate no longer force-blocks a Primary seconds after it sends a child a message, before the child has had any chance to read it. New setting devswarm.parentGateNeglectGraceMin (default 1 minute): a plain, native-inbox unread backlog younger than this never counts as neglect by itself, independent of whether the child is separately busy. Never applies to a store-only (mesh-direct send, or dead/foreign-descriptor) row, and never suppresses an unanswered child question or a corroborated stale/escalated verdict.
  • task-guard's IDLE NEGLECT check now recognizes an explicit owner-blocked marker instead of forcing a fake blockedBy dependency to silence it. A task with metadata.blockedOn (or top-level blockedOn) === 'owner' / 'user' / 'human' (case-insensitive), or a subject starting with "OWNER:" / "OWNER DECISION" (case-insensitive), is excluded from the ACTIONABLE-NOW set and never nags. New setting guards.taskGuardOwnerBlockedMarker (default on). Documented in docs/TASK-WORK.md, docs/KB.md, and the nudge text itself.
  • Static per-turn reminder blocks (VERIFY-FIRST, the DEVSWARM PRIMARY dispatch-tier and top-fan-out-tier suffixes) no longer repeat every single turn. They now follow the same once-per-session / once-after-compact-or- clear / once-every-N-turns cadence as the existing DevSwarm workspace-table dedupe, via lib/emit-dedupe.js's keepalive rule. New setting guards.injectionRepeatEvery (default 10 turns; 0 restores every-turn injection). task-tracker.js's own FULL/SHORT + freshness-note logic and devswarm-child-turn.js's COMMS OVERRIDE reassertion are unchanged (the latter is deliberately per-turn — DevSwarm's --system-prompt-file erases the child's system prompt at every spawn, so per-turn reassertion is the only lever against that erasure).
  • A Codex quota/rate-limit exhaustion is now recorded once and shared, instead of every lane rediscovering it independently. New PostToolUse hook codex-quota-detect.js (matcher Agent) detects a quota-exhaustion message in a codex:codex-rescue Agent result and records {available:false, until, reason} into ~/.anti-hall/codex-availability.json (lib/codex-quota.js, merged with the existing PATH-probe fields, never clobbering them). The SessionStart codex-availability.js hook now also surfaces a live outage even when the PATH probe alone found nothing to report, with the routing fallback ("Codex unavailable until X; route correctness review to Sonnet"). New setting guards.codexQuotaDetect (default on). The exact Codex CLI quota message wording was not found verified anywhere in this repo or machine at authoring time, so detection matches conservatively on quota/rate-limit exhaustion vocabulary rather than one fixed string.
  • Nudge-class Stop hooks now take a per-signature session ack, instead of re-blocking on the same confirmed-false condition. silent-agent-nudge.js and tasklist-guard.js had no user-triggered ack at all (only automatic same-snapshot dedup); devswarm-parent-gate.js already had one (its own intents/intentAcks state, driven by devswarm.js gate-intent --reason) but it is deeply coupled to that gate's own escalation shape, so it was left untouched rather than force-generalized. New shared hooks/lib/stop-ack.js gives the two gap hooks a documented skip-file entry (~/.anti-hall/stop-ack/<session>.json, keyed "<hook>:<signature>") the agent writes once the user has explicitly confirmed a condition is a false positive — that exact signature then stays advisory (never blocks again) for the rest of the session; a genuinely changed condition is a new signature and blocks normally. New setting guards.stopAck (default on).
  • A stale, already-fixed nudge no longer keeps blocking after claude plugin update has re-registered a newer build. installed_plugins.json (harness-owned) can be re-registered at a newer version while the CURRENT session's hooks keep executing the OLD build until a full restart — /reload-plugins does not pick this up (doctor.js's own harness- registration check; claude plugin update --help documents "restart required to apply"). New shared hooks/lib/stop-version-gate.js (reusing skills/update/scripts/update.js's own version-resolution exports, never reimplementing installed_plugins.json parsing) lets silent-agent- nudge.js, tasklist-guard.js, and devswarm-parent-gate.js's plain NEGLECT nag detect this and downgrade their block to advisory until restart. Deliberately NOT applied to devswarm-parent-gate.js's unanswered-question / truncation / escalation paths, which bypass the cap unconditionally by design, nor to any safety guard (command-guard/edit- guard/git-guard stay out of scope). Prospective only: a hook build that predates this file has no way to run the check. New setting guards.stopHookVersionDowngrade (default on).
  • Investigated (peer complaint #3, batching/re-fire): how Claude Code combines several Stop hooks that each return decision:block in one turn (one continuation vs. several) is not documented in this repo's own docs/KB-claude-code-hooks.md, and no existing test in this repo answers it either — recorded as unverified rather than guessed at. What WAS verified: silent-agent-nudge.js already had a per-exact-snapshot dedup plus a hard once-per-agent-per-session cap; tasklist-guard.js already short-circuits on hash === lastHash before blocking again. Both already satisfied "no re-block for an identical snapshot within a session" before this change — no fix was needed there. devswarm-parent-gate.js's own stable-kind cap (hooks/lib/stop-policy.js) also already resets only on an observed-clear condition, not on content churn.
  • DevSwarm directive text (mailbox wake cron, Monitor re-arm, comms override, drain nudge) now names a version-independent launcher instead of a version-pinned plugin-cache path. Every printed node <path> command was baked from the CURRENTLY RUNNING hook's own __dirname — a path like ~/.claude/plugins/cache/anti-hall/anti-hall/0.109.1/scripts/devswarm.js. That path was correct the instant it was printed, but crons, Monitors, and handovers keep the literal text around across releases, so after the next update the printed command pointed at an old (sometimes deleted) version directory and showed a stale version number in the text itself — forcing a manual recreate of every cron/Monitor and a handover edit on every release (peer report: downstream Primary). Fixed by installing two tiny, self- contained launchers under ~/.anti-hall/bin/ (devswarm.js, wake-watch.js) that resolve the CURRENTLY REGISTERED anti-hall install (installed_plugins.json -> the marketplace clone -> the path baked in at generation time) EVERY TIME THEY RUN and delegate to it with full argv/exit-code passthrough; devswarm-child-role.js, devswarm-parent- gate.js, devswarm-child-gate.js, and devswarm-child-drain.js now embed those stable paths in their directive text instead of the raw version- pinned one. Idempotent (only rewrites the launcher when its content actually changes) and fail-open throughout (an install failure, or devswarm.stableLauncher = false, falls straight back to the previous version-pinned path — byte-identical to pre-fix behavior).

0.110.0 (2026-09-26)

Features

  • Bug history for the defect channel. defect.js has three new maintainer verbs:
  • backfill [--repo <path>] [--dry-run] imports every fix: / fix(scope): commit from git history as a fixed record. The component comes from the source file with the most changed lines (tests and docs are ignored). fixedIn is the earliest release tag that contains the commit. The cause comes from keyword rules. The matching CHANGELOG bullet is linked when found. Records are keyed by commit sha, so a re-run adds nothing. They are stored in ~/.anti-hall/defects/history/ and never appear in list --open or the defect nudge.
  • recurring [--since <version|date>] [--top N] [--json] groups reported and imported fixes by component and by cause. It flags hotspots (a component fixed 3+ times, or the same component and cause 2+ times) and likely regressions (the same component and cause fixed again within 5 releases, or an explicit regressionOf).
  • similar <text…> [--component X] lists the 10 past fixes closest to a new bug.
  • report and rule accept optional --component, --cause (a fixed 12-class root-cause list) and --regression-of <fp>. Existing records without these fields still load unchanged.
  • The root-cause skill (Claude and Codex) now runs defect.js similar before an anti-hall fix. If the component is a hotspot, fix the class of bug, not just this instance. The defects skill documents the new verbs on both ports.
  • The command guard exempts the two new read-only verbs, defect.js recurring and defect.js similar, alongside report/list/show, so the root-cause skill's similar step runs on the main thread. backfill writes history records and stays gated.

Changed

  • One row-eligibility projection (companion/lib/row-eligibility.js). Every workspace row is now judged once for archived (anti-hall marker, DevSwarm app DB, active-list absence, with archivedBy provenance), held (devswarm.heldPartitions), ignored (archive-ignore marker) and live/busy/waiting-on-user. It has a memoized per-invocation context and a batch API. The parent Stop gate, the per-turn parent-inbox table (plus its archive-ready nudge and stale-registry filter), and the CLI's routing, roster and diagnose read it instead of combining the predicates one axis at a time. There is no behavior change.
  • Hygiene ratchet (tests/hygiene/archived-predicates-single-projection.test.js). A direct call to an archived/held predicate outside the projection and the helpers it wraps now fails CI. The remaining partition-level and app-DB write-guard call sites are allowlisted with exact counts and reasons.
  • One lock primitive. New companion/lib/lock.js replaces 12 hand-written cross-process locks: swarm-guard, repair-on-reload, settings, recovery per-id, supervisor sweep, ingest (plus its orphan sweep and legacy probe), migrate, pull/wake-watch, the store journal, log rotation and retention. It keeps recovery.js's design: write-then-link publish, with an O_EXCL fallback on filesystems without hard links (SMB/exFAT), atomic rename-aside reclaim, and token-checked release. It adds a host-scoped owner record, an mtime torn-read guard, and per-caller steal policy. Each caller keeps its timeouts, stale windows, retry budgets and fail-open/fail-closed contract.
  • A hygiene ratchet (tests/hygiene/lock-single-primitive.test.js) forbids new O_EXCL, linkSync or lock-unlink code outside lock.js.

Fixes

  • Lock reclaim race. Two processes that judged the same dead or stale lock holder both deleted it. The second deletion removed the first process's fresh lock, so both ran the critical section. Affected: supervisor sweep, migrate, pull/wake-watch, retention, log rotation, store journal (duplicate dedupe rows), ingest (two monitor consumers), settings (lost key), swarm-guard (spawns past the cap) and repair-on-reload (two doctor --repair spawns).
  • Lock torn-read steal. An empty lock file is normal while its live holder is still writing it. The supervisor sweep, retention, settings and repair-on-reload locks read that empty file as ownerless and stole it.
  • Lock blind release. The settings lock released by deleting the file without checking the owner token, so it could delete a successor's lock. It now checks the token.
  • Ingest orphan-sweep race. The ingest orphan sweep and legacy-lock probe re-read a lock and then deleted it. A fresh lock published between those two steps was deleted. Both now reclaim atomically.
  • Lock three-way reclaim race. Two processes could judge the same stale holder. A reclaimed and published, then B renamed A's live lock aside and C published into the empty path. B's restore then failed, so B deleted A's lock and A and C both held it. Reclaimers are now serialized by an O_EXCL <lock>.reclaim sidecar. It uses the same publish path, so it works without hard links, and it goes stale after 5 s or when its pid is dead. Under the sidecar the holder is re-read and re-judged before any rename.
  • Lock hostname drift. macOS renames the host when the network changes, so a live local lock holder was treated as belonging to an unknown machine. Migrate and pull could then steal it once it was stale by age, and recovery and the supervisor lost the immediate dead-holder reclaim. Lock records now carry the boot time and, on Linux, the pid namespace. A holder with a different hostname but a matching boot time and namespace is on this machine, so its pid is checked. Holders on other machines are still judged by age only.
  • Lock scratch sweep covers every lock directory. doctor --repair removed leftover lock scratch files only from devswarm/locks/. The shared lock primitive also leaves *.lock.tmp-*, *.lock.reap-* and heartbeat .hb. temp files in ~/.anti-hall/, logs/ and the store journal directories. The sweep now covers all of these plus the .reclaim sidecar. It removes only lock scratch files older than 15 minutes and lists every file it removed.
  • Handover's terminal line is now trigger-aware, not an unconditional compact signal. HANDOVER.md must record a Trigger: line (auto-threshold / user-request / restart-pending / task-boundary, plus context % from the statusline or ~/.anti-hall/auto-handover/<tag>.json's firedPct). The terminal line only claims "🟢 HANDOVER COMPLETE — GOOD POINT TO /compact NOW" for auto-threshold or an explicit compact/clear request; any other trigger below the threshold gets a neutral "📝 Handover saved (proactive, ...): no need to compact now" line instead. Fixes a field incident where a proactive handover below the 85% auto threshold was misread as a signal to stop/compact. Mirrored in the Codex port (plugins/anti-hall/codex/skills/anti-hall-handover/SKILL.md).
  • shell-scan-differential.test.js no longer misreads a timed-out/killed guard spawn as ALLOW. spawnGuard compared spawnSync's res.status directly against 2 (BLOCK); under load, a child killed by the 10s timeout or a signal returns status: null, which the bare comparison counted as ALLOW, producing false "SAFETY REGRESSION base=BLOCK new=ALLOW" failures. spawnGuard now retries a null status up to 2 more times with a longer (30s) timeout, and classification is a small pure function (classifyGuardResult) that returns 'INCONCLUSIVE' — never 'ALLOW' or 'BLOCK' — for a still-null result after retries; the test then fails with an explicit "INCONCLUSIVE (timeout/killed) — not a verdict" message instead of silently misclassifying. A genuine exit-0-where-base-blocked regression is unaffected — it returns a real status on the first attempt.
  • Canonical resolveHome() guard added for shared home-fallback helpers. os.homedir() is called directly at ~300+ sites across the plugin; a test that forgets to pass an explicit home/env can silently fall through to the REAL developer machine home instead of an isolated fixture (the recurring "tests never touch the real home" defect class). Added companion/lib/test-home-guard.js#resolveHome(explicitHome, env) — the same explicitHome || os.homedir() fallback in production, but it refuses (throws) under node --test (or explicit ANTIHALL_TEST=1) when that fallback would resolve to the real user home, with an ANTIHALL_ALLOW_REAL_HOME_TEST=1 escape hatch for a test that deliberately needs it. Migrated hooks/lib/settings.js (homeDir/homeFromEnv) and hooks/lib/jev-assist.js (homeDir) — the two shared helpers most other hooks/companion code routes settings/jev config reads through — to use it. Added tests/hygiene/homedir-call-site-ratchet.test.js, which counts remaining direct os.homedir() call sites and fails if the count grows, so new code can't reintroduce the gap while the rest of the call sites are migrated incrementally.

0.109.5

Fixes

  • DevSwarm MAILBOX WAKE directive now names the right id. The SessionStart and Stop-gate wake directives were telling the agent to poll an unregistered id (the raw session/builder id, or — in a repo whose Primary seat had never been registered — an id with no history at all), so the inbox check the directive told the agent to run always came back "unregistered" even though wake-watch itself had armed correctly on the real, resolved Primary id. Fixed by threading the resolved id through both directives and by registering a Primary seat on its very first use, not only when it is adopted from a previously closed one. First-use registration happens only under real DevSwarm (DEVSWARM_REPO_ID set) — forcing devswarm.supervisorMode=on alone never registers a seat.
  • A corrupt Primary seat descriptor is reported, not overwritten. An unparseable workspaces/<id>.json now reads as seat state unknown with a SessionStart warning naming the file, instead of being re-registered over.
  • DevSwarm wake-watch no longer needs a manual re-arm after every release. When a newer build's watcher is available, the running watcher now hands the stream off to it automatically instead of printing a re-arm instruction and stopping — Monitor keeps streaming through the handoff. Guarded so a build can only hand off to a newer version, and only once per watch chain, with a safe fallback to the old print-and-stop behavior if the handoff itself fails.
  • Wake-watch handoff hardening. After handing off, the old watcher no longer overwrites the new one's seen-state on exit, exits with the new watcher's own code (128+signal when it was killed), and forwards SIGTERM / SIGINT to it so the new watcher is never left orphaned holding the lock.
  • Edit guard now allows the harness plan file outside plan mode too. ~/.claude/plans/*.md is exempt from the delegation block unconditionally, not only while the session is in plan mode.
  • Tasklist guard no longer blocks Stop in plan mode. Plan mode cannot write the progress file the guard was demanding, which previously caused a stuck loop.
  • Speculation guard no longer flags "must be"/"should be" used as a requirement. Obligation phrasing in requirement or acceptance-criteria context (e.g. "the result must be idempotent") is exempted; real speculative claims are still caught. The exemption covers only true obligation verbs (measured/verified/tested/...): state claims like "should be done by now" or "should be deployed" are still flagged, every must-be/should-be in a reply is checked, and the requirement-label context counts only at the start of a line.
  • Command guard allowlists read-only/append-only defect commands. node scripts/defect.js report|list|show no longer requires approval; rule and archive remain gated.
  • Command guard's own-CLI allowlist is anchored to the segment start everywhere, not just for defect.js. jev-setup.js, settings.js, jev-report.js, doctor.js, phase.js, agent-watchdog.js, and devswarm.js previously matched anywhere in the segment, so a heavy command mentioning one as trailing args (npm run build -- node scripts/devswarm.js list) slipped through unblocked; every entry now goes through one shared anchoring helper so a future addition can't forget it.

Added

  • New setting: autoHandover.decisivePrompt (default true). At a turn-ending Stop point, once the current session's handover exists and is fresh, the agent is told to end its reply with one prominent line naming the exact /compact (or /clear, or Codex /new) command to run next — or, if the handover has gone stale since it was written, to refresh it first. Turn it off to revert to the plain handover wording.
  • Decisive prompt freshness uses the tasklist guard's own work detection. Subagent edits and Bash writes (cp, mv, rm, > redirects, ...) after the handover now mark it stale; with no readable transcript (including Codex) the line is a neutral "📝 Handover saved at ..." instead of 🟢; and "Next action" means done only when it reads exactly none/done/complete/nothing.

0.109.4

Fixes

  • DevSwarm Stop gate: archived workspaces no longer block the Primary. A workspace archived in the DevSwarm app, archived by anti-hall, or marked with archive-ignore no longer stops the Primary's turn, even if it still has unread mail. Its mail is also no longer counted toward the NEGLECT warning or escalation. At most one summary line is shown instead: "N archived workspace(s) still have unread mail (ignored)". This check also reuses the cached copy of the DevSwarm app's archived list.
  • Owner-held workspaces no longer block the Primary either. An id listed in the devswarm.heldPartitions setting is skipped by the Stop gate, even when it is an old twin of the Primary. The Primary's own current mailbox is never skipped.
  • "Waiting on a human answer" now requires a running session. A workspace is only reported as waiting on a human answer if its session is still running and its transcript shows the unanswered question. The gate used to report this for sessions that had already ended.
  • Active workspaces are unchanged. Unread mail with no transcript still blocks the Primary, as in 0.109.0.

0.109.3

Fixed

  • Silent-agent nudge no longer blocks every Stop for a finished agent. A background agent whose completion notice arrived as an attachment or queue-operation transcript entry was treated as still running and reported as "silent", and the Stop hook then blocked on every attempt. The nudge now reads all three completion shapes (reusing the idle gate's parser), treats an agent whose result was later delivered as finished, and hard-caps nudges at one per agent per session.

0.109.2

Fixes

  • Heartbeats reuse a cached copy of DevSwarm's archived-workspace list instead of opening the app database on every call.

0.109.1

Fixes

  • DevSwarm roster: a relaunched terminal can no longer revive an app-archived workspace. Field bug: two workspaces archived in the DevSwarm app (isActive=0, isHidden=1) had their terminal tabs left open; when the running claude process was killed, the tab's login shell relaunched claude within ~10s, and the new session's routine ensure + explicit register + heartbeat calls all flipped both rows back to active on the roster — one even broadcast "idle — awaiting task brief". Anti-hall's own resurrection guard only checks its OWN archived/<id>.json marker, which is never written when a workspace is archived from the DevSwarm app UI instead of anti-hall's own archive verb. register/ensure/ heartbeat now consult the DevSwarm app DB's own archived verdict first and refuse to reactivate a row it reports archived — "regardless of new heartbeats or registrations" (owner rule: the app DB is ground truth over anti-hall's own markers). heartbeat's base file write still always succeeds; only the liveness-verdict clear-to-alive and the mesh --summary broadcast are suppressed. The roster's archived hint gains a live session in archived workspace flag when a fresh heartbeat still exists for an archived row. If the app DB can't be read, behavior is unchanged (fail-open).
  • doctor reports (never kills) a live claude session left running in an app-archived workspace. New report-only check: claude session alive in an archived workspace: <title> (pid N), resolved from the app DB's own AI-terminal session id through the same harness session-file mapping liveness.js already used elsewhere. Never kills the process, never archives or deletes anything — it names the leak and the two safe ways to close it (close the DevSwarm tab, or hivecontrol workspace archive <full id>). One aggregated WARN, never one per workspace. (Auto-archive was already unaffected: gatherCandidates only considers isActive=1 builders, so an app-archived row was never an auto-archive candidate.)

0.109.0

Features

  • Post-handover new-work gate (on by default). Once context is past the auto-handover threshold and this session's handover has been written, the agent now sizes each new request before starting it. If the request would need more than about 5% of the context window, it offers you two choices: park it in the task list and the handover and start it after /compact or /clear, or go ahead anyway if you insist. Quick questions, finishing the task already in flight, and spawning a DevSwarm workspace pass straight through. Works the same on Claude Code and Codex.
  • One-shot budget reminder after a handover. If usage grows more than the budget past the point where the handover was saved, you get one reminder per handover: refresh the handover, and offer to park the rest of the work.
  • New settings: autoHandover.gateNewWork (boolean, default true) and autoHandover.gateBudgetPct (1-50, default 5), both also in /config.
  • New Jev integration postHandoverGate (advisory, default off). It logs whether Jev thinks a request fits in the remaining post-handover budget; it never changes what the gate says. It has its own jevIntegrations.postHandoverGate row, like the other Jev integrations. It ships off by default: an offline benchmark (n=299) found park-recall 17.6% vs 28.8% for the agent's own size judgment plus the measured budget backstop, no gain over the baseline.
  • Mechanical nudge on a silent background subagent. anti-hall already documented the heartbeat convention for a coordinator to notice and re-dispatch a stale subagent, but nothing forced it to happen and nothing wrote the heartbeat automatically. A new Stop hook (shared with the Codex port) instead watches the signal the harness always produces: every background agent launch and its eventual terminal notification. An agent counts as silent once its own output has gone stale (or never appeared) past the threshold with no terminal notification seen, and you get one nudge per stale snapshot — advisory only, it never stops or re-dispatches the agent itself. The ~/.anti-hall/agents/<id>.json heartbeat file, when a subagent does self-report it, is kept as an additional secondary signal.
  • New settings: guards.silentAgentNudge (boolean, default true) and guards.silentAgentNudgeMin (minutes, default 20), both also in /config.

Changed

  • DevSwarm auto-archive: a finished workspace's own status pings no longer keep it "active" forever. A done child workspace is woken periodically by its own mailbox cron, the Monitor watcher, and a Stop-hook heartbeat reminder — each wake used to reset the idle timer, so a finished child could never go idle and was never auto-archived. The idle timer now reads the child's own transcript turn by turn and skips only those wake/ping/heartbeat/status turns; any other activity (real work, read-only work included, a new direct message, or a commit) still resets the timer, and a turn that is still doing real work blocks the archive outright. If the transcript can't be read, the older, more conservative rule applies, so this change can only delay an archive, never make one happen early. Archives stay reversible, only by the exact workspace id, and never triggered from a hook.
  • New setting: devswarm.autoArchive.ignorePings (boolean, default true), also in /config.
  • devswarm.js spawn is faster and more deterministic. A cached workspace title from hivecontrol's raw branch-name default could clobber an already-confirmed, different title, giving a "sometimes branch, sometimes brief" spawn-title race — fixed. The roster/per-turn "finish" column now shows the actual done-rule state in plain words (done ✓ merged / done, merge unverified / done, not merged / working) instead of a raw gate-count ratio. Spawn now skips the redundant origin fetch when the remote-tracking ref is already fresh, fetches submodule updates on demand only when it does fetch, and puts a timeout on the underlying hivecontrol create call, reporting per-phase timings. It also now detects and reports (never auto-repairs) a submodule worktree that failed to create during an otherwise-successful workspace create.
  • New settings: devswarm.spawnFetchTtlSec (seconds, default 300) and devswarm.spawnCreateTimeoutMs (milliseconds, default 180000), both also in /config.

Fixes

  • The DevSwarm ingest daemon no longer freezes, dies silently, or reads healthy while a slow hivecontrol drains nothing. Under heavy machine load the monitor call timed out (spawnSync … ETIMEDOUT), but spawnSync's timeout only sends SIGTERM and then keeps blocking until the child actually exits. A hivecontrol that was slow to die froze the whole daemon, heartbeat included, for minutes. The daemon (companion/devswarm-ingest.js) now runs each monitor call with a non-blocking child_process.spawn: SIGTERM at the hard timeout, SIGKILL after a grace period, and a hard upper bound after that. While a call is in flight, a timer keeps the lock and liveness heartbeat fresh. The heartbeat is also written on the backoff paths that skipped it before (a blocked delivery WAL, a retryable store error), and it carries a new lastMonitorAttemptMs. Transient failures back off exponentially (2s, 4s, 8s …) up to a 5-minute cap instead of retrying every 2s. Every exit now leaves a reason in ~/.anti-hall/devswarm-ingest.log: SIGTERM/SIGINT/SIGHUP are logged (and the in-flight hivecontrol child is killed), as are uncaught exceptions and unhandled rejections (with the stack) and the final exit code. An exit with no line at all is now the signature of a SIGKILL. Health (monitorFaultFor / daemonHealth): when no monitor poll has succeeded since the daemon started, or since the last success, for longer than the new devswarm.monitorNoOkFailMin window (default 10 minutes), health reads FAILING even when no failures have been counted. That was the field case: lastMonitorOkMs stayed null with 0 failures for 15+ minutes and read as healthy. Inside that window a freshly started daemon reports "starting up" (startingUp: true, status still healthy, so no repair restarts it). The failing banner now says which condition tripped, the heartbeat age, and the last error in plain words.
  • New setting: devswarm.monitorNoOkFailMin (minutes, default 10), also in /config.
  • doctor --repair no longer reinstalls a healthy-but-slow ingest daemon, and the installer stopped churning launchctl on every workspace spawn. A monitor fault only fires once the daemon's base liveness signals already passed, so a merely slow/timed-out hivecontrol call (ETIMEDOUT, no resolvable spawn-error code) is a live daemon, not a broken one — doctor --repair used to reinstall/restart it anyway, which only interrupts an otherwise-healthy process. isMonitorConfigFault() now reinstalls only for a genuine config fault (ENOENT/EACCES/ENOTDIR); a slow/transient fault is reported via the new monitorSlowReason() and left alone — a genuinely dead daemon still reinstalls either way. Separately, install-devswarm-ingest.js (which runs on every DevSwarm workspace spawn) no longer unloads already-reaped legacy per-worktree launchctl units it never installed, and skips the unload+load of the live per-project unit entirely when the on-disk plist already matches what would be written and the label is already loaded.
  • jev report's triage rows leaked across --project/--by groups, and a settings write could slip past its own risky-change lock. jev-report.js filtered triageRows by time window only, never by the same groupKeyOf() used for every other row, so every project's or session's triage counts showed every OTHER project's/session's triage decisions mixed in; a project-less triage row now buckets under unknown like any other row instead of leaking everywhere. settings.js set()'s locked-key risky-change check read settings.json and decided before acquiring withSettingsLock — the same TOCTOU reset() already had fixed — so a concurrent writer could change settings.json between that stale read and this call's own write, letting a risky change through unconfirmed; the check now runs inside the same lock, reusing the same load.
  • The silent-background-agent nudge could stay silent, and the handover scan it triggers is no longer a full recursive walk. silent-agent-nudge.js's terminal-status match lacked a case-insensitive flag, so a differently-cased status was misread as still-silent; its task-id/status parsing was first-match-only against the whole text leaf, so when several agents finished together in one leaf only the first agent's id/status pair was ever read and every later agent stayed flagged as silent. Each <task-notification> block is now parsed separately, paired with its own id and status, and the displayed nudge is deduped by agent id across the transcript and heartbeat sources so an agent visible through both never produces two lines for the same thing. Separately, the handover lookup this nudge (and the post-handover gate) relies on used to do a full recursive walk of .anti-hall/handovers — every date dir × every session dir × every file — just to find one session's own handover; it's now bounded to <date>/<sessionId>/ per date dir via the new findNewestHandoverForSession(). findNewestHandover()'s cross-session fallback (used by handover-resume and the pre-compact snapshot) is unchanged.
  • The DevSwarm auto-archive idle gate could archive a workspace mid-work, or never fire at all. A safety review of the 0.109 idle gate found it treated every wake/ping turn as proof of idleness, when a false "ping" archives a child mid-work (an extra reset only delays, so every doubt now counts as real work): a background Agent/Bash launch with no matching final <task-notification> (completed/failed/stopped, matched by task id or tool-use id), or a turn_duration pending-background-agent count above zero, now blocks archive outright, and a ping turn never hides it. The mailbox-ping classifier first went too strict (a bare allowlisted node devswarm.js inbox|heartbeat|roster|mesh command, nothing else) — since real children pipe their pings through grep/head/tail/wc, that made the gate inert in practice, every wake turn counting as real work — so it now also accepts exactly one 2>&1 plus a chain of those read-only filters with plain-word arguments only; anything with a file redirect, <(, $(, a pipe reader like python3/ sed/awk/jq/tee/xargs, or a shell chain still counts as real work. Native mailbox-drained rows (no mtype) now count toward the inbound floor, and a cron-wake flag now only ever attaches to the very next meta prompt so a human prompt can't be misclassified as a wake. The wake text asks children to run mailbox commands plain, with no pipes or filters.
  • 6 DevSwarm spawn-review fixes: timeout detection, freshness, submodule fetch, and an honest merge label. A real spawnSync timeout (ETIMEDOUT, SIGTERM, null status) never reached cmdSpawn's timeout detection because the error branch returned before the signal check ran; it's now propagated with an explicit timedOut flag, and a timeout now reports plainly, including that a partial workspace may already exist for the branch (never auto-cleaned). Remote-ref freshness no longer consults FETCH_HEAD (it moves on ANY fetch of ANY ref), only the ref's own reflog/loose-ref mtime. A submodule-recursive fetch that fails now retries once with --no-recurse-submodules before being reported as failed. The submodule-worktree-failure regex now only matches fatal: lines actually tied to submodule worktree creation, instead of sweeping in an unrelated fatal error. The roster/finish-column label now distinguishes "done, merge unverified" from "done, not merged" instead of collapsing both into the same dishonest label.
  • DevSwarm no longer flags an ordinary process restart as a concurrent-instance split. The instance-split check only confirmed each instance nonce had a row somewhere in the 15-minute freshness window, not whether two nonces were ever alive at the same time — and a nonce legitimately changes on every restart, so an ordinary restart (old process's last heartbeat still under 15 minutes old, new process's first heartbeat landing minutes later) always read as two concurrent instances even though only one was ever alive. Each nonce's activity span is now padded by 60s and swept for the peak number of nonces overlapping at any point in time; only that peak counts toward instances/instance-split, shared identically by the roster and diagnose paths.
  • .anti-hall/handovers|progress|history paths could double when the session's cwd was already inside one of those directories. precompact-snapshot, handover-resume, tasklist-guard, task-lifecycle-log, progress-prune, and migrate-state all joined .anti-hall/... onto the raw session cwd instead of the git toplevel, so a cwd already under .anti-hall/handovers/ (or any repo subdirectory) doubled the path and handover-resume could never find what precompact-snapshot had just written. All of them now resolve through the repo's git toplevel first (a submodule and a DevSwarm child worktree each keep their own state), falling back to the raw cwd outside a git repo.
  • mcp-reaper can now also detect abandoned Codex app-server-broker.mjs helpers — report-only, never killed. These are spawned detached and unref'd on purpose, so PPID 1 is normal for a live one — not evidence of death like it is for the reaper's ordinary MCP-server matcher — so a helper is only listed as abandoned once its --cwd directory no longer exists, or no live claude/codex process's cwd is equal to, an ancestor of, or a descendant of it (realpath'd, and excluding the broker's own child processes; so a session at a workspace root still owns a broker whose --cwd is inside a submodule beneath it), and only once it's older than the new minimum age. Matched by an exact script-name + codex-plugin-path signature, kept fully separate from the reaper's generic MCP parent-death matcher so that invariant never loosens. Detected brokers are listed in the reaper log only — the reaper never terminates them. Opt-in, and only takes effect when the companion reaper is installed and running.
  • New settings: guards.reaperCodexBroker (boolean, default true; report abandoned brokers) and guards.reaperCodexBrokerMinAgeS (seconds, default 1800), both also in /config.
  • DevSwarm's wake-watch/doctor re-arm could name a path that doesn't exist on disk. The version check that picks the "newest" build took the max across the installed plugin list, the newest cache directory, and the marketplace clone's own plugin.json — but the marketplace clone can fast-forward before that new version is actually mirrored into the plugin cache, so the re-arm command it printed could point at a cache directory that was never created, crashing the moment it ran. Both the wake-watcher's version check and doctor's shared version-check helper now verify the target script file actually exists on disk before returning a path; when a newer version is merely known but not yet cached, they print an update-available notice once and keep running instead of exiting with nothing left watching.
  • A settings reset() could silently disarm a safety guard under a race, and a DevSwarm supervisor blocker label could go silently blank. settings.js reset now re-checks the safety-switch confirmation gate inside the same lock it uses for the write, closing a window where a concurrent writer could arm a guard and have reset quietly disarm it without --confirmed. Separately, the DevSwarm supervisor's blocker-label dedupe now reads back the mode field it already writes; previously that field was dropped on read, so every deduped (skipped) re-ask treated an already-promoted label as unset, silently blanking it for the rest of the re-ask interval.
  • Jev's own triage integration was invisible in jev report, and its agree% swung wildly across identical time windows. Triage decisions were logged to a separate file with a different schema, so the report's per-integration table never showed them at all rather than showing them with a bad value — they now appear as their own triage row. The agree% denominator now counts distinct, fresh decisions only (a cache-hit retry of the same decision no longer adds an extra vote), which was the actual cause of agreement swinging between 99%/77.5%/37% across windows on the same underlying data; the table now also prints the sample size inline. The same blocker-label read-back fix noted above is included here too, proven with a fixture repro (10 sweeps: label once, then null the other nine before the fix).
  • The DevSwarm ingest daemon now keeps hivecontrol's error output when a monitor call fails or is killed. The daemon used to run each monitor call with stderr discarded (stdio: [..., 'ignore']), so a non-zero exit or an external signal (not the daemon's own timeout kill, which already had its own stable message) carried no diagnostic text at all. Stderr is now piped and kept as a bounded last-2 KB tail, folded into the resolved error message on a non-zero exit or a signal with no prior spawn error; the existing ok/fail classification is unchanged, only the missing diagnostic text.
  • The progress/history guard now names the absolute path in its block message. The Stop-hook guard reads and writes the root-joined absolute .anti-hall/progress|history path, but its block message named the bare relative path — from a repo subdirectory cwd, following that message literally wrote into the subdirectory instead, a location the guard never checks. The message now names the absolute path.
  • Repo-root resolution no longer climbs to a parent repo when the session cwd was deleted, and no longer resolves to the home folder when ~ is itself a git repo. A cwd that no longer exists (for example a removed nested worktree) used to climb to a surviving ancestor repo and write that session's state there; it now falls back to the raw, nonexistent cwd like any other unresolvable case. Separately, a dotfiles repo checked out at $HOME could resolve its toplevel to the home directory itself, which would redirect every write into anti-hall's own global ~/.anti-hall/ store; that case now also falls back to the raw cwd.
  • Correction: Jev integration postHandoverGate defaults to off, not shadow. Replace the earlier "New Jev integration postHandoverGate (shadow only, default shadow)" bullet: the row exists (jevIntegrations.postHandoverGate), but it ships off. Set it to shadow to log Jev's view without changing what the gate says.
  • DevSwarm parent Stop gate: "busy" now needs real evidence. A child with unread mail only gets the non-blocking busy, N queued line when its own transcript was written in the last devswarm.parentGateBusyFreshMin minutes (default 5) and its latest turn is real work. A live process or a fresh heartbeat no longer counts, so an idle child sitting at its prompt with old mail blocks and escalates again. A missing or unreadable transcript (including a Codex child) also blocks.
  • Waiting children always block. A child stuck on an unanswered question, a plan approval (ExitPlanMode), or any tool call while its transcript has gone quiet (a permission prompt or a hung tool) blocks with a "waiting on a human answer" line. In a family of twin descriptors, one waiting member makes the whole family block.
  • Age cap on the busy advisory. Even a busy child blocks once its oldest unread is older than devswarm.parentGateBusyMaxAgeMin (default 60): ": busy but hasn't read mail in Xm". A busy pass no longer resets the forced-ack count, so escalation still fires after the usual number of blocks.</li> <li><strong>New settings:</strong> <code>devswarm.parentGateBusyFreshMin</code> (minutes, default <code>5</code>) and <code>devswarm.parentGateBusyMaxAgeMin</code> (minutes, default <code>60</code>), both also in <code>/config</code>.</li> </ul> <h2 id="01085">0.108.5<a class="headerlink" href="#01085" title="Permanent link">¶</a></h2> <h3 id="p0-anti-hall-could-move-a-git-tracked-planning-folder-in-child-worktrees">P0: anti-hall could move a git-tracked <code>.planning/</code> folder in child worktrees<a class="headerlink" href="#p0-anti-hall-could-move-a-git-tracked-planning-folder-in-child-worktrees" title="Permanent link">¶</a></h3> <p>anti-hall could move a git-tracked <code>.planning/</code> folder in child worktrees (and submodules). This is now never automatic, it is copy-only, and <code>doctor</code> shows how to restore.</p> <ul> <li><strong>Cause.</strong> <code>hooks/lib/doctor-repair.js</code> registered the GSD <code>.planning/</code> fold as an automatic repair, so <code>repair-on-reload</code> ran it in every session, DevSwarm child worktrees included. <code>scripts/migrate-state.js migrateGsdPlanning</code> copied each file into <code>.anti-hall/history/legacy/planning/</code> and then deleted the source, which moved git-tracked files.</li> <li><strong>Never automatic.</strong> The fold is gone from doctor's repair pass, so repair-on-reload, <code>doctor --repair</code> and the updater never run it. It is also no longer part of the capability scan's pending-migration check. It runs only as the explicit, human-typed <code>migrate-state.js --planning</code>.</li> <li><strong>Copy-only.</strong> Even when run explicitly it never unlinks, renames or removes a source, and never overwrites a different existing legacy copy. It skips the whole tree, and prints why, when <code>.planning/</code> is git-tracked, when the directory is a linked/child worktree rather than the main checkout, when it is inside a submodule, or when the git location cannot be confirmed.</li> <li><strong>Restore.</strong> <code>doctor</code> now reports (it never restores) every worktree of the repo, and their submodules, where tracked <code>.planning/</code> files are missing and a legacy copy exists. It prints <code>git -C <worktree> checkout -- .planning</code>. The opt-in <code>node scripts/migrate-state.js --restore-planning [--dir <worktree>]</code> restores only files that are missing, tracked, and whose legacy copy is byte-identical to <code>HEAD</code>. It never deletes the legacy copies.</li> <li><strong>Audit.</strong> No other automatic migration deletes or moves repo content. The remaining unlink/rm/rename calls in <code>migrate-state.js</code>, <code>migrations.js</code>, <code>doctor-repair.js</code> and <code>devswarm-migrate.js</code> act on temp-file renames, lock files, or <code>~/.anti-hall</code> state.</li> </ul> <h3 id="fixes_14">Fixes<a class="headerlink" href="#fixes_14" title="Permanent link">¶</a></h3> <ul> <li><strong>Every jev-assist call site now threads <code>sessionId</code>/<code>turnRef</code> into the logged row.</strong> Live speculation decisions (the <code>speculation</code> add-block integration, the only one that can actually add a block while "on") were logging <code>sessionId: null</code> for every row, making them unjoinable to the transcript that produced them. <code>hooks/lib/jev-assist.js</code> now threads <code>sessionId</code> through every <code>ask()</code>/<code>askSync()</code>/<code>askDetached()</code>/<code>consultRelax()</code> call, and adds an optional <code>turnRef</code> (the transcript's last-line ISO timestamp, or a line-count fallback via the new <code>turnRefFromTranscript()</code>) so a decision row can be pinned to which turn it was about. Every integration call site (<code>speculation-guard</code>, <code>claim-ledger</code>, <code>task-tracker</code>, <code>merge-gate</code>, <code>output-verify-guard</code>, <code>model-routing-guard</code>, <code>codex-nudge</code>, <code>tasklist-guard</code>) now passes both fields where a session/transcript is available; <code>devswarm-supervisor.js</code>'s background sweep intentionally stays session-less.</li> <li><strong><code>supervisorBlockerLabel</code> no longer re-asks/re-logs an unchanged input every sweep.</strong> The DevSwarm supervisor's ~90s liveness sweep re-asked (and re-logged) the exact same blocker-label decision on every tick as long as the child's jev-triage pending entry stayed unchanged — one real input produced 382 <code>jev-assist.ndjson</code> rows in 24h. A new per-workspace state file under <code>~/.anti-hall/devswarm/blocker-label-ask/</code> now hashes the input (childId + kind + ts) and skips the ask/log entirely when it matches the last one asked, re-asking only when the input changes or the configurable re-ask interval (<code>devswarm.supervisorBlockerLabelReaskSec</code>, default 6h) elapses.</li> <li><strong><code>jev report</code> now prints its effective time window and row counts.</strong> <code>--since</code>/<code>--until</code>/ <code>--exclude-window</code> already filtered every metric (including the agreement metric) and already read the rotated <code>.1</code> log alongside the live one — but the report never showed WHICH window it actually used or how many rows were counted vs excluded, so two runs against "the same window" could not be verified as comparable. <code>jev report</code> (and <code>jev report --by project|session</code>) now print a <code>window: <since> .. <until> exclude: <…> rows: N in window / M total (K excluded)</code> line (also included as <code>window</code> in <code>--json</code> output), via the new <code>describeWindow()</code>/<code>printWindow()</code>.</li> </ul> <h2 id="01084">0.108.4<a class="headerlink" href="#01084" title="Permanent link">¶</a></h2> <h3 id="features_5">Features<a class="headerlink" href="#features_5" title="Permanent link">¶</a></h3> <ul> <li><strong>Every Jev integration now has its own dedicated setting.</strong> A new <code>jevIntegrations</code> settings-schema section gives each of the 12 integration ids (<code>speculation</code>, <code>triage</code>, <code>newRequest</code>, <code>claimLedger</code>, <code>outputVerifyGuard</code>, <code>gitGuardSelfCredit</code>, <code>modelRouting</code>, <code>tasklistTrivial</code>, <code>codexNudgeSubstantial</code>, <code>mergeGateHedge</code>, <code>parentGateQuestion</code>, <code>supervisorBlockerLabel</code>) its own row/setting (<code>jevIntegrations.<id></code>), instead of only 5 of them living as <code>advanced</code> sub-keys of the <code>jev</code> section. Each is its own table row in <code>/anti-hall:settings</code> (<code>settings.js show --section jevIntegrations</code>) and its own <code>/config</code> row, titled "Jev integration · <name>". Defaults are unchanged: <code>speculation</code> and <code>triage</code> default <code>on</code>, every other integration defaults <code>shadow</code>.</li> <li><strong>Full back-compat, nothing deleted.</strong> A pre-existing <code>~/.anti-hall/jev.json</code> <code>integrations.<id></code> value, and a pre-0.108.4 <code>~/.anti-hall/settings.json</code> <code>jev["integrations.<id>"]</code> value, both keep working and forward-migrate automatically into the new <code>jevIntegrations.<id></code> key (<code>companion/lib/migrations.js</code> <code>migrateJevIntegrationsSection</code>, idempotent, fail-open, never deletes the old key). Resolution precedence is unchanged: env > settings.json > <code>/config</code> > jev.json legacy > default. <code>jev-setup.js mode <id> on|shadow|off</code> now writes the new canonical <code>jevIntegrations.<id></code> settings.json key (still also writes <code>jev.json</code> for back-compat).</li> <li>Docs: <code>docs/GUIDE.md</code>, the <code>jev</code>/<code>settings</code> skills (Claude + Codex), and the system-briefing operator guide (Claude + Codex) all cover the new section/keys.</li> <li><strong>New Jev integration: <code>findingDedup</code>, wired into deadly-loop.</strong> <code>scripts/finding-dedup.js</code> groups deadly-loop TRIO (Reviewer/Auditor/Critic) findings that Jev judges to describe the SAME underlying issue, via <code>jev-assist.js</code>'s <code>ask()</code> path (id <code>findingDedup</code>, trust <code>advisory</code>) — candidate pairs are same-file within ±40 lines or the same id recurring across rounds, capped at 200 pairs/run, concurrency 4, union-find grouped at confidence ≥0.85. Offline benchmark (2026-09, 3 projects, 30 days of real reviews): Jev answered the exact "same underlying issue?" question 65/65 correct at confidence ≥0.85, vs 45% precision for a same-file ±10-lines heuristic baseline — the 13th <code>jevIntegrations</code> settings row, defaulting <code>on</code> (unlike every other 0.108.4 integration, which defaults <code>shadow</code>) on the strength of that result. Fail-open throughout: Jev off/unconfigured/ erroring → no groups, exit 0. <code>deadly-loop</code>/<code>deadly-loop-multi</code> (Claude + Codex mirror) show its output as an advisory hint after each round's findings are collected — it never auto-collapses; the agent still decides.</li> <li><strong>Every feature is now controllable from settings.</strong> Each hook anti-hall registers (Claude and Codex) has an on/off switch whose default is the old behaviour. The hook checks it first and does nothing when it is off; a settings error always leaves the hook running. New sections: <code>safety</code> (4 keys), <code>context</code> (7: the verify-first injections, task tracker, handover resume, defect nudge) and <code>maintenance</code> (5: repair-on-reload, progress prune, pre-compact snapshot, task lifecycle log, session-end MCP reaper). <code>guards</code> gains 7 (<code>modelRouting</code> strict/advisory/off, <code>apiGuard</code>, <code>speculationGuard</code>, <code>claimLedger</code>, <code>taskGuard</code>, <code>tasklistGuard</code>, <code>scanThrottle</code>) and <code>devswarm</code> gains 12 (parent/child gates, the roster and inbox injections, reply tracker, comms and inbox-read guards, the wake watcher, app-DB sync, screenshot sync). The env-only switches <code>ANTIHALL_MODEL_ROUTING</code>, <code>ANTIHALL_REPAIR_ON_RELOAD</code>, <code>ANTI_HALL_SCAN_THROTTLE</code>, <code>ANTI_HALL_SESSION_END_REAPER</code> and <code>ANTIHALL_DEVSWARM_APP_SYNC</code> keep working as the env tier of those settings. <code>userConfig</code> grows from 39 to 76 <code>/config</code> rows (still no <code>options</code> field).</li> <li><strong>Safety guards need a confirmed change, not a hard refusal</strong> (owner decision: no guard needed — a human direct command, or a confirmation after a clear, plain warning, is enough). <code>safety.gitGuard</code>, <code>safety.commandGuard</code>, <code>safety.editGuard</code>, <code>safety.swarmGuard</code>, and the knobs that weaken a safety guard (<code>guards.stashGuard</code>, <code>guards.editGuardAllow</code>, <code>guards.allowSubagentMailbox</code>) work through <code>settings.js set</code>/<code>reset</code> like any other key, but need <code>--confirmed</code>. Without it, nothing changes and the call returns one short, factual, human-readable line — "Turning off <code><guard></code> means <code><what it protects, in one plain sentence></code>. Ask the user to confirm, then re-run with --confirmed." (calm facts, not alarming) — and <code>{ok:false, needsConfirmation:true, warning}</code> on <code>--json</code>. The warning text comes from a new per-key <code>safetyNote</code> in the schema. Normal precedence (env > <code>~/.anti-hall/settings.json</code> > <code>/config</code> > legacy > default) is restored for these keys — the confirmation is the protection now, not an ignore rule. The <code>settings</code> skill (Claude + Codex) and system-briefing: a direct user ask to change a guard IS the confirmation (apply with <code>--confirmed</code> right away); otherwise show the warning and ask (<code>AskUserQuestion</code> on Claude, a numbered yes/no on Codex) before applying — never infer consent, never confirm on the agent's own initiative. The <code>skip.json</code> escape hatch is unchanged, and <code>"all"</code> still does not cover git-guard.</li> <li><code>settings.js show</code> lists every section, marks the safety rows (<code>needs --confirmed</code>), and ends with the parts that have no switch on purpose (skip-guard, coordinator-detect, omc-detect, phase-tracker, fable-availability, codex-availability, emit-dedupe-reset, agent-watchdog, command-guard's data-safety sub-guards) and why.</li> <li><strong><code>devswarm.heldPartitions</code> (owner-held mesh partitions).</strong> A new csv settings key (env override <code>ANTIHALL_DEVSWARM_HELD_PARTITIONS</code>) lets an owner permanently exempt specific mesh partition ids from the per-turn "ORPHANED MESH" warning and from <code>reap-orphans</code>. Held ids are diverted out of <code>orphans[]</code> and into their own <code>heldPartitions[]</code> field by <code>computeSummary</code> — never dropped, still visible via <code>devswarm diagnose</code>/<code>devswarm healthcheck</code> as owner-held. <code>reap-orphans</code> refuses them even under <code>--apply</code> as an extra belt; the reaper already never auto-deletes anything (dry-run default, human-only, <code>--apply --max N</code> required).</li> </ul> <h3 id="changed_24">Changed<a class="headerlink" href="#changed_24" title="Permanent link">¶</a></h3> <ul> <li>A boolean env switch now accepts every true/false token (<code>0</code>/<code>off</code>/<code>false</code>/<code>no</code>), not only the one literal it used to check. The wake watcher's refusal vocabulary gains <code>disabled-by-settings</code>.</li> </ul> <h3 id="fixes_15">Fixes<a class="headerlink" href="#fixes_15" title="Permanent link">¶</a></h3> <ul> <li><strong><code>devswarm.js spawn</code> no longer hands a child stale tooling.</strong> hivecontrol branches the new workspace from a LOCAL branch name (<code>-s</code>, else the caller's current branch), so a Primary whose local <code>main</code> had fallen 25 commits behind <code>origin/main</code> spawned a child with an old deploy script that lacked a newer CI gate. When the source is the default branch (resolved from origin/HEAD), spawn now runs <code>git fetch origin <default></code> and fast-forwards the local branch to it before <code>create</code>: a guarded ref update when it is not checked out, <code>merge --ff-only</code> when it is checked out and clean. It never rebases, resets or forces, and touches no other branch. Offline, it warns and continues. When local <code><default></code> is behind and cannot be fast-forwarded (diverged, or checked out with local changes) spawn refuses with a one-line reason ("local main is N commits behind origin/main; spawning from it would give the child outdated tools") unless <code>--from-local</code> is passed; that flag is anti-hall's own and is not forwarded to hivecontrol. The result carries a <code>sourceCheck</code> object.</li> <li>New setting <code>devswarm.spawnFromOrigin</code> (default <code>true</code>, a /config row) turns the check off.</li> <li><strong><code>command-guard.js</code> no longer forces a subagent spawn just to read one value.</strong> Trivial READ-ONLY commands in the main coordinator thread — <code>git ls-remote</code>/ <code>fetch</code>/<code>merge-base</code>/<code>ls-tree</code>, <code>sqlite3 -readonly</code>, <code>gcloud/gh/kubectl describe|list|get|view</code> (and <code>gcloud logging read</code>), and anti-hall's own read-only CLI subcommands (<code>jev-setup.js status</code>, <code>settings.js show|get</code>, <code>hooks/doctor.js</code> without <code>--repair</code>/<code>--fix</code>, <code>jev-report.js</code> without <code>label</code>/<code>prune-audit</code>) — tripped the generic heavy-verb/<code>node *.js</code> heuristic and were blocked with "Heavy command detected", forcing a subagent spawn for a one-line lookup. Added a narrow, per-segment read-only allowlist evaluated BEFORE the heavy-command check (same place/discipline as the existing <code>devswarm.js</code> carve-out), plus a conservative <code>isSafeNodeEval()</code> classifier for <code>node -e</code>/<code>--eval</code> payloads (deny-listed write/spawn Node APIs only). State-changing variants (<code>git push/pull</code>, <code>settings.js set</code>, <code>sqlite3</code> without <code>-readonly</code>, <code>doctor.js --repair</code>, <code>jev-report.js label</code>, <code>gcloud ... delete</code>) stay blocked — the allowlist is scoped to the specific read-only subcommand, never the whole script/verb.</li> <li><strong><code>model-routing-guard</code>'s planning-shaped-on-haiku advisory false-positived on mechanical work.</strong> Row 4 matched the broad <code>COMPLEX</code> word list (bare <code>review</code>, <code>audit</code>, <code>design</code>, <code>plan</code>, <code>root cause</code>, <code>regression</code>, <code>logic</code>, <code>security</code>) anywhere in the spawn's description/prompt, including inside backtick-quoted config keys/CLI flags (<code>`jev.audit.snippets`</code>) and inside ledger content being copied verbatim. A 316-spawn field sample from this project's own transcripts measured a 24% false-positive rate on haiku spawns that were correctly mechanical (status checks, report reads, ledger appends, defect filing, CI watching). Row 4 now uses a stricter <code>PLANNING_INTENT_RE</code> (an actual planning verb phrase: <code>design a/the</code>, <code>plan a/the</code>, <code>architecture</code>, <code>brainstorm</code>, <code>deep/code/security review</code>, <code>root cause analysis</code>, <code>security audit</code>, etc.), matched with backtick-quoted spans stripped, and suppressed when the corpus marks itself read-only/mechanical/verbatim/fixed-command (<code>READONLY_SUPPRESS_RE</code>). Measured false-positive rate on the sampled corpus: 24.1% → 0%; recall on genuine planning-shaped fixtures (design/plan/review/ audit/root-cause tasks) unchanged at 100%. Rows 1-3's <code>COMPLEX</code>-anywhere veto is untouched — it stays broad because being generous there only prevents a block, the safe direction.</li> <li><strong><code>jev-assist.test.js</code> no longer reads the real machine's home.</strong> The <code>getMode</code> test for <code>codexNudgeSubstantial</code>/<code>tasklistTrivial</code> called <code>getMode(id, cfg)</code> without the 3rd <code>home</code> argument; both ids have a <code>settings-schema.js</code> <code>integrations.<id></code> entry, so <code>getMode</code> resolved them via <code>schemaIntegrationMode(id, home) -> settings.js get(..., { home })</code>, and a missing <code>home</code> fell back to <code>os.homedir()</code> — reading the real <code>~/.anti-hall/settings.json</code>/<code>jev.json</code>. Failed on any machine with a customized Jev mode for those ids. Fixed by passing an isolated <code>home</code> from <code>makeHome()</code>. Extended <code>tests/hygiene/settings-home-injection.test.js</code> with a <code>getMode</code> regression guard that points <code>HOME</code> at a poisoned settings.json and proves <code>getMode(id, cfg, home)</code> never reads it.</li> <li><strong>A child's <code>primary-<hash></code> label no longer comes back as a live "ghost" roster row under the Primary's session.</strong> Root cause: <code>reconcile</code> (run by <code>update</code> or doctor-repair inside a Primary session) spawns <code>inbox pull <id></code> for each row with cwd set to that row's worktree, but it passed the Primary's own environment through. Inside that subprocess the cwd made the child's label look like the caller's own row, so auto-ensure re-created the label's tombstoned descriptor and promoted it to the Primary's <code>CLAUDE_CODE_SESSION_ID</code>. The sweep subprocess now drops <code>CLAUDE_CODE_SESSION_ID</code>/<code>DEVSWARM_BUILDER_ID</code> and never promotes a session. Additional safeguards: <code>inbox pull</code> of a child label (or of a label whose <code>retired/</code> redirect names another id) now pulls under the canonical id and writes nothing under the label. <code>heartbeat</code>/<code>register</code>/<code>ensure</code> refuse a retired label even when no row is left to check. <code>roster</code> and the per-turn table fold a label row into its canonical row (via the sender alias, the retired redirect, or the DevSwarm app's builder for that worktree) and carry its unread count over. <code>repair-child-sender-labels</code> now folds a label whose live session is running on a different worktree instead of leaving it pending forever, and it moves the label's leftover descriptor into <code>archived-retired/</code> rather than deleting it. Messages sent to the alias are forwarded, and they stay readable under the alias.</li> <li> <p><strong><code>cmdArchive</code> claimed hivecontrol "has no teardown command" — false on DevSwarm >= 2.5.3.</strong> A live substrate test (2026-09-25) proved <code>hivecontrol workspace archive <id></code> exists and works (sets isActive=0/isHidden=1). When the capability gate allows it, <code>devswarm.js archive</code> now ALSO archives the workspace in the DevSwarm app, explicit id always, retrying once on the known-flaky "Could not confirm terminal process boundary" error; dormant/failed falls back to an accurate manual-step instruction instead of the old false claim. Also fixed: the post-archive "child session still live" warning fired off a heartbeat file alone, which can be stale (e.g. the workspace was already deleted in the app); it now cross-checks the app DB's own builder rows and suppresses the warning when there is no row for that id at all.</p> </li> <li> <p><strong><code>output-verify-guard.js</code> false positive on zero-count summaries.</strong> A clean <code>cargo test</code> summary line ("test result: ok. 5 passed; 0 failed") was flagged as a mixed pass/fail result because the dual-signal regex matched the literal "0 failed" text. Count-bearing fail/pass patterns now require a non-zero count; plain marker patterns (<code>FAIL</code>, <code>PASS</code>, <code>--- FAIL:</code>) are unaffected. Covers pytest, jest, cargo, go test, and mocha summary lines.</p> </li> <li> <p><strong><code>devswarm-parent-inbox.js</code> re-instructed re-sending an already-pending archive-request.</strong> A child the Primary had already sent <code>archive-request</code> to kept triggering the CHILD NOT DRAINING nag and the cooldown'd ARCHIVE-READY reminder every turn, even hours later, even while the child's session was still live — the child was simply waiting on its own user, not neglected (field report: a downstream project, 399105fe/e75cade3). Both are now suppressed while <code>computeSummary</code>'s <code>archive_request_only_unread</code> is true, resuming automatically once the request is answered/drained or the child resumes other work, or after a new <code>devswarm.archiveRequestRenagHours</code> settings key (default 24h) elapses. The dead-session ARCHIVE-READY reminder stays exempt (<code>archive-request</code> auto-archives a dead target on its next run, so it is never redundant there); a genuinely <code>stuck</code>/escalated status is unaffected.</p> </li> <li> <p><strong>Jev cost tracking logged <code>costUsd: null</code> forever ("not configured").</strong> <code>computeCostUsd</code> only ever computed a real per-token cost from an owner-populated <code>jev.prices</code> table, so every row stayed null until an owner manually filled it in. It now falls back to a BUILT-IN default rate (new <code>jev.priceUsdPerMInput</code>/<code>jev.priceUsdPerMOutput</code> settings, default <code>0.042</code>/<code>0</code> — verified: typesafe.ai, vercel.com/ai-gateway/models/jev, openrouter.ai/typesafe) whenever <code>jev.prices</code> has no matching entry, so real cost is populated out of the box. <code>jev-report.js --by project|session</code> now also prints a real-cost summary (total + per-integration) scoped to that project/session; the gateway credit balance is still read exactly as before.</p> </li> <li> <p><strong>Auto-archive gate (h)'s no-re-archive record depended solely on a general-purpose log file.</strong> <code>~/.anti-hall/logs/devswarm-auto-archive.ndjson</code> is a LOG other log-rotation/ pruning paths in this codebase could legitimately remove, which would silently re-open the hole gate (h) exists to close (an owner-unarchived workspace getting auto-re-archived once the focus/idle windows pass again). Moved to a durable, never-rotated/pruned state file (<code>~/.anti-hall/devswarm/auto-archived.json</code>, <code>{id: [{doneHead, at}]}</code>, append-only). An idempotent forward-migration seeds it from existing ndjson records (wired into both <code>doctor --repair</code> and <code>update.js</code>'s post-pull pass); the ndjson log itself is still written on every archive and still read as a fallback for any pre-migration record.</p> </li> <li><strong><code>settings.js reset</code> could silently disarm a human-armed safety guard.</strong> <code>reset()</code> never asked for <code>--confirmed</code> on the assumption that removing an override is always the safe direction — false for <code>guards.stashGuard</code>, whose default is <code>false</code>: resetting a human-armed <code>true</code> fell back to the default and disarmed it with no warning. <code>reset()</code> now computes the effective value the key will have AFTER its settings.json override is removed (env, <code>/config</code>, legacy, default) and gates it through the same risky-direction check as <code>set</code>; a reset that leaves the value unchanged or restores a safe default (e.g. <code>safety.gitGuard</code> back to <code>true</code>) still needs no confirmation.</li> <li><strong><code>cmdArchive</code>'s new app teardown could reach <code>hivecontrol workspace archive</code> with a wrong target, and from a hook.</strong> DevSwarm 2.5.3's archive/delete verbs default to the CURRENT workspace when no id is given, and the phantom-descriptor retire in <code>devswarm-child-turn.js</code> (a UserPromptSubmit hook) calls <code>cmdArchive</code> with a truncated id. The app call now runs only when the app DB (read-only, fresh read) holds a builder with that EXACT full id that is open and whose <code>builderType</code> is known and not <code>primary</code>; a truncated id, a <code>primary-<hash></code> label id, an unknown builderType, a closed builder, or an unreadable app DB means no spawn at all (not even the capability probe). <code>cmdArchive</code> takes <code>{appArchive:false}</code> for a local-only archive, and the hook passes it, so no hook ever spawns hivecontrol or spends its timeout budget on it.</li> <li><strong>A settings.json <code>jev.integrations.<id></code> value lost to a conflicting jev.json value during migration.</strong> <code>migrateSettingsFromLegacy</code> ran the jev.json legacy loop first, which claimed the empty <code>jevIntegrations.<id></code> key, so the later settings.json forward-migration saw the key as set and skipped it — inverting the precedence (settings.json outranks jev.json). <code>migrateJevIntegrationsSection</code> now runs first. Nothing is deleted; both old values stay where they were.</li> <li><strong>The roster fold dropped a ghost row's unread when the canonical row's <code>directUnread</code> was null.</strong> Folding a <code>primary-<hash></code> ghost into, e.g., an archived child row (whose <code>directUnread</code> is null) kept the null and lost the ghost's unread directs and its hints. The canonical row now receives both (unread summed from 0, hints unioned), so the roster's total unread is the same before and after the fold.</li> <li><strong><code>devswarm-parent-inbox.js</code> lost a folded ghost row's unread.</strong> A <code>primary-<hash></code> ghost folded under its canonical row was skipped outright, so its unread directs never reached the canonical row's nag. The fold is now resolved in a pre-pass and the ghost's unread is added to the canonical row (table, attention, own-checkout path alike), whichever order the two rows appear in.</li> <li><strong><code>finding-dedup.js</code> reported false duplicate groups, including with Jev off.</strong> The same finding id recurring across rounds was pushed into the union-find id list twice, so it formed a <code>[X, X]</code> "group" with no Jev edge at all, and a cross-round pair of one id was unioned with itself. Findings are now keyed by id+round (exact repeats dropped, first wins), a finding is never paired with itself, an id that recurs across rounds is reported as <code><id>@round<N></code>, and Jev off means no groups and no per-pair calls. The Jev cache key now also includes each finding's round and text, so a reused id never gets another finding's cached answer.</li> <li><strong><code>model-routing-guard</code>'s read-only suppression hid genuine planning on haiku.</strong> The Row-4 fix above suppressed the advisory whenever the spawn said "read-only", so a "read-only code review / security audit of X" on haiku got no advisory. Row 4 is now suppressed only when the spawn is read-only AND mechanical (fixed commands, "run exactly", "return ≤N lines") AND carries no review/design/audit/analysis verb. The intent set also covers "review this PR/diff", "audit the/this …" and root-cause asks ("find/identify/diagnose the root cause", "root cause why …"); a bare "root cause:" noun in ledger content does not count. Re-measured: 0/316 false positives on the same 316 real mechanical haiku spawns (0.0%; the pre-0.108.4 guard fired on 75/316, 23.7%), and 18/18 on a genuine-planning set that includes the read-only review/audit, PR/diff review, "audit the" and root-cause cases (the earlier 0.108.4 fix caught 8/18).</li> <li><strong>Docs described the safety-key gate wrongly.</strong> The <code>safety</code> section description in the schema still said <code>settings.js set/reset</code> refuses these keys and a settings.json value is ignored, and <code>llms.txt</code> said they are never changed via <code>settings.js set</code> or <code>settings.json</code>. The schema, the CLI usage text, <code>llms.txt</code>, <code>docs/GUIDE.md</code>, the <code>settings</code> skill and the system briefing (Claude + Codex) now say the same thing as the code: a risky <code>set</code> and a <code>reset</code> whose fallback value is risky need <code>--confirmed</code>, re-arming never does, and a settings.json value counts like any other. They also note that nothing mechanically stops an agent from writing settings.json directly — the owner chose consent over an extra guard.</li> </ul> <h2 id="01083">0.108.3<a class="headerlink" href="#01083" title="Permanent link">¶</a></h2> <h3 id="features_6">Features<a class="headerlink" href="#features_6" title="Permanent link">¶</a></h3> <ul> <li><strong>Every non-advanced setting is now a row in Claude Code's native <code>/config</code> panel.</strong> <code>plugin.json</code> <code>userConfig</code> grows from 19 to 39 entries (existing keys unchanged). Titles are prefixed with the section ("Guards · Stash guard") since <code>userConfig</code> has no grouping field. Enum settings are <code>type: "string"</code> fields whose description lists the allowed values (a <code>userConfig</code> <code>options</code> picker would break loading on Claude Code versions before v2.1.271, per the plugin manifest reference, and anti-hall is a public plugin), and numbers carry the schema's <code>min</code>/<code>max</code>. A drift test (<code>tests/hooks/settings-schema.test.js</code>) fails if <code>userConfig</code> stops matching the schema's non-advanced set, or if any field declares <code>options</code>. Each exposed key round-trips through <code>CLAUDE_PLUGIN_OPTION_<KEY></code> and <code>pluginConfigs</code> and shows Source <code>/config</code>. <code>/config</code> rows need Claude Code ≥ 2.1.269; older versions still work via the skill. Codex has no <code>userConfig</code> equivalent and still uses the <code>anti-hall-settings</code> skill.</li> <li>Three settings used to read only their env var and skipped the settings chain. They now go through <code>settings.js</code> (env > file > <code>/config</code> > default), so the new <code>/config</code> rows take effect: <code>guards.emitDedupe</code> (<code>hooks/lib/emit-dedupe.js</code>), <code>jev.judgeModel</code> (<code>hooks/speculation-judge.js</code>, <code>hooks/lib/jev-triage-worker.js</code>) and <code>defects.defaultProj</code> (<code>scripts/defect.js</code>).</li> <li>The <code>settings</code> skill (Claude and Codex) now points the user to <code>/config</code> by default. It applies a named change with a single <code>set</code> and prints the full table only when asked.</li> </ul> <h3 id="fixes_16">Fixes<a class="headerlink" href="#fixes_16" title="Permanent link">¶</a></h3> <ul> <li><strong>Auto-archive never archived anything (<code>candidates:10, wouldArchive:[]</code>).</strong> Gate (a) "done" required <code>archive_ready</code> (every required gate: <code>done,merged,tests_passed</code>), and those rows are only written by a manual <code>devswarm.js gate --set</code>. Gate (a) now also passes on the child's structured done-report: the <code>done</code> gate on its own (<code>devswarm.js gate <id> --set done</code>). <code>tests_passed</code> is not required, because gate (b) still has to prove the merge, and gates (c)-(g) are unchanged. Chat text such as "DONE" never counts. Only the auto-archive decision changed; <code>archive_ready</code> means the same for the parent gate and the merge gate.</li> <li><strong>A stale <code>done</code> gate could auto-archive unmerged commits.</strong> The gate was never cleared, so a child that reused its branch after a squash-merged PR kept <code>done</code>, the PR fallback matched the old merged PR, and new commits were eligible. The <code>done</code> verb now records the worktree HEAD with the gate (<code>set_by</code> <code>devswarm-done@<sha></code>, projected as <code>doneHead</code>), and gate (b)'s app PR row counts only for a done-report tied to HEAD and only when git ancestry can't be determined. When git says "not an ancestor" the lane stays blocked, so a reused branch with new commits never auto-archives. The DevSwarm app DB's <code>pull_requests</code> table has no head/commit sha column (15 columns, checked on a live DB), so a merged PR row can't be tied to HEAD and a squash-merged child still needs a manual archive. A <code>done</code> set any other way (<code>gate --set done</code>, no sha) still works for the Primary, but needs the git proof.</li> <li><strong>Gate (b) called merges "unproven" that <code>gate --set merged</code> had just verified.</strong> The verb checked HEAD against the remote default branch (<code>origin/HEAD</code>'s target, e.g. <code>origin/main</code>); gate (b) checked the LOCAL source branch first and stopped there when it said "not an ancestor", so a stale local <code>main</code> (behind <code>origin/main</code>) hid every merge. Both now call one function, <code>devswarm-git-truth.js</code> <code>gitMergeProof</code>: HEAD must be an ancestor of the remote default branch (<code>origin/HEAD</code>'s target), and a local branch ref never counts, so a local <code>main</code> with unpushed commits can't prove a merge either. The verb records the HEAD it verified at in the <code>merged_verified</code> row (<code>set_by</code> <code>devswarm-merged@<sha></code>, projected as <code>mergedVerifiedHead</code>). When git can't decide, a <code>merged</code> gate verified at the current HEAD proves gate (b); an unverified <code>merged</code> gate, or one verified at an older HEAD, never does. A resolved "not an ancestor" always blocks. Gate (b) proves against the remote now, so a repo with no <code>origin</code> never auto-archives.</li> <li><strong>Nothing made a child set the <code>done</code> gate, so auto-archive still never fired.</strong> New child verb <code>devswarm.js done [<id>] [--summary "..."]</code>: it sets the <code>done</code> gate on the caller's own workspace id (never <code>merged</code>/<code>tests_passed</code>) and sends the Primary one <code>[[ANTIHALL_DONE]]</code> message (<code>kind:'done'</code>). A re-run on the same HEAD adds nothing; a foreign id or the Primary checkout is refused. The child SessionStart directive (<code>devswarm-child-role.js</code>, Claude and Codex) now tells a child to run it once its work is merged, and the roster shows such a child with the <code>done</code> + <code>archive-pending</code> hints. Auto-archive then retires it once the merge is proven and gates (c)-(g) pass, so the user no longer archives finished workspaces by hand.</li> <li><strong>Auto-archive treated a child as the Primary.</strong> Gate (e) counted any <code>primary-<hash></code> descriptor id sharing a builder's worktree as the Primary, and a legacy child descriptor carries exactly that label, so a standard child was never archived. The app DB's <code>builderType</code> is now the authority; without that column, or when the row's value is NULL, empty or whitespace-only (gate (e) and <code>primary-seat.js</code> both trim it), the Primary seat's main-checkout rule decides; a <code>primary-</code> id prefix never does. An unreadable app DB still archives nothing.</li> <li><strong>A stale anti-hall archived marker overrode the DevSwarm app.</strong> A workspace open in the app (<code>isActive=1</code>, <code>isHidden=0</code>) but still holding <code>archived/<id>.json</code> kept counting as archived and triggered a screenshot ask. The supervisor's app-DB sync now trusts the app for a marker anti-hall wrote from app evidence (<code>archivedBy</code> <code>devswarm-app</code>, <code>devswarm-app-deleted</code> or <code>devswarm-ui-sync</code>; same id and worktree, older than 10 min): it first restores the workspace to active with the same logic <code>unarchive</code> uses (puts the descriptor back in <code>workspaces/</code> if it is missing, revives the registry row), and only then moves the marker to <code>archived-retired/<id>.<ms>.json</code> with a <code>retired</code> record (when, by whom, why, <code>restored</code>). It is never deleted. If a restore step fails, the marker stays in <code>archived/</code>, so the workspace is never left in neither directory. A marker from anti-hall's own <code>archive</code> verb is never restored automatically, at any age: an open app row there only means the owner hasn't archived it in the app yet. It stays archived, is reported as <code>openButMarkedArchived</code>, and <code>unarchive</code> undoes it. The new migration <code>retire-stale-archived-markers</code> (<code>doctor --repair</code>) applies the same rules to existing markers. The screenshot ask now fires only when the app DB is unreadable. Nothing is unarchived in the app itself.</li> <li><strong>Auto-archive re-archived a workspace the owner had just unarchived.</strong> The owner's only undo is unarchiving in the DevSwarm app (there is no CLI), but the <code>done</code> gate at that HEAD stayed set, so the sweep archived it again once the focus and idle windows passed. Each successful auto-archive now logs <code>{id, doneHead, at}</code> in <code>~/.anti-hall/logs/devswarm-auto-archive.ndjson</code>, and new gate (h) (<code>h-rearchive</code>) never auto-archives the same workspace again at that HEAD. A new <code>done</code> at a new HEAD makes it eligible again. The log is append-only and no retire/restore path touches it.</li> <li><strong>A local branch named <code>origin/main</code> could prove a merge.</strong> <code>gitMergeProof</code> shortened <code>refs/remotes/origin/<b></code> to <code>origin/<b></code> before calling <code>rev-parse</code>/<code>merge-base</code>, and git resolves that short name to a local branch literally called <code>origin/<b></code> first. It now passes the full <code>refs/remotes/...</code> ref to git (an explicit <code>origin/...</code> ref too); <code>via</code> still shows the short name.</li> <li><strong>The screenshot ask never named a real conflict.</strong> The parent-inbox scoping of <code>openButMarkedArchived</code> kept an entry only when its <code>worktreePath</code> equalled the Primary's own checkout (<code>gitTop</code>), but every conflict is a CHILD worktree, so all of them were dropped. It now scopes by repo identity like the D29 filter: the entry's worktree must resolve to the Primary's <code>repoKey</code>. Entries without a <code>worktreePath</code>, or whose worktree no longer resolves, still fail closed. The test fixture had stamped the conflict with the Primary's own path, which hid the bug; it now uses a real child worktree plus a foreign-repo entry that must stay hidden. <code>doctor</code> was not affected (it compares the app-DB <code>repositoryId</code>).</li> <li><strong>An unresolvable <code>origin/HEAD</code> fell back to guesses.</strong> <code>gitMergeProof</code> then proved the merge against <code>origin/<sourceBranch></code> (a workspace's source branch need not be the default branch), and gate (b) could then accept an app PR row matched by branch name alone. It now fails safe: <code>gitMergeProof</code> returns <code>merged: null</code>, <code>via: 'default-branch-unknown'</code>, and gate (b) blocks with <code>default-branch-unknown</code> unless a <code>merged</code> gate was verified at the current HEAD. Run <code>git remote set-head origin --auto</code> in such a checkout to restore auto-archive.</li> <li><strong>The Stop-time handover pause nag could repeat the same text with no progress.</strong> <code>hooks/auto-handover-pause-nag.js</code> stored <code>nagStepPct</code> and <code>lastNagPct</code> but never compared them. It now re-nags only when context has risen at least <code>nagStepPct</code> points since the last nag (inside the quiet window too, advancing the baseline), or at a quiet pause after <code>nagQuietMin</code>, and never repeats the identical percentage within the same step (<code>lastPauseNagPct</code>). Shared by the Codex port.</li> <li><strong><code>recordOutcome()</code> (<code>hooks/lib/jev-assist.js</code>) never populated <code>project</code> on its <code>type:'outcome'</code> rows</strong>, unlike every decision row <code>ask()</code>/<code>askSync()</code>/ <code>askDetached()</code> log via <code>finalize()</code>. <code>triage</code>'s answer-latency join (<code>hooks/lib/jev-triage.js</code>'s <code>recordAnswered</code>) and <code>speculation</code>'s evidence-added/user-override outcomes are the two current callers, so their <code>jev-assist.ndjson</code> rows silently grouped under <code>unknown</code> for <code>jev report --by project</code>/<code>--project <name></code>. <code>recordOutcome()</code> now applies the same cwd-basename fallback (<code>defaultProject()</code>) <code>finalize()</code> already uses, and accepts an explicit <code>project</code> override for future callers. Rows logged before this fix cannot be repaired (no source cwd to recover) and keep grouping under <code>unknown</code> — this is expected, not a bug.</li> <li><strong><code>jev.audit.snippets</code> never wrote a snippet for a shadow-mode decision, even with snippets on.</strong> <code>maybeWriteAuditSnippet</code> (<code>hooks/lib/jev-assist.js</code>) gated on the <code>changed</code> direction, which is null-by-construction whenever <code>mode !== 'on'</code> — and shadow is the DEFAULT mode for every integration under evaluation, so the exact would-change decisions an owner/agent needs to label had no snippet to show. It now also writes when the row's <code>wouldChange</code> signal (the same would-change computation <code>jev-report.js</code> already reads) is truthy, marking that row <code>shadow: true</code> so <code>label</code> readers can tell a would-have-changed decision from an actual one. <code>jev-report.js</code>'s <code>readAuditSnippet</code> already read by hash only, with no filter on <code>changed</code>, so shadow rows surface via <code>label <hash></code> unchanged.</li> <li><strong><code>jev-report.js</code>'s per-integration table showed a bare <code>0.0%</code> in the <code>changed%</code> column (and <code>0 changed</code> in the headline) for label-only integrations</strong> (a <code>choice</code>/string classifier with no boolean baseline to diff against, e.g. <code>newRequest</code>, <code>supervisorBlockerLabel</code>) — indistinguishable from "Jev never changed anything here" when the real story is "there is nothing boolean to compare". The table now prints <code>n/a (N distinct)</code>, a new <code>label-only integrations</code> note below the table (and the <code>headlines</code> section) spells out the real signal — <code>label-only: no boolean outcome to compare; N distinct decisions (M fresh)</code>, deduped by content hash so repeated cache hits of the same decision count once. KEEP/REMOVE/REVIEW verdict gating is unchanged.</li> <li><strong><code>jev-report.js</code> never counted owner-delegated tp/fp labels on a label-only integration's would-change decisions, so such an integration could NEVER reach KEEP/REMOVE.</strong> For a <code>choice</code>/string classifier (e.g. <code>newRequest</code>), <code>effectiveDirection</code> was forced to <code>null</code> for every row, so its would-change hashes never entered <code>changedHashByFresh</code> — the map the tp/fp/precision loop and the <code>known === 0</code> label-only short-circuit both read. Labels were read but silently dropped, and <code>bucket.labeled > 0 && known === 0</code> always won first, permanently returning <code>REVIEW (label-only, no outcome signal yet)</code>. Fixed: a would-change (<code>wouldChange</code>/<code>changed</code> truthy), fresh, hashed choice row now joins a label-candidate set that feeds the same <code>humanTP</code>/<code>humanFP</code>/<code>autoTP</code>/<code>autoFP</code>/<code>labeledSample</code> pipeline a boolean integration's changed decisions use — reported additively via JSON as <code>labelWouldChangeUnique</code>. <code>changed%</code>/<code>changedUnique</code> stay untouched (0; a choice answer has no added/relaxed/changed semantics), the label-only short-circuit now only fires when <code>labeledSample</code> is also 0, and KEEP for a choice integration is gated on its own would-change rate in place of the (always-0) changed-decision rate. KEEP/REMOVE thresholds themselves (≥20 labelled, ≥50/≥200 calls, ≥60%/≥80% good-outcome) are unchanged.</li> </ul> <h3 id="docs--tests">Docs / tests<a class="headerlink" href="#docs--tests" title="Permanent link">¶</a></h3> <ul> <li><strong>Codex parity audit of the six Jev-bearing hooks named in the 0.108.3 census</strong> (<code>output-verify-guard.js</code>, <code>model-routing-guard.js</code>, <code>speculation-guard.js</code>/<code>speculation-judge.js</code>, <code>tasklist-guard.js</code>, <code>devswarm-parent-gate.js</code>, <code>claim-ledger.js</code>): verified against the live <code>plugins/anti-hall/codex/hooks/hooks.json</code>, <code>plugins/anti-hall/codex/install-codex.js</code>, and <code>tests/hygiene/manifest-drift.test.js</code>'s <code>CLAUDE_ONLY_ALLOWLIST</code> that a decision already exists for all six — the census was stale. Five (<code>speculation-guard.js</code>, <code>speculation-judge.js</code>, <code>tasklist-guard.js</code>, <code>devswarm-parent-gate.js</code>, <code>claim-ledger.js</code>) are already registered under <code>Stop</code> in both hooks.json files with platform-neutral payloads (<code>session_id</code>/<code>cwd</code>/<code>transcript_path</code>, no Claude-specific <code>tool_name</code> reads). <code>model-routing-guard.js</code> and <code>output-verify-guard.js</code> are correctly documented Claude-only — Codex has no <code>PreToolUse</code> Agent/Task-tool call (subagent spawn is a separate <code>SubagentStart</code>/<code>SubagentStop</code> event with no pre-spawn <code>tool_input</code> to classify), and the Codex shell tool's <code>PostToolUse</code> <code>tool_response</code> shape is unverified, so wiring either would register a trigger that never fires as intended. Added <code>tests/codex/codex-jev-hooks-parity.test.js</code>, spawning each of the five already-wired hooks with a Codex-shaped Stop payload to prove they run and fail open (exit 0) — closes the runtime-behavior gap <code>tests/hygiene/manifest-drift.test.js</code> (structure-only) does not cover. Added a Claude/Codex parity column to the Jev "All integrations" table in <code>plugins/anti-hall/skills/jev/SKILL.md</code> and the equivalent note in the Codex mirror <code>plugins/anti-hall/codex/skills/anti-hall-jev/SKILL.md</code>.</li> <li><strong><code>openButMarkedArchived</code> (app-state.json) was HOME-GLOBAL and reached the Primary's screenshot ask unfiltered.</strong> <code>syncAppState</code> (<code>scripts/devswarm.js</code>) collects this conflict list across EVERY repo the DevSwarm app knows about, but <code>hooks/devswarm-parent-inbox.js</code>'s <code>uiSyncAsk</code> passed it straight through — the Primary of one repo could be asked for a screenshot about workspaces belonging to entirely different repos. The writer now stamps each entry with <code>repositoryId</code> and its canonical <code>worktreePath</code>. <code>companion/lib/doctor-devswarm.js</code> scopes the list by <code>repositoryId</code> (the current repo resolved via <code>companion/lib/identity.js</code> <code>resolveContext</code> + the app-DB's <code>repositoryForWorktree</code>); the parent-inbox hook, whose ask fires only when the app DB is unreadable, scopes by the <code>repoKey</code> of the entry's <code>worktreePath</code> (see the screenshot-ask fix above). An entry missing the field its reader needs (written by 0.108.0-0.108.2) or an unresolvable current repo fails CLOSED: it is never shown. <code>syncAppState</code> regenerates app-state.json on every supervisor tick, so pre-fix entries self-repair on the next sync — no migration needed.</li> <li><strong><code>primary-session-drift.js</code> picked "newest session" by transcript START time, not last activity.</strong> Three sessions started within ~90s could rank an abandoned transcript (no writes since) as "newer" than the session still actively being written to, because <code>sessionsForWorktree</code> sorted by the transcript's first embedded timestamp. It now sorts by the transcript file's mtime (last activity, already collected for the MAX_SCAN bound), and <code>anchorSessionDrift</code> only surfaces "a newer session exists" when that most-recently-active session is actually more recently active than the CURRENT running session's own last activity.</li> </ul> <h3 id="docs_5">Docs<a class="headerlink" href="#docs_5" title="Permanent link">¶</a></h3> <ul> <li><strong>Every <code>docs/*.md</code> file now linked from both README.md and docs/KB.md.</strong> <code>docs/KB.md</code> was missing <code>KB-claude-code-hooks.md</code>, <code>KB-devswarm-app-db.md</code>, <code>archive/devswarm-layered-recovery-history.md</code>, and two <code>superpowers/</code> design/plan docs; the root <code>README.md</code>'s Documentation section only linked ~5 of 52 tracked docs. Both now link the full set, grouped by topic. New hygiene test <code>tests/hygiene/readme-doc-links.test.js</code> guards this going forward (every git-tracked <code>docs/*.md</code> linked from README.md AND <code>docs/KB.md</code>; every relative docs/ link in README.md resolves).</li> <li><strong>Auto-archive gate (d) blocked a finished lane on its own broadcast backlog.</strong> Field evidence: <code>d-unread (to=197 from=0)</code> on a done, merged, clean, idle child — all 197 were mesh-wide broadcast/FYI rows in the child's own inbox, not mail addressed to it, so gate (d) could never clear and the Primary couldn't ack another partition's cursor either. <code>devswarm-lifecycle.js</code>'s <code>unreadFact</code> was summing <code>w.unread</code> (direct) AND <code>w.broadcastUnread</code> (the shared mesh partition, <code>computeSummary</code>'s own split) into gate (d)'s blocking count. Gate (d) is now DIRECT-only: it blocks only on unread DIRECT rows addressed to the child, or unread rows FROM the child in another partition (e.g. its done report) — never on broadcast/mesh-wide rows. The plan's blocker detail now reports the split, e.g. <code>to_direct=0 to_broadcast=197 from=0</code>. Scoped to gate (d) alone: the separate <code>prune-archived</code> unread check is unchanged.</li> </ul> <h2 id="01082-2026-09-25">0.108.2 (2026-09-25)<a class="headerlink" href="#01082-2026-09-25" title="Permanent link">¶</a></h2> <h3 id="fixes_17">Fixes<a class="headerlink" href="#fixes_17" title="Permanent link">¶</a></h3> <ul> <li><strong>DevSwarm mailbox-wake Monitor instruction assumed <code>persistent: true</code> is always accepted.</strong> Some Claude Code builds' <code>Monitor</code> tool schema is <code>{command, description, timeout_ms, ws}</code> with <code>additionalProperties:false</code>, so passing an unsupported <code>persistent</code> field fails validation outright. The wake-arm instruction (<code>hooks/lib/devswarm-wake.js</code>'s <code>monitorArmLine</code>, <code>skills/update/scripts/update.js</code> and <code>companion/lib/doctor-devswarm.js</code>'s <code>armCmd</code>, and the stale-build re-arm line in <code>companion/lib/devswarm-wake-watch.js</code>) now tells the agent to arm with <code>persistent: true</code> only if its Monitor tool supports that field, otherwise set <code>timeout_ms</code> to its maximum and re-arm on the tool's final/expired event — never running two watchers at once (the watcher's own lock still guards that), with the cron mailbox job staying the fallback either way. <code>docs/KB-claude-monitor-tool.md</code>'s two illustrative Monitor examples carry the same caveat. No Codex-side mirror exists for this instruction (the Codex port has no Monitor tool), so dual-platform parity is unaffected.</li> <li><strong>Read receipts (<code>inbox read-primary</code> / <code>ack-primary</code>) were keyed by the literal caller id, so a Primary registered under two aliases on the SAME worktree — a DevSwarm-native builder-id UUID row and anti-hall's own minted <code>primary-<hash></code> row, the exact pair <code>resolveMeshPartitionIds</code> already widens reads across — could read via one alias and never ack via the other.</strong> <code>writeReadReceipt</code>/<code>readReadReceipt</code> (<code>scripts/devswarm.js</code>) now file the receipt under a new <code>canonicalReceiptId</code> — the SAME identity/alias-family resolution <code>meshPartitionIds</code> already uses (the lowest sorted id in the resolved mesh group), so any alias in the family can look it up. Backward compatible: a lookup always falls back to the literal-id directory too, so a receipt written before this fix (or when family resolution is unavailable) is still found. <code>ackCommand</code> is unaffected — it still names the id the read was addressed to. Also fixed a related latent bug this surfaced: <code>applyReadAckOps</code>'s <code>'own'</code> op branch committed the ack against the outer <code>ack-primary</code> caller's id instead of <code>op.partition</code> (the partition the op was actually computed against at read time) — harmless when read and ack use the same id, but silently acked the WRONG partition's cursor for a cross-alias ack. Now uses <code>op.partition</code>, matching the <code>'sibling'</code> branch's existing pattern. A new forward-migration, <code>fold-read-receipts</code> (<code>companion/lib/migrations.js</code>, applying <code>foldReadReceiptsAllStores</code>), repairs receipts written BEFORE this fix existed — it copies (never moves/deletes) each literal-id-only receipt into its resolved canonical alias-family directory too. Idempotent, fail-open, no-delete, runs from <code>doctor --repair</code> and <code>update</code> like every other default migration.</li> <li><strong>The same cross-alias bug also hit <code>applyReadAckOps</code>'s <code>'nd'</code> op branch (the NDJSON descriptor channel's own cursor, a third namespace alongside the <code>'own'</code>/<code>'sibling'</code> store ops)</strong>: it always committed against the outer <code>ack-primary</code> caller's id instead of <code>op.partition</code> (already recorded at read time by the NDJSON ackOps push site), so a cross-alias ack-primary filed the nd ack under the wrong alias's cursor. Fixed to mirror the <code>'own'</code> branch (<code>ndPartitionId = op.partition != null ? op.partition : id</code>); falls back to <code>id</code> for a legacy op with no <code>partition</code> recorded. <code>meta.callerId</code> keeps the real caller for audit/logging — only <code>commitNdAck</code>'s partition-key argument changed.</li> <li><strong>Default ceiling removed: <code>autoHandover.maxTokens</code> now defaults to <code>0</code> (off) — the trigger fires at 85% of the session's REAL context window, not a fixed 170000-token assumption.</strong> <code>autoHandover.pct</code> (default 85%) is already measured against this session's actual context window (<code>hooks/lib/context-pct.js</code>: the statusline's real <code>context_window</code>, a Codex rollout's <code>model_context_window</code>, or an estimate), so it already scales correctly to a 1M+ window; a fixed absolute <code>maxTokens</code> default fired an EARLY, unrelated handover on a genuinely large window instead (the exact G3 gap this ceiling was meant to close, re-opened by hardcoding its value — see <code>docs/KB-handover-research.md</code>). <code>maxTokens</code> is now opt-in only (<code>ANTIHALL_AUTO_HANDOVER_MAX_TOKENS</code>, <code>settings.json</code> <code>autoHandover.maxTokens</code>, or <code>scripts/auto-handover-config.js set maxTokens <n></code> — all still work exactly as before) for a user who wants an absolute token floor regardless of window size. <code>hooks/lib/settings-schema.js</code> and <code>hooks/lib/auto-handover-config.js</code>'s <code>DEFAULT_MAX_TOKENS</code> both moved from <code>170000</code> to <code>0</code>. Also fixed the fire-directive wording (<code>hooks/lib/auto-handover-text.js</code>): a token-ceiling fire (only reachable when a ceiling is opted in) now leads with "CONTEXT: ~NK tokens ≥ your configured autoHandover.maxTokens ceiling (NK)" instead of the pct-framed "CONTEXT AT ~N%" line, which read confusingly small/unrelated next to "REQUIRED" when the session was nowhere near the pct threshold.</li> </ul> <h2 id="01081-2026-09-25">0.108.1 (2026-09-25)<a class="headerlink" href="#01081-2026-09-25" title="Permanent link">¶</a></h2> <h3 id="fixes_18">Fixes<a class="headerlink" href="#fixes_18" title="Permanent link">¶</a></h3> <ul> <li><strong>jev report: shadow-mode yield was invisible to KEEP/REMOVE verdicts.</strong> A shadow-mode row's <code>changed</code> field is always <code>null</code> by construction (<code>jev-assist.js</code>'s <code>finalize()</code> only applies a change when <code>mode==='on'</code>), so reading only <code>row.changed</code> meant a shadow integration's changed-decision rate was always 0 — it could hit REMOVE at >=200 calls no matter how good Jev's shadow answers actually were. <code>finalize()</code> now also logs <code>wouldChange</code> — the same trust-rule outcome computed WITHOUT the mode gate — and <code>jev-report.js</code> reads it for shadow/off rows. Label-only integrations (a <code>choice</code> classifier with no boolean baseline, e.g. <code>newRequest</code>) are excluded from changed-rate entirely, regardless of mode.</li> <li><strong>jev report: REMOVE/KEEP could fire with zero labelled outcomes.</strong> A changed-rate<1% or a high failure rate alone used to earn REMOVE with no labelled outcomes behind it (observed: 3 changed / 304 fresh hit REMOVE via changed-rate alone). KEEP and REMOVE now both require a labelled sample (tp+fp, human+auto) of at least <code>MIN_LABELED_FOR_VERDICT</code> (20); below that, <code>REVIEW (needs labels: n/20)</code>. Changed-rate<1% is now a low-yield note only, never a REMOVE trigger by itself. The label-only guard now runs BEFORE REMOVE/KEEP are evaluated (it used to run after REMOVE, so it could never actually fire once REMOVE's changed-rate<1% path caught a label-only row first).</li> <li><strong>jev report: <code>--since</code>/<code>--until</code>/<code>--exclude-window</code></strong> exclude a known-bad run from the report by <code>ts</code>, without editing <code>jev-assist.ndjson</code> directly — e.g. excluding an accidental <code>supervisorBlockerLabel</code> run.</li> <li><strong>DevSwarm: <code>acquireLock</code> TOCTOU let two sessions both adopt the same closed Primary seat.</strong> <code>companion/lib/recovery.js</code>'s per-id advisory lock published its holder via <code>openSync(p, 'wx')</code> immediately followed by a SEPARATE <code>writeSync(fd, ...)</code> call; a second <code>acquireLock</code> that hit EEXIST in the window between those two syscalls could read the lock file while it was still empty, misread the unparseable/null holder as unconditionally stale, and steal a lock its rightful, live, milliseconds-old owner was still writing — breaking the "ONE Primary per project" invariant under real concurrent adoption. Fixed by publishing the lock via write-then-link (full content written to a private temp file, then made visible atomically with <code>linkSync</code>), which can no longer expose partial content. This was the real root cause of the <code>devswarm-primary-seat.test.js</code> "concurrent adoption" test flaking once on macOS CI; that test was also hardened (barrier + bounded retry instead of a bare timing-dependent race) as defense in depth.</li> <li><strong>DevSwarm: two code-review follow-ups to the <code>acquireLock</code> fix above, same release.</strong></li> <li><em>Regression the write-then-link fix itself introduced:</em> a <code>linkSync</code> failure for any reason OTHER than EEXIST (EPERM/ENOTSUP/EXDEV/ENOSYS — e.g. <code>~/.anti-hall</code> mounted on SMB/exFAT, where hardlinks don't work even though a normal local disk supports them) used to be treated as fail-open/no-lock, meaning EVERY acquire on such a filesystem returned null and <code>withIdLock</code> refused every DevSwarm mutation it gates, permanently. <code>acquireLock</code> now falls back to the pre-fix create-then-write (<code>openSync(p, 'wx')</code> + <code>writeSync</code>) for that one attempt instead.</li> <li><em>Pre-existing, not introduced by the fix above:</em> the stale-holder RECLAIM step used a blind <code>unlinkSync(p)</code>, which deletes WHATEVER currently sits at that path, not specifically the dead/stale holder just read. Two callers that both read the same dead holder concurrently could both decide to steal it; whichever unlinked second deleted the first's brand-new, legitimately-published lock, and both then "won". Fixed by moving the file aside atomically with <code>renameSync</code> first, re-reading the moved copy, and only discarding it if its token still matches the one read before the rename — a token mismatch (a real racer's fresh lock) restores it and respects it instead.</li> <li><code>doctor --repair</code> now sweeps stale <code>*.lock.tmp-*</code> / <code>*.lock.reap-*</code> scratch files (both fixes' own temp artifacts, normally self-cleaning) under <code>devswarm/locks/</code> older than 15 minutes — a belt-and-suspenders backstop for a crash/kill mid-acquire, never touching an actual lock file.</li> </ul> <h2 id="01080-2026-09-24">0.108.0 (2026-09-24)<a class="headerlink" href="#01080-2026-09-24" title="Permanent link">¶</a></h2> <blockquote> <p><strong>UPGRADE NOTE — coming from 0.107.x or earlier:</strong> run <code>claude plugin update anti-hall@anti-hall</code> once, then restart Claude Code. The 0.107.x <code>update.js</code> has no post-pull re-exec, so it cannot run 0.108.0's harness registration; without this step Claude Code keeps loading the old build. From 0.108.0 on, <code>update</code> does it for you. The Codex port loads from the marketplace clone and needs no extra step.</p> </blockquote> <h3 id="new-features">New features<a class="headerlink" href="#new-features" title="Permanent link">¶</a></h3> <ul> <li><strong>Auto-handover (on by default).</strong> When the main agent's context first crosses <code>autoHandover.pct</code> (85%), <code>hooks/auto-handover.js</code> tells it, without asking first, to write an anti-hall handover itself, tell the user and list the saved paths, and urge <code>/compact</code> or <code>/clear</code>. After that: a short reminder every <code>nagStepPct</code> (5) further points, and a Stop-time reminder at a quiet pause (no open tasks, no recent spawn) at most once per <code>nagQuietMin</code> (15) minutes (<code>hooks/auto-handover-pause-nag.js</code>). Context % comes from the statusline's real <code>context_window</code> figure (persisted by <code>statusline/phase-bar.js</code>), a Codex rollout's <code>model_context_window</code>, or a transcript estimate whose window is the env override, the session's last-seen statusline window, or "inferred 1M" once usage passes 200k. With a genuinely unknown window it sends one soft advisory instead of the mandatory directive. <code>ANTIHALL_AUTO_HANDOVER_PCT</code> overrides the threshold (<code>0</code> = off). Shared with the Codex port.</li> <li><strong>Absolute token ceiling</strong> <code>autoHandover.maxTokens</code> (default 170000 = 85% of a 200k window; env <code>ANTIHALL_AUTO_HANDOVER_MAX_TOKENS</code>; CLI <code>max-tokens <n></code>; <code>0</code> = off): fires on whichever comes first, so a 1M session hands over at ~170k instead of ~850k. The token count is real, so a ceiling crossing gets the full directive even with an unknown window. The directive prints an exact <code>/compact focus: …</code> line; the handover skill ships a CLAUDE.md "Compact Instructions" snippet.</li> <li><strong>Stop-side fire:</strong> a long autonomous turn that never reaches a new prompt gets the directive once at a Stop (shared latch, never while <code>stop_hook_active</code>).</li> <li><strong>PreCompact safety net</strong> (<code>hooks/precompact-snapshot.js</code>, Claude + Codex): before every compaction a mechanical <code>.anti-hall/handovers/<date>/<session>/PRECOMPACT-<n>.md</code> (pwd, git state, task list from the transcript, last 10 user messages verbatim, newest handover pointer). Always exits 0 and prints nothing, so it never blocks compaction.</li> <li><strong>Resume:</strong> <code>handover-resume.js</code> names the snapshot, adds git facts measured at resume (HEAD, commits since the handover, dirty files) and asks the agent to read back goal, next action and session rules before acting. All messages are platform-aware (Claude: <code>/anti-hall:handover</code>, <code>/compact focus: …</code>; Codex: <code>anti-hall-handover</code>, <code>/compact</code> or <code>/new</code>, <code>AGENTS.md</code>).</li> <li><strong>Handover content:</strong> a verbatim "Session rules" slot; seq N>1 carries forward rules and verified/not-verified rows verbatim. Research: <code>docs/KB-handover-research.md</code>.</li> <li><strong>One settings store for everything.</strong> <code>~/.anti-hall/settings.json</code>, with a declarative schema (<code>hooks/lib/settings-schema.js</code>) and API (<code>hooks/lib/settings.js</code>); front ends <code>/anti-hall:settings</code> (Codex: <code>anti-hall-settings</code>) and <code>scripts/settings.js show|get|set|reset [--json]</code>. Precedence: env → settings.json → <code>/config</code> (a <code>plugin.json</code> <code>userConfig</code> value that differs from its manifest default) → legacy file (e.g. <code>jev.json</code>; outranks <code>/config</code> until the one-time migration is stamped) → default. Every setting is wired to the code that reads it, including auto-archive, retention and the Jev budget/prices/audit settings; resolvers that take an <code>env</code> derive <code>home</code> from that env (<code>settings.getWithEnv</code>). Dotted keys can be written flat (<code>"autoArchive.mode"</code>) or nested. A corrupt settings.json is backed up before any write. The old <code>auto-handover-config</code> skill is folded into <code>/anti-hall:settings</code>; <code>scripts/auto-handover-config.js</code> stays as a CLI alias.</li> <li><strong>Repairs run on every reload.</strong> <code>hooks/repair-on-reload.js</code> (SessionStart + a UserPromptSubmit fallback, since <code>/reload-plugins</code> is not shown to re-fire SessionStart) starts one detached <code>doctor.js --repair --migrations-only</code> whenever a repair is not yet stamped for the running version. It runs only the data migrations and store repairs; statusline, Codex hooks and the supervisor are never installed by a reload (that stays behind a user-typed <code>doctor --repair</code>). The legacy-state and GSD migrations do write the project's <code>.anti-hall/history/legacy/</code>, and GSD <code>.planning/</code> files are deleted once their copy is verified. A stamp at the running version or newer counts as done (the newest cached doctor stamps its own version), runs are at most hourly per running version (a failed spawn never starts the wait), only the newest 5 logs are kept, and the child runs at nice 19. Lock-guarded, never blocks, no-op when nothing is pending. <code>ANTIHALL_REPAIR_ON_RELOAD=off</code> disables it.</li> <li><strong>Version alerts say update or reload.</strong> If the plugin-cache mirror already holds a newer version than the running one, the agent is told to reload (with that version's changelog headline); if the remote is newer, to update and then reload. Both are explicit "tell the user now" directives, once per session per case, on Claude and Codex.</li> <li><strong>Jev metrics.</strong> Real per-call cost (gateway-reported, else tokens × <code>jev.prices</code>; a cache hit costs $0; never invented); <code>jev-report.js</code> shows cost windows (<code>--window 24h|7d</code>), precision from <code>label <hash> tp|fp</code> (plus derived AUTO labels), yield, cost efficiency, overhead and a one-line headline per integration; the Vercel AI Gateway credit balance (report/status time, 15-min cache). Opt-in budget watch (<code>jev.budget.mode=watch</code> + <code>usdPerDay</code>/<code>usdPerWeek</code>/<code>minCreditUsd</code>) only warns and never disables Jev. Opt-in audit snippets (<code>jev.audit.snippets</code>) keep a redacted ~200-char snippet for decisions Jev changed; <code>jev-report.js prune-audit --days N</code> trims them.</li> <li><strong>Jev: five more integrations</strong> (settings <code>jev.integrations.<id></code>, default <code>shadow</code>; <code>jev-setup.js mode <id> on|shadow|off</code>):</li> <li><code>gitGuardSelfCredit</code> (git-guard): paraphrased AI self-credit in commit/PR text the regex misses ("written with help from Claude"); <code>askSync</code> 1.5 s, add-block only — it can never relax git-guard.</li> <li><code>parentGateQuestion</code> (parent gate): an unread child message triage already labelled a question counts as awaiting a reply; cache-only, no network.</li> <li><code>supervisorBlockerLabel</code> (supervisor report): waiting-on-parent vs wedged from cached triage labels; advisory, after poke/escalate already ran.</li> <li><code>tasklistTrivial</code> (tasklist-guard) and <code>codexNudgeSubstantial</code> (codex-nudge): is the nudge about to fire worth it? Logged fire-and-forget in shadow; in <code>on</code> asked synchronously (1.5 s cap, fail-open to the nudge) and a confident "trivial" skips it.</li> <li><strong>Jev report:</strong> <code>--by project|session</code> and <code>--project <name></code> (rows carry the cwd basename and session id); <code>--weekly</code> compact 7-day scorecard; the classifier-call latency (p50/p95 ms) and the triage reply turnaround are separate figures, locked by tests. <strong>Weekly notice:</strong> <code>hooks/jev-weekly-scorecard.js</code> (SessionStart, Claude + Codex, Jev enabled only, never in a DevSwarm child) names, at most once a week, one integration whose report says KEEP or REMOVE but whose mode does not match yet; it never changes a mode (<code>jev.weeklyNotice</code>, default true). <code>jev-setup.js status</code> lists every integration.</li> <li><strong>DevSwarm: the app database is the ground truth.</strong> One read-only snapshot of the DevSwarm app's own SQLite DB (<code>companion/lib/devswarm-app-db.js</code>) feeds the per-turn table, <code>roster</code>, identity, <code>doctor</code> and the supervisor: archived/open state, full titles (app renames propagate), sidebar order, <code>[pinned]</code>, <code>[on screen]</code> (nags for the focused workspace are suppressed for 2 min), PR state (<code>PR #N merged, checks failed</code>, only when synced after the branch last moved), and <code>[⚠ brief not delivered]</code>/<code>[⚠ brief withheld]</code>. It never reads message bodies, brief text or credential tables. The Primary's anchor self-ack and <code>register-primary</code> takeover follow the app's session map when Claude's own transcript confirms it. A pinned schema makes <code>doctor</code> warn when the app drops a column anti-hall reads. Evidence per field: <code>docs/KB-devswarm-app-db.md</code>.</li> <li><strong>DevSwarm: supervisor app sync every tick.</strong> Archived markers for workspaces the app archived or deleted (never a delete), title refresh, <code>app-state.json</code> (session map, drift, pending app deletions, a report-only message-gap cross-check); about 0.1 s per sweep. <code>devswarm.js app-state [--json]</code> and <code>app-sync</code>; <code>ANTIHALL_DEVSWARM_APP_SYNC=0</code> disables it.</li> <li><strong>DevSwarm: capability gate.</strong> Every DevSwarm surface anti-hall touches — each <code>hivecontrol workspace</code> verb, every app-DB table and column, and the app files it stats — is checked by minimum version and runtime detection (<code>--help</code> parsing, <code>PRAGMA table_info</code>, existence), cached per hivecontrol build (<code>companion/lib/devswarm-capabilities.js</code>). A missing surface puts its feature to sleep and <code>doctor</code> names it; without DevSwarm everything sleeps silently. The gate is default-deny: a <code>hivecontrol</code> invocation that maps to no registered capability is refused with a dormant reason (never spawned, never thrown); only explicit <code>ungated</code> entries (<code>--version</code>, <code>--help</code>, <code>workspace --help</code>) skip the check.</li> <li><strong>DevSwarm: auto-archive of done workspaces</strong> (default on, DevSwarm ≥ 2.5.3). The supervisor archives a child only when all are proven: finish gates set, branch merged (git ancestry or the app's PR row), clean worktree, no unread either way, not the Primary, not viewed in the app for 10 minutes, idle ≥ <code>idleMin</code>. It always names the workspace id, tells the Primary with an undo hint, and replaces the archive-ready nag for those workspaces. <code>devswarm.js auto-archive</code> shows the plan.</li> <li><strong>DevSwarm: owner-approved prune.</strong> <code>devswarm.js prune-archived --older-than <days></code> lists archived workspaces with evidence and a 15-minute plan nonce; deletion runs only with <code>--confirm-ids <exact ids> --plan <nonce></code> from an interactive caller. Each row is re-checked, logged (<code>logs/devswarm-prune.ndjson</code>) and tombstoned; no store rows are deleted. A hygiene test keeps every automated path away from it.</li> <li><strong>DevSwarm: message retention.</strong> The supervisor archives old message bodies to <code>~/.anti-hall/devswarm/archive/<store>/<yyyy-mm>.ndjson.gz</code>, then clears them (rows, read positions, hashes and seq numbers stay, so unread counts and gates do not change). Only bodies older than <code>retention.days</code> (30) that every reader has read, outside the newest <code>keepPerPartition</code> (200) rows and not an open question; a store over <code>maxStoreMB</code> (100) is pruned oldest first; <code>VACUUM</code> reclaims space; the archive is never evicted unless you set <code>archiveMaxMB</code> (default 0 = no cap; <code>doctor</code> warns past 500 MB). <strong>Pruning is automatic:</strong> the first run on a machine only writes a dry-run report (<code>~/.anti-hall/devswarm/retention-dry-run.json</code>, flagged by <code>doctor</code>); after that, sweeps prune for real. To keep every body, set <code>devswarm.retention.days</code> to 0 before then. <code>devswarm.js retention status | run [--dry-run] | restore --store X --month yyyy-mm</code>.</li> <li><strong>DevSwarm: screenshot sync.</strong> When the app DB is unreadable or disagrees, the per-turn hook asks once per session for a sidebar screenshot; <code>devswarm.js sync-ui</code> plans it against the app DB and the owner confirms. It archives only what the app DB proves, never unarchives by itself, never deletes.</li> <li><strong>Operator guide.</strong> <code>/anti-hall:system-briefing</code> (and the Codex mirror) is now the agent-facing guide: glossary, hard rules, every skill, CLI verb and setting with its default, plus the live inventory. The SessionStart foundation points agents at it.</li> </ul> <h3 id="fixes_19">Fixes<a class="headerlink" href="#fixes_19" title="Permanent link">¶</a></h3> <ul> <li><strong>P0: tests could register real launchd jobs.</strong> <code>doctor-repair</code>'s installer spawn passed a caller's partial env (<code>env: {}</code>) straight to <code>spawnSync</code>, which replaces the child environment: the installer ran with no <code>HOME</code> (so the real home) and no test marker, and loaded a real supervisor/ingest job. Installer spawns (doctor, <code>devswarm</code> self-heal, <code>update</code>) now merge onto the parent env and always carry the test markers; the supervisor, ingest and reaper installers route <code>launchctl</code>/<code>systemctl</code>/<code>crontab</code> through one seam that refuses them under a test; the statusline and Codex installers refuse user config outside a temp dir under a test. Test helpers set <code>ANTIHALL_TEST_ISOLATION=1</code>, and a hygiene test fails any hand-built env passed to a doctor/installer spawn without it. The reaper installer gained the same temp-HOME guard as supervisor/ingest (<code>ANTIHALL_REAPER_ALLOW_TMP_HOME=1</code> to opt out).</li> <li><strong>P0: <code>update.js</code> ran the old version's post-pull stages.</strong> After pulling a newer version it now re-execs the freshly pulled <code>update.js --post-pull-only</code>, so stages that exist only in the new version run in the same update; any failure falls back to the local results with a note. <code>repair-on-reload</code> likewise uses the newest cached <code>doctor.js</code>.</li> <li><strong>P0: the harness never loaded an updated build.</strong> Claude Code loads anti-hall from <code>installed_plugins.json</code>'s <code>installPath</code>, which <code>update</code> never touched. <code>update</code> now runs <code>claude plugin update anti-hall@anti-hall</code> whenever that record is behind (reported as <code>harnessRegistered</code>; never auto-accepts a confirmation — it prints the exact command instead), and then says <strong>RESTART</strong> Claude Code (a registry change is not picked up by <code>/reload-plugins</code>). From 0.108.0 on, the post-pull re-exec means the newly pulled <code>update.js</code> performs the registration even when a later release changes it; a registration done by the parent is never hidden by the child's no-op. This does NOT cover the hop from 0.107.x or earlier (that <code>update.js</code> has no re-exec) — see the upgrade note above. When <code>claude</code> is not on <code>PATH</code> it retries with <code>CLAUDE_CODE_EXECPATH</code> (the running Claude Code binary); if registration still fails, <code>action</code> gives the exact manual command plus a restart instead of <code>/reload-plugins</code>. The re-exec child's timeout is now the post-pull budget + the registration timeout + a 60 s margin (was a fixed 120 s, below the default 90 s budget + 20 s registration). <code>doctor</code> warns when <code>installed_plugins.json</code> lags the newest cache/marketplace version.</li> <li><strong>Identity: a child's messages were labelled as a Primary.</strong> Every sender's <code>from</code> was <code>primary-<worktree hash></code>, children included, so a resumed Primary could see another "Primary" and stand down. <code>send</code>, <code>archive-request</code> and the <code>merge</code> broadcast now label a child worktree with its workspace id; only the Primary checkout (the app's <code>primary</code> builder, else the main worktree) sends as <code>primary-<hash></code>. <code>register-primary</code> refuses in a child worktree (<code>not-primary-checkout</code>; <code>--force</code> overrides). <code>heartbeat</code>, <code>register</code> and <code>ensure</code> refuse a child worktree's label (<code>child-label-id</code>, naming the real id), so a placeholder row can no longer look like a live twin. Questions and broadcasts sent under an old child label are attributed to the child (<code>mesh read</code> keeps the stored <code>fromLabel</code>).</li> <li><strong>Identity: the Primary anchor follows the running session</strong> after a <code>/clear</code>: <code>inbox tick</code>/<code>heartbeat</code> re-point it (Primary checkout only, never while the recorded session is still running), and the briefing names a newer session on the Primary's worktree (read-only).</li> <li><strong>New: the Primary seat survives a session change.</strong> At SessionStart a new session in the Primary checkout adopts <code>primary-<hash></code> (same partitions and cursors) when the holder has closed, and is told which handover to read. While the holder is still live nothing is taken over; the conflicting session is warned and its <code>send</code>, <code>inbox ack</code>, <code>ack-primary</code>, <code>spawn</code> and <code>merge</code> are refused (<code>primary-seat-conflict</code>) until <code>devswarm.js primary takeover</code>; <code>primary status</code> shows the seat. The refusal sits on the cursor-write doors themselves, so no path (<code>drain-primary-legacy</code>, <code>--legacy-ack-now</code>, <code>mesh read</code>, <code>roster --ack</code>, …) lets a non-holder advance the Primary's read positions. Liveness: harness session pid, then the app's active terminal, then heartbeat age; unknown → warn only.</li> <li><strong>Stale-build children looked neglected.</strong> <code>heartbeat</code>/<code>inbox tick</code> stamp the running anti-hall version; the roster, parent-inbox table and nag, and <code>doctor</code> show <code>stale anti-hall <v>: restart this session</code> instead of <code>not-draining</code> when a workspace runs an older build than this machine has (an app-archived row stays plain <code>archived</code>). <code>devswarm-wake-watch.js</code> exits with one line when a newer version is installed.</li> <li><strong>The parent inbox flagged a just-sent message within seconds.</strong> The unread-only nag now waits for a grace window (<code>devswarm.inboxGraceSec</code>, 120 s) unless the child heartbeats first; stuck/not-draining signals are unaffected.</li> <li><strong>Three directories grew without bound</strong> (measured 133 MB / 34k files, 62 MB, 27 MB): per-session <code>devswarm/child-gate/</code> state is swept after <code>devswarm.childGateRetentionDays</code> (14); that sweep and the existing <code>reaped/</code> sweep now also run from a periodic supervisor housekeeping sweep (<code>devswarm.housekeepingSweep</code>, every <code>housekeepingSweepSec</code> = 1 h); the supervisor rotates its own log at <code>supervisorLogRotateBytes</code> (10 MB, 2 generations).</li> <li><strong>Version alert missed multi-release days</strong>: the 24 h cache TTL served a stale <code>latest</code> all day; now 2 h, plus the network-free plugin-cache check above.</li> <li><strong><code>update.js --check</code> trusted a lagging <code>installed_plugins.json</code></strong>; it now takes the higher of that and the newest cache dir and says when to <code>/reload-plugins</code>.</li> <li><strong>Three <code>doctor</code> false failures</strong>: the parent-inbox self-test resolved its repoKey from the plugin dir (and a skip counted as a failure); a <code>register-primary</code> Primary row was poked/escalated like a child (the supervisor now skips Primary rows); doctor's own context-footprint probe no longer triggers <code>repair-on-reload</code>. doctor's version-alert self-test matches the new directive wording.</li> <li><strong>Parent-inbox</strong>: the app's own primary builder row on the Primary's checkout (proven by <code>builderType</code>) folds into the own-unread line instead of a false child nag; an app-archived row with a stored not-draining flag renders plain archived.</li> <li><strong><code>jev report</code></strong>: "agreement" compared Jev to a trust constant — it now uses the caller's own heuristic verdict (<code>compare</code>) and reports <code>n/a</code> without one; changed decisions and cost are counted once per content hash (a fresh call plus its cache hits); TypeSafe's <code>usage.input_tokens</code>/<code>output_tokens</code> are read.</li> <li><strong><code>settings.js show --all</code></strong> dropped four Jev advanced keys and listed <code>budget.*</code> twice (headline/advanced split by position); object settings now render as JSON.</li> <li><strong>command-guard allowed read-only <code>--help</code>/<code>-h</code></strong>: <code>hivecontrol workspace read-messages --help</code>, <code>monitor -h</code> and <code>message-child --help</code> only print usage, so a segment carrying a bare <code>--help</code>/<code>-h</code> token is no longer blocked as a destructive read or native send; a real call chained on another segment still blocks.</li> <li><strong>The supervisor installer could register a real unit for a test or scratch HOME.</strong> <code>install-devswarm-supervisor.js</code> now has the same guards as the ingest and reaper installers: forced dry-run under <code>node --test</code> or with HOME under a temp root (<code>ANTIHALL_SUPERVISOR_ALLOW_TMP_HOME=1</code> for a deliberate one). The shared temp-root check also recognises macOS <code>/var/folders</code> when <code>TMPDIR</code> is not set.</li> <li><strong>Settings writes could lose each other's keys.</strong> <code>set</code>/<code>reset</code> read-modify-write the whole <code>settings.json</code>; they now hold <code>settings.json.lock</code> (bounded wait, a dead or stale holder is reclaimed, busy → <code>{ok:false, lockBusy}</code>, never throws). Known limitation, documented: a <code>/config</code> value equal to its manifest default counts as unset.</li> <li><strong>Jev audit snippets redact more:</strong> AWS key ids, any <code>*KEY</code>/<code>*SECRET</code>/<code>*PASSWORD</code>/ <code>*TOKEN</code>= identifier, JWTs, <code>user:pass@</code> URLs and PEM blocks.</li> <li><strong>edit-guard's handover redirect</strong> now names <code>.anti-hall/handovers/**</code> as the exempt place to write.</li> <li><strong><code>update.js</code> split-store summary</strong> now states how many already-handled rows were re-delivered as unread, and why (no behaviour change).</li> <li><strong>DevSwarm 2.5.3</strong>: <code>workspace archive|delete [idOrBranch]</code> default to the CURRENT workspace and have no <code>--yes</code>; anti-hall now always passes the explicit workspace id and makes no call without one. Help fixtures are the real 2.5.3 text.</li> <li><strong>Tests could touch the real home.</strong> <code>update.js</code> <code>runUpdate</code>, doctor's <code>runRepairs</code> and <code>runMigrations</code> now refuse (throw) under <code>node --test</code> when their home is the real user home (<code>companion/lib/test-home-guard.js</code>); this caught two more test files calling <code>runUpdate</code> against the real home, now isolated.</li> <li><strong>Docs</strong>: full <code>docs/README.md</code> index and link check (<code>tests/hygiene/docs-links.test.js</code>); <code>llms.txt</code> is now the condensed catalog of every hook, skill, script and setting; <code>tests/hygiene/docs-coverage.test.js</code> fails the build when anything shipped is undocumented. GUIDE's 560-line DevSwarm version history moved to <code>docs/archive/devswarm-layered-recovery-history.md</code>; GUIDE keeps a current-state summary.</li> </ul> <h3 id="data-repairs-update--doctor---repair-idempotent-never-delete">Data repairs (update + <code>doctor --repair</code>, idempotent, never delete)<a class="headerlink" href="#data-repairs-update--doctor---repair-idempotent-never-delete" title="Permanent link">¶</a></h3> <ul> <li><strong><code>repair-child-sender-labels</code></strong> (update + <code>doctor --repair</code>): maps each child's old <code>primary-<hash></code> label (with or without a registry row) to its id in <code>sender-aliases.json</code>; a non-live phantom <code>primary-<childhash></code> row has its unread mail forwarded to the child's partition, then is retired with a redirect and an archived marker (a live one is retried next run). Stored messages are not rewritten; nothing is deleted.</li> <li><strong>Settings migration</strong>: legacy config (<code>~/.anti-hall/jev.json</code>, including nested <code>budget</code>/<code>audit</code> keys) is forward-migrated into <code>settings.json</code> once per version; the legacy file is never deleted, and file-only <code>prices</code> keeps being read from it.</li> <li><strong>Escalated Primary verdicts</strong>: <code>doctor --repair</code> clears an already-stale <code>escalated</code> verdict left on a <code>register-primary</code> Primary row.</li> <li><strong>App-archived/deleted workspaces</strong>: the supervisor's app sync writes the archived marker for workspaces the app archived or deleted (marker only; nothing is deleted).</li> <li><strong>State retention</strong>: <code>doctor --repair</code> sweeps auto-handover latches and context-% state older than 30 days (<code>ANTIHALL_AUTO_HANDOVER_STATE_RETENTION_DAYS</code>, <code>ANTIHALL_CONTEXT_PCT_STATE_RETENTION_DAYS</code>).</li> <li><strong>Repairs now also run on reload</strong> (<code>repair-on-reload.js</code>), not only on update.</li> </ul> <h3 id="owner-decisions">Owner decisions<a class="headerlink" href="#owner-decisions" title="Permanent link">¶</a></h3> <ul> <li>Auto-handover is on by default at 85% or 170k tokens, whichever comes first.</li> <li>Auto-archive (<code>devswarm.autoArchive.mode</code>) defaults to <code>on</code> (<code>dry-run</code>/<code>off</code> available).</li> <li>Retention defaults: 30 days, 100 MB per store, 200 newest per partition kept, archive on, no archive cap (<code>archiveMaxMB</code> 0 — archived bodies are evicted only when you set one).</li> <li>Spawn titles are no longer truncated (the 60-char cap is gone).</li> <li>Deleting DevSwarm workspaces is never automated: only <code>prune-archived</code> with an owner-approved exact list.</li> </ul> <h2 id="01071-2026-09-24">0.107.1 (2026-09-24)<a class="headerlink" href="#01071-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: phantom "N unread" on the Primary's own mailbox.</strong> When a store's reader-position import ran without the legacy inbox cursor, it seeded the NDJSON floor at 0, so a fully read legacy inbox counted as unread again. The NDJSON floor is now never lower than the legacy descriptor cursor. The <code>repair-reader-floors</code> repair also raises import-seeded rows and the floor to that cursor. It only raises values and deletes nothing, and running it twice changes nothing.</li> <li><strong>Fixed: <code>update</code> skipped the reader-floor repair outside a DevSwarm session.</strong> It is a store repair, so it now runs on every update; <code>doctor --repair</code> runs the same pass.</li> <li><strong>Fixed: a resumed Primary disagreed with itself by one message.</strong> A resumed session gets a new session id, but the anchor row kept the old one. So the Primary's own builder partition looked like a foreign live sibling and was never acked. The anchor now counts as the caller's own when its recorded session has no running process and the caller's session does, which is the same rule <code>register-primary</code> uses. Existing watermarks heal on the next read, with nothing lost.</li> <li><strong>Fixed: workspaces archived in the DevSwarm app kept nagging.</strong> <code>hivecontrol workspace list all</code> lists archived builders too, so detecting them by absence never fired. Archive state now comes from the app's own database: a read-only query that falls back to the old path if the file or schema is missing, with an <code>ANTIHALL_DEVSWARM_APP_DB</code> override (<code>off</code> disables it). Archived rows and twin rows on an archived worktree no longer nag or raise archive-ready prompts, and they are still listed in the table. A new <code>mark-app-archived</code> repair writes the archived marker for those workspaces. It never overwrites an existing marker and never touches descriptors.</li> <li><strong>Safety: <code>reconcile-active</code> no longer archives a workspace just because it is missing from <code>--active</code>.</strong> A scrolled or partial list would archive live workspaces. A candidate is now archived only when the DevSwarm app database confirms it is archived. Anything else is kept and listed in <code>keptNotArchivedInApp</code>, and an unreadable app DB archives nothing. The skill docs no longer suggest a roster screenshot as the source.</li> </ul> <h2 id="01070-2026-09-24">0.107.0 (2026-09-24)<a class="headerlink" href="#01070-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: phantom unread.</strong> The v0.106.0 reader-position import had declared every live session on the machine a reader of every partition, and the DevSwarm gate double-counted the Primary's own mailbox against that; an updater repair now retires the affected floors on already-updated stores.</li> <li><strong>Fixed: one mailbox for a session's two partitions</strong>, with a repair for stores where the split had already duplicated acks.</li> <li><strong>Changed: one persisted storage backend per store</strong>, plus a merge of any store that had already split before that marker fix shipped (so a child on the non-chosen backend could not see broadcasts) — <code>doctor</code> (plain) now reports split stores found; <code>--repair</code> / <code>update</code> merge them. Children now wake on broadcasts.</li> <li><strong>Changed: no more archive-ready nag when only your own request is outstanding</strong>, plus an ignore list; spawn titles are applied to new workspaces.</li> <li><strong>Fixed: an app-archived workspace stops reading as active</strong>; a dead archive-ready workspace is now archived directly instead of nagged forever.</li> <li><strong>Changed: per-turn hook blocks (urgent/stale/limit-conservation) are re-emitted only when their content changes</strong>, not every turn. The archive-ready reminder keeps its own cooldown instead (it must re-surface identical text once that cooldown elapses, which content-change dedupe would otherwise suppress).</li> <li><strong>Fixed: tasklist-guard no longer re-nags after a real progress update</strong> — a scratchpad write never counts as project work, but a transcript-observed progress write always does, regardless of file mtime.</li> <li><strong>Changed: edit-guard allows the session's own tmpdir scratchpad.</strong></li> <li><strong>Changed: output-verify is scoped to test runners.</strong></li> <li><strong>Added: Jev</strong>, a shared LLM-assist layer with shadow metrics and a <code>jev</code> skill ("activate jev") that asks for a Vercel AI Gateway or TypeSafe key, installs it, enables Jev, and runs a test call (also: status / disable / mode). Ships shadow integrations for the claim ledger, merge-gate hedges, new-request classification, and output-verify (scoped to test runners), speculation-outcome tracking, model-routing in relax-only (shadow) mode, and a <code>jev report</code> command (KEEP / REVIEW / REMOVE) over the metrics log.</li> </ul> <h2 id="01060-2026-09-24">0.106.0 (2026-09-24)<a class="headerlink" href="#01060-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li><strong>Changed: DevSwarm reader positions live in one SQLite table</strong> with one unread count and one writer. This retires the phantom-unread and floor-drift class of defects, and closes several pre-existing message-loss paths: a rehome/fold race, the orphan forward, a torn journal import, and a lock steal.</li> <li><strong>Changed: row state comes from one reducer</strong>; every surface derives it the same way.</li> <li><strong>Changed: <code>doctor</code> is read-only by default</strong> — repairs run only with <code>--repair</code> / <code>--fix</code>. One migration registry now owns every one-time sweep and never stamps a version done while work is left (errors, deferred stores, pending rows, forward failures).</li> <li><strong>Added: advisory Jev triage labels</strong> for mesh messages (opt-in via <code>~/.anti-hall/jev.json</code>); labels are advisory only and never change delivery or any unread count.</li> <li><strong>Added: crash-safe delivery through a read WAL.</strong> Every destructive native read is fsynced before it is ingested and replayed until closed, closing the pull-crash and ingest lock-contention losses. The remaining native-dequeue window is documented.</li> <li><strong>Changed (action needed): <code>inbox read-primary</code> no longer acks.</strong> After handling the mail, run the <code>ackCommand</code> it returns (<code>inbox ack-primary <id> --receipt <rid></code>). <code>inbox drain-primary-legacy</code> keeps the old read-and-ack behavior for one release.</li> <li><strong>Changed: one shared Stop policy</strong> for the gates, honoring <code>stop_hook_active</code>, with per-kind caps.</li> <li><strong>Changed: a send to a stale twin is rerouted only on proof</strong> that the twin is stale.</li> <li><strong>Fixed: a torn journal row no longer corrupts the next append.</strong></li> <li><strong>Fixed: heal-orphan-partitions no longer stamps a pass done</strong> when its deadline deferred orphans; the deferred ids count as pending and the next run finishes them.</li> <li><strong>Fixed: the ingest daemon writes through the locked partition door</strong> (the per-id lock plus a registration recheck). A busy or unregistered partition leaves the batch pending in the WAL for replay instead of writing unlocked.</li> </ul> <h2 id="01053-2026-09-24">0.105.3 (2026-09-24)<a class="headerlink" href="#01053-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: the optional Jev check now backs speculation-guard</strong> (the always-on hedge check) instead of the opt-in LLM judge, and is trusted only to block. Its question was worded so a hedged guess could never count as speculative, so a clearly speculative reply got a confident allow; with the corrected question it scored 10/10 on labelled replies (8/10 confidently). Any other Jev outcome falls through to the regex check; with Jev off, the hedge check sees exactly the same text as before (the de-duplicated text extraction is used only as Jev's input, not the regex path). The LLM judge (<code>ANTIHALL_SEMANTIC_JUDGE</code>) is back to its pre-Jev form.</li> <li><strong>Fixed: Jev's time limit (default 1.5s, clamped to at most 3s) now covers the whole response, including the body</strong> — previously the deadline only bounded the initial request/headers, so a server that stalled after sending headers could hang past the configured timeout.</li> </ul> <h2 id="01052-2026-09-24">0.105.2 (2026-09-24)<a class="headerlink" href="#01052-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li><strong>Changed: command-guard and git-guard share one shell-scanning library</strong> (<code>hooks/lib/shell-scan.js</code>) for heredoc and substitution parsing. Duplicated parsing had produced repeated guard defects; behaviour is unchanged, proven by a 138-command differential against the previous release plus an independent 65-command adversarial check (0 differences).</li> <li><strong>Fixed: the Codex port was missing the progress-prune SessionStart hook.</strong> It is now registered.</li> <li><strong>Added: a hygiene test</strong> that fails when the Claude and Codex manifest versions differ, or a platform-neutral hook is missing from the Codex hook list.</li> </ul> <h2 id="01051-2026-09-24">0.105.1 (2026-09-24)<a class="headerlink" href="#01051-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li> <p><strong>Fixed: the orphaned-mesh warning no longer lists archived, closed workspaces with no live identity family.</strong> <code>devswarm.js</code> assigned <code>module.exports</code> after <code>main()</code> had already run and exited, so the orphan policy's lazy self-require got an empty module when devswarm.js was run as the CLI entrypoint; the archived-workspace check failed open and every such partition stayed in the orphan list. Exports are now assigned before <code>main()</code> runs, and the policy retries a load that came back unusable instead of caching it. A regression test drives the real CLI entrypoint (the prior equivalence test only exercised <code>require()</code>).</p> </li> <li> <p><strong>Fixed: statusline renders its own renderers in-process instead of forking a shell plus two node processes, each of which spawned git.</strong> The 3s outer watchdog was shorter than the 10s inner spawn timeouts it wrapped, so under load the output silently truncated before those inner budgets were reachable. Inner work is now bounded at 2.5s (git calls at 1.5s) under the unchanged 3s watchdog; a foreign base command still spawns. Measured: p50 153ms to 104ms at rest, 245ms to 163ms under load; output is byte-identical.</p> </li> </ul> <h2 id="01050-2026-09-24">0.105.0 (2026-09-24)<a class="headerlink" href="#01050-2026-09-24" title="Permanent link">¶</a></h2> <ul> <li> <p><strong>Added: optional Jev backend for <code>speculation-judge</code>.</strong> An opt-in Jev (TypeSafe System One, via the Vercel AI Gateway) judge can now run before the existing Haiku judge — <strong>off by default</strong>, enabled via <code>~/.anti-hall/jev.json</code> (<code>{"enabled": true}</code>) or <code>ANTIHALL_JEV=1</code> (<code>ANTIHALL_JEV=0</code> always wins). A verdict at or above the confidence threshold (0.85) is used directly in either direction; a low-confidence verdict or any failure (no key, timeout, HTTP error, bad response) falls back to the unchanged Haiku judge. Disabled, behavior is byte-identical to before Jev existed and writes no log; the key is never logged. Documented in <code>docs/KB-jev-classifier.md</code> with the benchmark behind the defaults.</p> </li> <li> <p><strong>Fixed: <code>git-guard</code> now scans commit messages passed via <code>-F -</code> / a heredoc, or <code>-F <file></code>, for AI self-credit trailers.</strong> Previously only the inline <code>-m</code> / <code>--message</code> form was inspected, so the same trailer could land via <code>-F</code> unblocked. The guard now extracts heredoc bodies from the raw command and reads <code>-F</code> files (relative to a preceding <code>cd</code> in the same command), running the existing trailer check over them. Purely additive — command segmentation and every existing block behave exactly as before; an unreadable file still fails open.</p> </li> </ul> <h2 id="01040-2026-09-24">0.104.0 (2026-09-24)<a class="headerlink" href="#01040-2026-09-24" title="Permanent link">¶</a></h2> <p><strong>First structural step of a mesh redesign.</strong> Recurring project-context defects traced back to ~16 separate "which project am I in" resolvers scattered across hooks and the CLI, each answering the worktree/repoKey/meshId question slightly differently. This release collapses the resolution path into one module (<code>companion/lib/identity.js</code>) and moves the first batch of callers onto it.</p> <ul> <li> <p><strong>Fixed: a lane in a submodule inside a linked worktree minted a phantom <code>modules-<hash></code> key and its inbox refused with <code>project-context-mismatch</code>.</strong> <code>devswarm-repokey</code>'s resolvers now shim over <code>identity.resolveContext</code>, which keys every submodule kind to the outermost superproject instead of stopping at the submodule's <code>.git</code> file. Nested-repo git answers are cached per path (5-minute TTL); the registry path guard compares realpaths.</p> </li> <li> <p><strong>Fixed: inbox reads refused when git timed out under load while the caller's cwd was inside the registered worktree.</strong> The read guard no longer refuses when the caller's project key can't be resolved (e.g. a git timeout) but the caller's cwd is inside the workspace's registered worktree; a cwd that resolves to a genuinely different project is still refused. <code>resolveContext</code> gains <code>missingPath:'ancestor'</code> for caller cwds (row paths keep <code>'null'</code>).</p> </li> <li> <p><strong>Fixed: <code>doctor --quiet</code> took 745s and 41,587 git spawns on a 385-store home.</strong> <code>resolveCallerWorktree</code> spawned git twice per registry row with no memo, and doctor's three all-store sweeps called it for every row of every store, including rows whose worktrees were long deleted. It now takes <strong>27s and 514 spawns</strong>, with identical resulting registry/cursor/message state. Deleted worktree paths resolve to null and are never folded onto an enclosing repo.</p> </li> <li> <p><strong>Fixed: reading a store created an empty store directory.</strong> <code>openStore({readOnly:true})</code> never creates a store: a missing store returns null, and an existing sqlite store opens read-only. Used by the unread read path and doctor's enumeration sweeps — 213 empty store dirs had piled up before this.</p> </li> <li> <p><strong>Fixed: the installer would install a launchd unit for a temp working directory.</strong> An e2e run installed a real ingest unit whose <code>WorkingDirectory</code> was a scratchpad fixture repo under <code>/private/tmp</code>; once the fixture was cleaned up, launchd failed to spawn it (<code>EX_CONFIG</code>) and restarted it <strong>34,584 times</strong>. The installer already refused a temp <code>HOME</code> but not a temp worktree. It now forces a dry run for a <code>WorkingDirectory</code>/main worktree under a temp root unless <code>ANTIHALL_INGEST_ALLOW_TMP_HOME=1</code>, and <code>doctor</code> reports (never removes) any other worktree's installed unit that points under a temp root, with the exact bootout + quarantine command.</p> </li> <li> <p><strong>Changed: statusline and wake-watch identity resolution.</strong> The statusline resolves its toplevel through <code>identity</code> (no git spawn). <code>wake-watch</code> reads <code>CLAUDE_CODE_SESSION_ID</code> (the variable Claude Code actually sets) before falling back to the legacy <code>CLAUDE_SESSION_ID</code>.</p> </li> <li> <p><strong>Added: <code>task-lifecycle-log</code> hook (Claude-only).</strong> Appends <code>TaskCreated</code>/<code>TaskCompleted</code> events to the session history ledger. Log-only: never blocks, never injects context. Codex has no task lifecycle hook events, so this is not mirrored to the Codex port.</p> </li> <li> <p><strong>Added: <code>doctor</code> read-only <code>identity-rekey-candidates</code> report.</strong> Lists stores written under the old wrong project keys, with message counts. Nothing is moved automatically. On the author's machine, only 2 real messages were affected.</p> </li> <li> <p><strong>Added: <code>tests/mesh-invariants.harness.test.js</code>, a model-based test harness for mesh invariants.</strong> Known defects are <code>todo</code> in normal runs and fail under <code>ANTIHALL_HARNESS_STRICT=1</code>, so they stay visible without breaking CI until fixed.</p> </li> <li> <p><strong>Pending (landing next, B3):</strong> the six hook-local <code>findGitToplevel</code> copies (<code>command-guard</code>, <code>parent-gate</code>, <code>child-gate</code>, <code>child-turn</code>, <code>parent-reply-tracker</code>, <code>parent-inbox</code>) will resolve through <code>identity</code> too. A hook whose cwd was deleted will resolve from its nearest existing ancestor; nested submodules will no longer get a phantom <code>meshId</code>; child descriptors will keep the literal git toplevel as <code>worktreePath</code>.</p> </li> </ul> <p><strong>Known, not yet fixed.</strong> The MIN-floor phantom unread count and the pull-crash message-loss defect are now reproduced by the harness above (<code>ANTIHALL_HARNESS_STRICT=1</code>); both will be addressed in the next phases of the mesh redesign.</p> <ul> <li><strong>Fixed: a message written through a quoted heredoc was blocked as a heavy command.</strong> The command guard extracted backtick and <code>$(...)</code> substitutions from the whole raw command without recognising heredocs, so prose such as a backticked <code>pytest tests -k x</code> inside a <code><<'EOF'</code> message body was treated as an executed command (field report: blocked as "verb: pytest"). Bodies of quoted-delimiter heredocs are now skipped, as bash does; an unquoted <code><<EOF</code> body still expands, so substitutions there are still checked. The guard also now sees heavy commands behind the <code>taskpolicy</code> and <code>xargs</code> wrappers.</li> </ul> <h2 id="01030-2026-09-23">0.103.0 (2026-09-23)<a class="headerlink" href="#01030-2026-09-23" title="Permanent link">¶</a></h2> <ul> <li> <p><strong>Fixed: UserPromptSubmit context repeated up to 27× in one delivered turn.</strong> Claude Code runs the UserPromptSubmit hook once per QUEUED prompt — hooks fire at queue time, not at delivery — and then delivers the queued prompts together, so every hook's identical block repeated N times in one turn. Measured on a field transcript: bursts span a median 26s / max 137s (too wide for a fixed time window to collapse), 12.8M chars injected total, and the worst single turn carried 93,220 chars with the DevSwarm workspaces table repeated 11 times. Fix (<code>hooks/lib/emit-dedupe.js</code>): a block is suppressed while its previous copy is still undelivered, checked against the transcript itself (a UPS <code>hook_additional_context</code> attachment matching the emitted content by sha1, never by prefix) — not against a timer. State resets on every SessionStart source (startup/resume/<code>/clear</code>/compaction, via the new <code>emit-dedupe-reset.js</code> hook) so a genuine context loss re-sends everything the fresh context no longer holds. The unchanging parts (the DevSwarm workspaces table, the orphaned-mesh banner) are additionally re-sent only on change (volatile ages like "3m ago" are ignored) or every 10 delivered turns, whichever comes first — heartbeat staleness alone no longer forces a re-send. Fail-open: any error in the dedupe path emits, same as before. Kill switch: <code>ANTIHALL_EMIT_DEDUPE=0</code>. As a side effect of instrumenting this, segment-read errors in <code>devswarm-parent-inbox.js</code> are now logged to <code>~/.anti-hall/logs/parent-inbox-segment-errors.ndjson</code> — the orphan banner was observed to vanish for one turn during the investigation and the cause is not yet known; this gives it a place to leave evidence next time.</p> </li> <li> <p><strong>Changed: Codex models are no longer pinned.</strong> <code>gpt-5.4</code>, <code>gpt-5.4-mini</code>, and <code>gpt-5.3-codex(-spark)</code> were silently removed from Codex's own model catalog and now return 400 on use — every anti-hall Codex routing reference to those slugs went dead at once, including the default "cheap seat" named across 23 shipped and doc files. Routing is now by category — <strong>frontier</strong> / <strong>workhorse</strong> / <strong>fast</strong> — resolved at call time from Codex's own live catalog (<code>~/.codex/models_cache.json</code>, matched on description text and <code>priority</code>, never a hardcoded slug map: <code>plugins/anti-hall/companion/lib/codex-models.js</code>). When a category cannot be resolved (missing cache, no match), the caller omits <code>-m</code> entirely and Codex falls back to its own configured default rather than failing. <code>deadly-loop</code>/<code>ship-it</code>'s Critic seat also now accepts <code>args.codexCriticModel</code> to override the resolved model per call.</p> </li> <li> <p><strong>Changed: the DevSwarm roster's <code>finish</code> column showed <code>0/3</code> for every workspace that had never had a gate set, not just the ones genuinely at zero.</strong> Only the manual <code>devswarm.js gate</code> verb ever writes a gate row; an untouched workspace has an empty gate map, which is a different state from "every required gate checked and failing." <code>finishingRate()</code> (<code>hooks/devswarm-parent-inbox.js</code>) now renders <code>—</code> when the gate map is empty/absent for that workspace, and the real ratio once any gate has ever been set — unchanged from before in that case.</p> </li> <li> <p><strong>Process:</strong> <code>RELEASING.md</code> — a release commit must now pass CI on an <code>rc-v<version></code> tag before it moves onto <code>main</code>. The marketplace fast-forwards <code>main</code> on every push, so a red main commit is installed immediately; 0.102.2 shipped exactly that way (see the 0.102.3 entry below).</p> </li> </ul> <h2 id="01023-2026-09-23">0.102.3 (2026-09-23)<a class="headerlink" href="#01023-2026-09-23" title="Permanent link">¶</a></h2> <p><strong>0.102.2 was pushed to <code>main</code> but failed CI and was never tagged.</strong> Because the marketplace clone fast-forwards <code>main</code>, it was installed anyway on machines that ran an update in the window. This release is the first green one carrying the 0.102.2 changes, plus two fixes to them. Installs that picked up 0.102.2 must move to a new version number to receive these — the updater only copies a version into the plugin cache when that version's directory is absent.</p> <ul> <li> <p><strong>Fixed: the 0.102.2 cursor-parity change was inert for every child descriptor.</strong> It derived the instance nonce from each descriptor's own worktree (<code>cwd: d.worktreePath</code>), but <code>deriveInstanceNonce</code> matches the CALLING process's ancestry against the cwd it is given. For any descriptor whose worktree is not the gate's own cwd, that match fails and it falls back to <code>self:<parentPid>:<startMs></code> — a namespace no reader declares — so the count dropped straight back to the cross-instance MIN floor the change was meant to escape. Measured: one process, one reader, two cwds, two different nonces. The gate now derives the nonce from its own cwd: the nonce identifies the READER, not the partition. It had only ever worked for the gate's own row.</p> </li> <li> <p><strong>Fixed: the 0.102.2 vacuity test could not run in CI.</strong> It rebuilt the pre-fix hook with <code>git show <sha>:<path></code>, which fails on the depth-1 clone CI uses, so all four matrix jobs failed. The test now reconstructs the old naive resolver inline and asserts that it and the shipped <code>resolveWorktreeNoSpawn</code> DISAGREE on the same real submodule fixture — which is the defect itself, proven with no dependency on git history. Verified on a simulated shallow clone before release.</p> </li> </ul> <h2 id="01022-2026-09-21">0.102.2 (2026-09-21)<a class="headerlink" href="#01022-2026-09-21" title="Permanent link">¶</a></h2> <p>Two fixes from a defect filed by a peer session, plus the diagnostic that would have let that session self-diagnose one of them.</p> <ul> <li> <p><strong>Fixed: the parent gate reported unread counts that contradicted the CLI, and escalated on them.</strong> Root cause was one divergence, not several: the gate derived the Primary's own id from a pure-fs walk (<code>findGitToplevel</code>) while the family grouping key used <code>canonicalMeshId</code> -> <code>resolveCallerWorktree</code> (git rev-parse). <strong>Inside a git submodule those disagree</strong>, because a submodule's <code>.git</code> is a FILE — <code>statSync</code> succeeds and the fs walk stops at the submodule instead of reaching the superproject. From a cwd inside a submodule the gate minted an own id that was a bare path hash of that directory: no descriptor, no registry row, <code>known:false</code> to the CLI. It then synthesised a self-row for that id with no store-existence check, and an unconditional survivor-force transplanted the real Primary's unread count onto it — printing a <code>read-primary <phantom-id></code> command that could not run. The same wrong id defeated the self-sent filter (<code>row.sender === own.id</code>), so a child's count swung between its filtered and unfiltered value depending on which cwd the gate fired from. Reported as 325 unread on a nonexistent id while three independent CLI reads returned 0, and a child shown as 1 that a direct peek put at 82. Three changes: the gate now resolves its own id through the same canonical derivation the family key uses; the survivor is forced to the self-row only when that id actually appears in <code>registryRows</code>; and the gate's <code>unionUnread</code> call — the only one in the repo passing neither <code>storeBaseCursor</code> nor the per-instance <code>#nd-</code> cursor — was migrated onto the v0.102.0 cursor namespace so gate and CLI agree by construction. NOT a v0.102.0 regression: the resolver line dates to v0.73.0 and the survivor-force to v0.100.0. v0.102.0 only added the <code>cached</code>/<code>live-resolved</code> label (which mislabels provenance on a number that did not come from the cache) and made the child count cwd-sensitive.</p> </li> <li> <p><strong>Added: <code>resolveWorktreeNoSpawn</code>, a zero-spawn submodule-aware toplevel resolver</strong> (<code>companion/lib/devswarm-repokey.js</code>). It tells the two <code>.git</code>-file shapes apart without invoking git — a submodule's gitdir carries a <code>.git/modules/<name></code> segment, a linked worktree's carries <code>.git/worktrees/<name></code> — and does not stop walking at a submodule boundary. Measured cold in a fresh process, which is how a hook always runs: <strong>3.0 ms</strong> against <strong>380 ms</strong> for the spawning path. So the fix removes a ~127x cost rather than adding one, and keeps the 15,700-line <code>scripts/devswarm.js</code> require off the common hot path. The spawning resolver remains as the fallback.</p> </li> <li> <p><strong>Fixed: <code>REFUSED TO ARM: lock-held</code> gave no way to tell a live holder from a stale lock.</strong> The refusal now names the holder's <code>sessionId</code> and absolute acquire time alongside the existing <code>pid</code>, <code>age</code> and <code>version</code>. A peer session read a bare <code>lock-held</code> as a stale lock and filed it as a defect before checking <code>ps</code>; the lock was live and correctly held, and the single-consumer guarantee had worked exactly as designed. The new <code>sessionId</code>/<code>ts</code> fields are ADDITIVE to the lock JSON — a lock written by a pre-0.102.2 watcher parses with them absent and degrades to 'unavailable' rather than crashing. That backward compatibility is load-bearing: long-lived watchers pin the plugin version that armed them, and five 0.101.1 watchers were resident on the author's machine at release time.</p> </li> </ul> <p><strong>Note on deployment.</strong> Long-lived companions (wake-watch, supervisor, ingest) run from a version-pinned cache path and keep running the version that armed them until their session restarts. Hooks re-execute from disk per event, but the entry point resolves through a cache dir bound at session start. So this release takes effect per session on <code>/reload-plugins</code> or a restart — not the moment it is tagged.</p> <h2 id="01021-2026-09-20">0.102.1 (2026-09-20)<a class="headerlink" href="#01021-2026-09-20" title="Permanent link">¶</a></h2> <p>Three fixes, each with a proven root cause. All were found while auditing a live install, not from a report.</p> <ul> <li> <p><strong>Fixed: the periodic supervisor never drained <code>heal-registry-rows</code>, so that stage only advanced when a human happened to run <code>update</code> or <code>doctor</code>.</strong> <code>DEFERRED_SWEEP_STAGES</code> listed only <code>fold-all-stores</code>, <code>heal-orphan-partitions</code> and <code>fold-archived-rows</code>, and the supervisor runs exactly one of those per tick. <code>heal-registry-rows</code> was wired into <code>update.js</code>'s <code>healRegistryPostUpdate</code> but was absent from the rotation, so nothing ever selected it. Observed on a live machine: 204 stores pending, with each explicit <code>update</code> run rehoming ~6 before hitting its 20s budget. A deferred stage whose convergence depends on someone remembering to run a command does not converge. The stage is now in the rotation and is selected like the other three.</p> </li> <li> <p><strong>Fixed: the central logger ignored <code>ctx.home</code>, so tests wrote into the real <code>~/.anti-hall/logs/devswarm.jsonl</code>.</strong> <code>anti-hall-log.js</code>'s <code>logDir()</code> resolved from <code>ANTI_HALL_LOG_DIR</code> or fell straight back to <code>os.homedir()</code>, never consulting the home passed to a verb. Any test exercising a failing CLI verb reached <code>logVerbOutcome</code> -> <code>alog.logError</code> and appended a real line to the user's own log. This is the fourth instance of the same class (see the HOME-isolation fixes in 6687c11). Two test files proven to hit the path now set an isolated <code>ANTI_HALL_LOG_DIR</code>, and <code>logDir()</code> throws a tagged error when <code>NODE_TEST_CONTEXT</code> is set without it — mirroring the existing <code>resolveHomeGuarded</code> pattern in <code>devswarm-store.js</code>. Production logging stays fail-open: the throw is swallowed by <code>writeEntry</code>'s own catch and surfaced on stderr only.</p> </li> <li> <p><strong>Fixed: <code>doctor</code> reported a false FAIL under CPU load, and made the test suite flaky.</strong> The unconditional Statusline section spawned <code>statusline.js</code> with <code>timeout: 5000</code>. Under load that spawn can lose the contention race with no bug and no hang; <code>spawnSync</code> then SIGTERMs the child and returns <code>status: null</code>, which fell through to the "produced no output" branch, incremented <code>fail</code>, and flipped the exit code to 1. Every <code>runDoctor()</code> call reaches that section regardless of flags, which is why the failing test file moved between runs. Commits fe0d901 and b99eafb fixed this same class twice by raising the OUTER test-harness timeout to 60000ms; neither touched this INNER product-code spawn. Both it and the shared <code>runHook()</code> helper (backing ~15 other self-test sections) now use 30000ms, and a SIGTERM'd spawn is reported as a load-related warning naming the real cause rather than counted as a failure. A genuine "ran but produced no output" result is still a failure.</p> </li> </ul> <h2 id="01020-2026-09-20">0.102.0 (2026-09-20)<a class="headerlink" href="#01020-2026-09-20" title="Permanent link">¶</a></h2> <p>Six DevSwarm mesh defects were reported by three independent sessions. Two of the six did not survive verification and are recorded here so they are not re-chased.</p> <ul> <li><strong>Fixed: <code>inbox read-primary</code> and <code>inbox tick</code> used different NDJSON cursors, so a child could be woken forever by unread mail that <code>read-primary</code> could never clear.</strong> <code>read-primary</code> read and acked the raw descriptor cursor, while <code>count</code>/<code>read</code>/<code>ack</code>/ <code>tick</code> used the per-instance cursor from <code>resolveNdCursorPath</code> (<code>cursors/<id>#nd-<short>.json</code>). Nothing reconciled the two: the per-instance file is seeded from the descriptor once at creation, <code>projectNdDescriptorCursor</code> only pushes instance -> descriptor, and <code>reconcileOrphanCursor</code> explicitly excludes the NDJSON cursor. <code>read-primary</code> now uses the same resolver as every other verb and projects the MIN back to the descriptor afterwards, preserving the sibling invariant that stops one reader advancing past another reader's unread mail. A doctor repair step (<code>reconcileStuckNdCursors</code>) raises already-stuck per-instance cursors toward <code>min(descriptor, real inbox line count)</code> via <code>ackTo</code>, which re-clamps to a freshly read line count and refuses to lower — it can never skip mail.</li> <li><strong>Fixed: a Primary's own outbound messages counted as the Primary's own neglect.</strong> The store path already filtered rows whose sender is the Primary itself; the NDJSON path could not, because the NDJSON wire carried no sender field at all. Rows drained by <code>devswarm-pull.js</code> now carry <code>sender</code>, and the gate applies the same filter. The field is appended after <code>_h</code> is computed, so dedupe hashes and line positions are unchanged, and rows written before this release (which have no <code>sender</code>) parse and count exactly as before.</li> <li><strong>Fixed: a blocked Primary could not tell whose mail it was blocked on.</strong> The Primary's own row is synthetic and keyed on its worktree; identity-family collapse then groups every descriptor sharing that worktree and SUMS them under one id. So <code>primary-<id> — N unread</code> could be a sum over sibling descriptors while the Primary's own contribution was zero — which is exactly what two reporters hit. The gate now names each contributor, its count, whether it is archived or its worktree is gone, and the exact command that clears it (<code>inbox ack <id> --ack-as-owner</code>). Counting and every blocking decision are unchanged.</li> <li><strong>Fixed: <code>send</code> failed with <code>lockBusy</code> under ordinary concurrency.</strong> A bounded retry (3 attempts, jittered backoff) now wraps the per-workspace lock. The retry is safe because the timestamp and message hash are computed once before the loop and <code>withIdLock</code> never runs the append when the lock is busy, so the append happens at most once. <code>send</code> already exited non-zero on <code>lockBusy</code>; that was verified, not assumed.</li> <li><strong>Added: byte-length fields on <code>inbox read-primary</code>.</strong> Nothing in this plugin truncates message bodies, but the ack is content-blind: <code>read-primary</code> emits and acks in one call, so if anything downstream of the CLI's stdout clips bytes, the cursor has already moved. Each row now carries <code>bodyLength</code> (UTF-8 bytes) and the payload carries <code>totalBodyBytes</code>, so a clipped consumer can detect a short read, plus a hint naming the recovery path. The ack was deliberately NOT made conditional on a delivery receipt: that would break the positional invariant the partition and sibling cursor arithmetic depend on.</li> <li><strong>Added: age-based GC for stale summary projections.</strong> <code>summaries/*.json</code> accumulated one file per (repo x project) ever derived with nothing pruning them. GC is keyed on <code>generatedAt</code> (default 30 days, <code>ANTIHALL_DEVSWARM_SUMMARY_RETENTION_DAYS</code>) and never removes a summary whose store directory still exists.</li> </ul> <p><strong>Reported but NOT defects</strong> — verified against the code and recorded so they are not re-investigated: - "The gate double-counts unread across ~150 duplicate summary shards." It does not. The gate opens exactly one summary file per repoKey and enumerates children from descriptors, which carry no duplicate ids. The "duplicate shards" are per-repo partitions of one multi-repo workspace: the same workspace id legitimately appears under each repo it spans, derived at different times. - "Mesh messages truncate silently and a truncated message is unrecoverable once acked." Nothing truncates bodies anywhere in send, store or read — every cap found is a row count. Acked mail is re-servable: <code>inbox messages <id></code> is read-only and not unread-scoped, rows persist in both the NDJSON inbox and the store, and an ack is a watermark, never a delete. The real gap was that the receiving agent had no way to know that, which the byte-length fields and the recovery hint above address.</p> <p><strong>Deliberately not done:</strong> - Excluding archived or gone-worktree descriptors from the gate's count. This was attempted and reverted: 45 existing tests encode the opposite contract on purpose — archiving suppresses the liveness axis only, never the unread axis, and a gone worktree's mail is real and drainable. Excluding it would have hidden real mail. - Letting <code>reap-orphans --apply</code> run non-interactively. It deletes mesh partitions and its human gate is deliberate. The attribution fix above removes the need to delete anything to clear a false block.</p> <p><strong>Known, not fixed in this release:</strong> - <code>resolveSelfId</code> projects through the main worktree while the gate's own id does not, so if a Primary ever runs from a linked worktree the two ids diverge and the self-sent filter no-ops. It fails toward counting, never toward hiding. - Rows written before this release carry no <code>sender</code>, so a pre-existing backlog of a Primary's own outbound still counts until it drains. - <code>gcStaleSummaries</code> has a narrow window where deleting a summary as a store directory appears could read as "no unread" for one cycle.</p> <h2 id="01011-2026-09-19">0.101.1 (2026-09-19)<a class="headerlink" href="#01011-2026-09-19" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: DevSwarm ingest daemons ignored SIGTERM, so duplicates piled up.</strong> <code>devswarm-ingest.js</code> registered SIGTERM/SIGINT listeners, which disables Node's default terminate-on-signal, but its loop is fully synchronous (<code>spawnSync</code> and an <code>Atomics.wait</code> sleep), so the event loop never ran the handler: SIGTERM was ignored forever. <code>launchctl unload</code> therefore never stopped the old daemon, and each reload started another beside it. Observed live: 12 duplicate daemons for 5 units, 7-12 days old, stoppable only by SIGKILL. The listeners are removed so the default disposition applies. This is safe because the lock is already reclaimed immediately from a holder whose PID is confirmed dead. A regression test sends a real SIGTERM to a real child process and requires it to exit; it fails against the old handler. Three older tests that "proved" the handler worked by firing a simulated signal on a fake process object were removed: they passed while a real signal could never reach that code.</li> <li><strong>Fixed: a daemon that lost its lock kept running as a second consumer.</strong> The loop discarded <code>release.heartbeat()</code>'s result. <code>heartbeat()</code> now returns <code>true</code> (refreshed), <code>false</code> (definitive loss: the lock file is gone, or it parsed cleanly with another owner's token) or <code>'error'</code> (a transient read, parse, write or rename failure). The loop exits only on <code>false</code>; a transient error is logged once and the daemon keeps running, so a filesystem blip cannot take down a healthy daemon.</li> <li><strong>Fixed: the installer reloaded a unit without waiting for the old daemon.</strong> <code>macInstallProject</code> now reads the old daemon's PID, unloads, and waits (bounded, 10s) for it to exit before loading the new plist. If it outlives the deadline it is SIGKILLed, but only after re-reading that PID's command line and confirming its shape: the executable is <code>node</code> and an argument ends in <code>/companion/devswarm-ingest.js</code>. A PID reused by any other process, including <code>tail -f</code> on <code>devswarm-ingest.log</code> or an editor holding the script open, is never killed. <code>stopLegacyUnitEntry</code> now deletes a legacy lock file only after confirming its holder is dead. Every step fails open and none can block indefinitely.</li> </ul> <h2 id="01010-2026-09-19">0.101.0 (2026-09-19)<a class="headerlink" href="#01010-2026-09-19" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed (P0): <code>UserPromptSubmit hook timed out after 10s</code> on every prompt in repos with many DevSwarm workspace rows.</strong> <code>devswarm-parent-inbox.js</code> spawned a <code>git</code> process PER workspace row, per prompt — in two separate loops. The #36 repo-scope filter (<code>repoKeyOfWorktree</code>) ran <code>git rev-parse --git-common-dir</code> for each row, and the unanswered-question family collapse (<code>resolveMeshId</code> -> <code>canonicalMeshId</code> -> <code>resolveCallerWorktree</code>) spawned <code>git</code> twice per distinct worktree. The in-process <code>repoKeyCache</code> never hit, because it is keyed per worktree path and every prompt is a fresh process. CPU profile: <code>spawnSync</code> was 74% of the hook's time. Under machine load the hook crossed its 10s timeout and its injection was silently discarded. Both loops now derive repo identity from git's own on-disk worktree metadata with no spawn (<code>gitCommonDirNoSpawn</code> / <code>repoKeyForWorktreeFast</code> in <code>companion/lib/devswarm-repokey.js</code>): <code>.git</code> as a directory (main checkout) or a <code>gitdir:</code> file plus <code>commondir</code> (linked worktree). A worktree path that no longer exists spawns nothing. The spawn-based path remains only as a fallback for submodule and malformed shapes. repoKey/meshId output verified byte-identical to the git-derived values for main-checkout, linked-worktree, submodule and missing-path shapes, and the #36 filtering decision is unchanged for every row. Measured on the worst-affected repo: 8,935 ms before; 1.4-3.1 s after at machine load ~100. This hook's own header forbids spawning on the hot path; it now honours that.</li> <li><strong>Removed: graphify and Obsidian integration, entirely.</strong> Owner decision to retire both tools.</li> <li>Deleted hooks <code>graphify-session.js</code> (SessionStart), <code>graphify-guard.js</code> (PreToolUse on Bash/Grep/Glob) and <code>graphify-reminder.js</code> (Stop), plus their tests, and all their registrations in the Claude <code>hooks.json</code>, the Codex <code>hooks.json</code>, and <code>codex/install-codex.js</code>.</li> <li>Removed the injected <code>E2. GRAPHIFY-FIRST</code> orchestration rule (<code>verify-first-orch.js</code>), the graphify-first phrase in <code>verify-first-full.js</code>, doctor's Graphify section (later sections renumbered), graphify steps in the <code>orchestration</code>, <code>ship-it</code> and <code>deadly-loop</code> skills and both Codex skills, and <code>.graphifyignore</code>.</li> <li><code>scan-throttle.js</code> now ships with <strong>no built-in patterns</strong>. Its only built-in was graphify. It remains a generic throttle driven by <code>ANTI_HALL_THROTTLE_PATTERNS</code> and matches nothing unless that is set.</li> <li><strong>Migration for existing Codex installs.</strong> <code>install-codex.js</code> writes hook registrations into the user's own <code>~/.codex/hooks.json</code> (or project <code>.codex/</code>), so existing installs still registered the deleted hooks. <code>update.js</code> (<code>codexGraphifyHooksMigratePostUpdate</code>) and a new AUTO-SAFE doctor repair step (<code>codex-graphify-cleanup</code>) remove only those stale groups. Every other registration and top-level key is preserved; the file is backed up to <code>.bak-<timestamp></code> before any write; a file with nothing to migrate is left byte-identical; a malformed config is left alone; the file is never deleted. The doctor step exists because the existing codex-refresh check is per-EVENT and would report "already wired" while a stale graphify group remained.</li> <li>Claude side needs no migration: Claude Code loads anti-hall's hooks from the plugin's own shipped <code>hooks.json</code>, and nothing writes them into user settings.</li> <li>Historical CHANGELOG entries and dated plans/audits in <code>docs/</code> are left as they were written.</li> <li><strong>Known follow-up:</strong> <code>scripts/devswarm.js</code> <code>resolveCallerWorktree</code> still spawns <code>git</code> before its pure-fs fallback (backwards for a hot path) and always spawns a second <code>--show-superproject-working-tree</code> check. It is a fixed once-per-turn cost, not per-row, so it does not cause the timeout above, but it still costs the own-reader, child-turn and parent-gate hooks every turn. Fix: try the fs resolver first and skip the superproject spawn unless <code>.git</code> is <code>.git/modules/</code>-shaped.</li> </ul> <h2 id="01000-2026-09-11">0.100.0 (2026-09-11)<a class="headerlink" href="#01000-2026-09-11" title="Permanent link">¶</a></h2> <ul> <li><strong>New (LEDGER-ONLY): <code>claim-ledger.js</code>, a Stop hook that measures confident-but- unverified factual assertions the lexical <code>speculation-guard</code> cannot see</strong> — a claim with NO hedge word at all (e.g. "the spawn is still running", "you are on task 3 of your queue"). It builds a cumulative evidence string from everything the session actually observed (tool results, tool inputs, user prompts, hook attachments), extracts checkable tokens from the final assistant message (numbers with unit nouns, hex SHAs, <code>task N of</code>, <code>N days ago</code>, state words in a turn with zero tool calls), and records any token with no referent in that evidence to <code>~/.anti-hall/claim-ledger/<session>.jsonl</code> (class <code>hard</code>/<code>soft</code>). This release it NEVER blocks, never emits a decision, never prints, and always exits 0 — the point is to measure the real false-positive rate before anyone considers a blocking tier. Tuning carries over from a 5-transcript/1326-message prototype: per-turn evidence flagged 269 (20%, useless); cumulative evidence + a no-tool-call gate flagged 22. Reads only a 2 MB transcript tail at a file offset (0.17s against a 254 MB transcript). <code>speculation-judge.js</code> is untouched and remains opt-in and inert.</li> <li><strong>Fixed (P0, D4): <code>devswarm.js</code> had NO subcommand that recognized <code>--help</code>/<code>-h</code> — a help request fell straight through to real dispatch.</strong> <code>migrate -h</code> genuinely ran the migration; <code>merge --help</code> genuinely forwarded to hivecontrol AND sent a live, unconditional mesh broadcast to other people's sessions (surfaced when a read-only-fenced diagnostic agent triggered exactly this). <code>run()</code> now intercepts a help request — <code>--help</code>, <code>-h</code> anywhere among the positionals, or a bare <code>help [verb]</code> — BEFORE the switch statement, so it covers every verb including the two raw-argv-tail pass-throughs (<code>spawn</code>/<code>merge</code>), with zero store opens, zero filesystem writes, zero child processes. The verb list backing <code>help</code> output is derived from <code>run()</code>'s own switch statement rather than hand-typed, closing a pre-existing drift (<code>reconcile-registry</code>/<code>wake-directive</code> were real, dispatched verbs missing from the old hand-typed error-message list). Each verb's usage line names its concrete side effects when actually run, so a caller can tell before running one whether it's safe to explore.</li> <li><strong>Fixed (D1): archived DevSwarm workspaces were consuming roster table slots that live workspaces needed.</strong> An <code>archived</code>-label row competed for <code>MAX_TABLE_ROWS</code> slots on equal footing with every live row, so a project with several archived workspaces could push genuinely live ones into the <code>+N more</code> overflow line. An archived row is now dropped BEFORE sort/cap (never after), default ON via new env var <code>ANTIHALL_ROSTER_HIDE_ARCHIVED</code> (<code>0</code> restores the old behavior) — a row still genuinely coordinating (<code>not-draining</code>, rank 1.5) is a different label and is never caught by this filter, so nothing that still needs attention is ever hidden; the hidden count is always named via a <code>+N archived</code> note, never silently dropped. The table cap itself is now configurable via new env var <code>ANTIHALL_ROSTER_MAX_ROWS</code> (default 12, was hardcoded).</li> <li><strong>Fixed (D2): the per-turn broadcast feed was inflating the injection and repeating itself verbatim every turn.</strong> The advisory <code>recent[]</code> broadcast feed rendered <code>r.summary</code> — the FULL message body — verbatim on every turn with no memory of what had already been shown; a single sent broadcast could re-inject in full for the rest of the session (the real cause of a reported 10-12KB per-turn injection). Bodies are now capped to 200 chars with an ellipsis; a broadcast older than new env var <code>ANTIHALL_BROADCAST_MAX_AGE_MS</code> (default 24h) is dropped; and each row is deduped per session (by a stable <code>from</code>+<code>ts</code>+<code>summary</code> key) against a bounded, 200-key state file (<code>~/.anti-hall/devswarm/parent-inbox-broadcast-seen/<session>.json</code>) so a broadcast injects at most once per session. Any dedup-state read/write failure fails OPEN to the age-capped set — never a hard crash, never a silent full suppression.</li> </ul> <p><strong>KNOWN FOLLOW-UPS (not fixed in this release, flagged honestly):</strong> - <code>devswarm.js</code> still silently ignores unknown CLI flags rather than rejecting them. The design for a fix is settled — a mechanically-derived per-verb allowlist, warning-first (not a hard block), with <code>spawn</code>/<code>merge</code> exempt as documented pass-throughs — but deferred to a follow-up release so help support and flag-rejection don't land blind together in one release. - <code>inbox messages --tail N</code> under truncation returns <code>ok:false</code> with no <code>messages</code> key. That's unambiguous for a caller that checks <code>ok</code> first, but ambiguous for one that does <code>(r.messages || []).length</code> without checking <code>ok</code> — it would read a truncated response as an empty mailbox. No in-repo consumer does this today. - <code>ownReaderDelta</code> (<code>companion/lib/devswarm-own-reader.js</code>) has two paths that silently leave the raw (phantom) count uncorrected: a <code>stale:true</code> result returns <code>null</code> with no correction applied, and a missing instance file (the nonce derives from <code>{home, cwd}</code>, so a drain issued from a different cwd won't find it) returns <code>{delta: 0}</code> with no signal that the count is uncorrected. Not touched in 0.100.0. - Phantom unread in the per-turn DevSwarm injection is NOT fixed in this release. The parent-inbox projection can over-report unread relative to the authoritative <code>inbox count</code>/<code>inbox tick</code> — trust the CLI over the injected banner. A fix was written and REVERTED before release: it used the legacy shared-pair cursor as a live per-call floor in <code>computeSummary</code>, the exact pattern <code>scripts/devswarm.js:1668-1692</code> documents as tried, proven live to lose mail across a mixed-version fleet ("the old build received 3, and a DECLARED 0.99 instance then received 0"), and reverted. The defect over-reports (visible, harmless); the reverted fix risked under-reporting (silently hidden mail) — not a trade worth making. Investigated root cause, retained for a future safe fix: <code>commitInstanceAck</code> writes the legacy cursor file BEFORE the store cursor row, so the projection under-counts consumption in that window. The safe shape is to fix it at the WRITE ORIGIN in <code>commitInstanceAck</code> (keep file and store in agreement by construction where both were just written and provenance is known) rather than re-deriving a floor later on the read side from files whose writing build cannot be determined. - <code>computeSummary</code>'s main per-workspace loop has no exception handling around <code>messageCount()</code>/<code>cursorValue()</code>. A single workspace's cursor read failure crashes the whole call, taking down the live table, own-unread, and every other workspace's row for that turn, propagating uncaught through <code>deriveSummary</code>. - The journal backend's <code>cursorValue()</code> never throws on a genuine read error — <code>readAll()</code> swallows non-ENOENT errors and returns <code>[]</code>, so the cursor returns 0 silently, yielding <code>unread = total</code> (a false-maximum alarm). <code>getReadError()</code> exists to surface this; nothing consults it. An exception-based guard cannot catch this path.</p> <h2 id="0992-2026-09-11">0.99.2 (2026-09-11)<a class="headerlink" href="#0992-2026-09-11" title="Permanent link">¶</a></h2> <ul> <li><strong>Known issues, carried (pre-existing, shared with every other caller of the same primitives — NOT fixed this release, flagged by the Critic during the f061789267c1 own-reader review):</strong></li> <li><code>readInstanceBaseline</code> can WRITE <code>cursors/<id>#base.json</code> on what is conceptually a READ path (the ambiguous-branch live resolution in <code>companion/lib/devswarm-own-reader.js</code> calls <code>siblingBaseCursor</code>, which calls <code>readInstanceBaseline</code>, which seeds the baseline file on first touch). Every other caller of <code>siblingBaseCursor</code> (<code>inbox count</code>/<code>read</code>/ <code>ack</code>, <code>devswarm-child-drain.js</code>, <code>devswarm-child-gate.js</code>) already has this same side effect; it is not new here.</li> <li><code>openStoreForUnread</code> (called from the same ambiguous branch) passes no <code>backend</code>, so it inherits whatever the environment resolves to rather than an explicit choice — again shared with every existing caller of <code>openStoreForUnread</code>, not introduced by this fix.</li> <li><strong>Resolved the ambiguous stale-cache case with one live read</strong> (Critic GO-with-fix, on the P1 fix below): a bare UNKNOWN for <code>ownCursor === entry.total</code> left the reporter's exact configuration blocked forever on a healthy summary. That one branch now opens the store (never the common path) and computes the real number via the live message count minus <code>siblingBaseCursor</code> — a genuinely-drained reader now opens the gate, a reader behind real new mail still blocks, and a live read that cannot run still fails to UNKNOWN. Also fixed: <code>buildReason</code> no longer sends an operator into daemon healthcheck/logs triage for this cause — it now names the cache-vs-live mismatch and prescribes <code>inbox count</code>/<code>read-primary</code>.</li> <li><strong>Fixed (P1): the f061789267c1 own-reader fix below could HIDE real unread mail on a stale cache.</strong> <code>companion/lib/devswarm-own-reader.js</code> compared a LIVE instance cursor against a CACHED summary snapshot; once a reader's own live position caught up to or passed the snapshot's own <code>total</code>, the subtraction floored to 0 even though mail could have arrived after the snapshot that the reader had not actually seen — worse than the phantom the original fix replaced, since a Stop-gate hiding real mail fails in the dangerous direction. <code>ownReaderUnread</code>/<code>ownReaderDelta</code> now return <code>null</code>/ <code>{stale:true}</code> in that regime; every consumer (parent-gate, parent-inbox, child-turn) treats <code>null</code> as UNKNOWN — parent-gate routes it into its pre-existing <code>unknown:true</code> fail-safe (blocking) path, the two report-only nudge surfaces fall back to the raw pre-fix number.</li> <li><strong>Fixed (P2): <code>doctor --repair-ingest-orphans --repair-test-stores</code> (a combined invocation) silently ran only the first flag's section</strong> — the early-exit fix below called <code>emitVerdictAndExit()</code> unconditionally inside each flag's block, so the first one to run always exited before a later flag's block was reached. Each exit is now gated on no later repair flag also being set; a combined invocation runs every requested section, then exits once.</li> <li><strong>Fixed: <code>doctor</code>'s resurrected-registry-rows warning was ambiguous about the total row count.</strong> Two independent readers misread "N candidate(s), M needing manual review" as "M of N need review" (implying N was the total). <code>candidates</code> and <code>unhealable</code> are disjoint under the dry-run call this check always makes, so the message now states the sum up front: <code>"44 resurrected registry row(s): 24 repairable, 20 need manual review (no safe forward target)"</code>.</li> <li> <p><strong>Fixed: <code>doctor --repair-ingest-orphans</code> / <code>--repair-test-stores</code> / <code>--repair-resurrected</code> never exited after their own section</strong>, so their verdict line was buried behind every later report section and the full unconditional summary — field-observed at line 562 of 571 total output lines for <code>--repair-resurrected</code>. Each of the three now calls a shared <code>emitVerdictAndExit()</code> right after its own section instead of falling through.</p> </li> <li> <p><strong>Fixed (P0): the parent gate (and roster/reminder surfaces) could show a Primary or child a phantom unread backlog it had already drained</strong> (defect f061789267c1 / a77b85571dfa). 0.99.0's per-instance-cursor fix (defect 8b211241bbe9) deliberately made the SHARED cursor pair track the MIN across every live <code><id>#inst-<nonce></code> instance file, so no reader's mail is ever lost — correct for that pair's own cross-instance-safety contract. But <code>devswarm-store.js</code>'s <code>computeSummary()</code> sizes <code>workspaces[id].unread</code> from <code>total - <that shared min></code>, and every per-reader display (<code>devswarm-parent-gate.js</code>'s Stop-hook gate, <code>devswarm-parent-inbox.js</code>'s "Primary's OWN inbound unread" segment, <code>devswarm-child-turn.js</code>'s per-turn mesh-direct nudge, and the live-store reads in <code>devswarm-child-drain.js</code>/<code>devswarm-child-gate.js</code>) was reading that same min-floor number as if it were ITS OWN read position. A slower or stale sibling instance file is routinely still present (evicted only after the 7-day <code>gcInstanceCursors</code> window — well inside any ordinary multi-day gap), so a reader that had genuinely drained everything could still be blocked on mail it had already read. Live proof: a fresh reader's own instance file at 929, a 3-day-old sibling's at 859, total 930 — the shared floor gave <code>unread:71</code> though the fresh reader's true position left only 1 row unread. A cheaper GC cadence does NOT fix this: the stale file sits inside the SAME (unchanged) 7-day window no matter how often GC runs (see <code>tests/companion/devswarm-own-reader.test.js</code>'s GC-cadence test, a standing proof against re-proposing that shortcut). Fix: a new shared helper, <code>companion/lib/devswarm-own-reader.js</code>, computes each of the five per-reader surfaces' own number by subtracting this reader's own lead over the shared floor (<code>ownCursor - entry.cursor</code>, always</p> <blockquote> <p>= 0 by construction) from the already union-computed <code>unread</code>/<code>directUnread</code> — never re-deriving unread from scratch, never opening the store DB on the Stop-hook's cheap-read path, and failing open to the pre-fix number on any resolution failure (so a legacy/older summary shape, or a nonce-derivation failure, is byte-identical to before). <code>devswarm-child-drain.js</code> and <code>devswarm-child-gate.js</code> already had a live store handle open, so those two instead pass <code>scripts/devswarm.js</code>'s own <code>siblingBaseCursor</code> (the exact primitive <code>inbox count</code>/<code>read</code>/<code>ack</code> already use) as <code>storeBaseCursor</code> — the precise fix, not an approximation. No persisted-shape change: every input this fix reads (<code>entry.cursor</code>, <code>entry.unread</code>) was already part of the existing summary projection, so no forward-migration is needed in <code>update.js</code> or <code>doctor</code>. Monitoring/roster surfaces that show OTHER workspaces' unread (the child rows in <code>devswarm-parent-inbox.js</code>'s table, <code>roster</code>/<code>diagnose</code>'s CLI output, <code>doctor-runtime.js</code>'s stuck-ingest sweep, <code>liveness.js</code>'s drain-activity check) are deliberately UNCHANGED — those are cross-instance monitoring signals ("has ANY reader of this workspace drained it"), not a per-reader read position, and the min-floor is the CORRECT, conservative answer there.</p> </blockquote> </li> </ul> <h2 id="0991-2026-09-08">0.99.1 (2026-09-08)<a class="headerlink" href="#0991-2026-09-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed (P1): the store migration re-registered archived workspaces</strong> (defect df54edf54804). <code>migrateToStore</code>/<code>migrateOne</code> upserted the registry row for every id found in <code>workspaces/</code> unconditionally, never consulting <code>archived/<id>.json</code> or the registry tombstone <code>devswarm.js archive <id></code> had already appended — <code>removeRegistry</code> is not a permanent marker (the sqlite backend hard-deletes the row, the JSONL backend appends an unconditional <code>remove</code> op that a later upsert simply outraces), so migration reviving the row was a genuine resurrection, not a no-op. A field report (a downstream project) had 4 workspaces archived on 0.97.1 come back <code>archivedInApp: false</code> after the 0.99.0 update, re-entering the parent-inbox table and parent gate. Root cause: a still-running child terminal can recreate <code>workspaces/<id>.json</code> for an already-archived id via the explicit <code>register</code>/<code>register-primary</code> verb (the auto-<code>ensure</code> path's resurrection guard, field defect a48db2e0ea08, deliberately does not cover it), and migration then blindly trusted that recreated descriptor. The migration now defaults to NEVER resurrecting an archived id (<code>archivedSkipped</code> in its report), and migrates a recreated descriptor as live only with positive proof of a genuine later reuse: a differing sessionId from the archived marker, a demonstrably newer descriptor file (by mtime — <code>archived/<id>.json</code> is a hardlink of the id's OWN pre-archive descriptor, so its mtime is whatever that descriptor's last write was BEFORE it was archived, never later — archiving itself never rewrites it), and CURRENT liveness proof for the descriptor's own sessionId. That proof is the shared <code>isSiblingPartitionLive</code> predicate (<code>companion/lib/liveness.js</code>), not a bare fresh-heartbeat check — a bare heartbeat check reproduces the exact root cause that predicate's own header documents and replaced elsewhere: a Primary never writes a heartbeat at all, <code>register</code> writes none, and a child mid-long-turn's heartbeat goes stale well before the session does. Every other uncertainty (a corrupt/unreadable marker, an unresolvable mtime) fails toward skip, never toward resurrection, and the check is fully idempotent across repeated migration runs. Residual, by design (unchanged from 0.97.0/7e1ae67): a live child that re-registers an archived id after its Primary archived it is still treated as a legitimate re-registration by <code>register</code>/<code>register-primary</code> themselves — <code>devswarm.js archive</code> now warns loudly when the target still has a fresh heartbeat at archive time ("child session still live... it may re-register"), and the per-turn parent-inbox table labels a superseded-but-not-confirmed-live row <code>archived-superseded (live child)</code> instead of silently falling back into the ordinary dormant/escalated ladder as if it had never been archived — discriminated by worktree path (not a bare marker-file existsSync), so a genuinely new, unrelated workspace that merely reuses an old archived id at a DIFFERENT worktree is never mislabelled. <strong>Extended (worktree-group siblings):</strong> <code>cmdArchive</code>'s <code>retireArchivedWorktreeGroup</code> tombstones the STORE REGISTRY row of every sibling id sharing the archived id's physical worktree, but never writes those siblings their own <code>archived/<id>.json</code> marker or touches their descriptor file — so the direct per-id marker check above could not catch a sibling like this, and neither store backend leaves anything queryable to distinguish "tombstoned sibling" from "never registered" (sqlite hard-deletes the row with zero trace; the JSONL backend's <code>remove</code> op just makes the id absent from a <code>listRegistry()</code> read). The gate now also matches a sibling's worktree path against every OTHER id's archived marker (indexed once per run) and defaults to skipping it too, migrating it as live only on the SAME liveness proof — but never on mtime: a sibling's descriptor file is never touched by the archive of a different id, so its mtime has no relationship to that unrelated marker's and proves nothing either way. <strong>Decision logic extracted</strong> into a new shared module (<code>companion/lib/devswarm-archive-gate.js</code>) and reused by <code>scripts/devswarm.js</code>'s <code>healOrphanPartitions</code> doctor repair, which was the SAME class of bug (a bulk re-registration path consulting only a bare <code>hasArchivedCounterpart</code> marker check) and would otherwise silently re-adopt a group sibling the migration correctly refuses. <strong>Field aftermath, forward hygiene:</strong> an install that already ran the pre-fix migration once is left holding the resurrected rows regardless (one a downstream project install measured ~43 legacy-slug rows across a whole retired worktree-group family). <code>scripts/devswarm.js</code>'s new <code>reRetireResurrectedRows</code>/<code>reRetireResurrectedRowsAllStores</code> forward any unread mail into a same-worktree archived id, then remove ONLY the resurrected registry row (never a file), for a row that has a PROVEN archive link (its own marker, or a worktree-group match to a DIFFERENT id's marker) AND is NOT live by <code>isSiblingPartitionLive</code> — deliberately stricter than, and not a duplicate of, the pre-existing <code>foldArchivedRegistryRows</code> migration (an older, unrelated fix whose own safety gate protects any row with a self-consistent-sessionId descriptor regardless of actual liveness, by design, so it does not — and was never meant to — catch this shape). A row with NO archive link at all is NEVER a candidate (that stays healOrphanPartitions' job — a live-run repro found the first version wrongly removed such a row with its unread mail simply stranded); a candidate whose only "archive link" is its own marker (no other id at that worktree to forward into) or whose forward attempt fails or leaves mail unaccounted for is classified <code>unhealable</code> and left in place, reported, never guessed at. Each candidate's classify+forward+remove runs under the SAME per-id lock (<code>withIdLock</code>) heal/fold/group-retire use, with the row re-read fresh inside the lock, so a concurrent register/ensure is never raced. Decision logic is shared with <code>companion/lib/devswarm-orphan-policy.js</code> (the parent-inbox "orphaned mesh" warning suppressor), which now calls the SAME <code>resolveArchiveGate</code> heal calls instead of the bare marker check it used to mirror — keeping its own documented NON-DRIFT contract intact (verified against <code>tests/companion/devswarm-orphan-policy-equivalence.test.js</code>). Removal is HUMAN-INITIATED ONLY, via a NEW explicit, opt-in doctor flag — <code>doctor --repair-resurrected [--apply]</code> — same posture as <code>--repair-ingest-orphans</code>/<code>--repair-test-stores</code>: default (no <code>--apply</code>) is a dry-run that prints the plan and writes nothing, <code>--apply</code> executes it. <strong>R3 fix:</strong> this pass was FIRST wired into doctor's default AUTO-SAFE repair pass (<code>migrationFix</code>), which meant a BARE <code>doctor</code> invocation — the exact command the anti-hall-activate skill runs — removed resurrected rows with NO operator intent, contradicting its own "human-initiated only" documentation. It has been REMOVED from the default repair pass entirely and the new flag added to <code>DO_REPAIR</code>'s own exclusion list (alongside <code>--repair-ingest-orphans</code>/<code>--repair-test-stores</code>), so the NEW <code>--repair-resurrected</code> pass itself can never fire under a bare/<code>--fix</code>/ <code>--repair</code>/<code>--dry-run</code> doctor run — that scoping covers only this new pass, not resurrected rows in general: the pre-existing <code>fold-archived-rows</code> migration (<code>foldArchivedRegistryRows</code>, unchanged since before 0.99.1) still runs automatically inside <code>DO_REPAIR</code> on a bare doctor invocation, and retires any descriptor-less row whose worktree carries a matching archived marker via its own forward-then-tombstone (unread mail forwarded first, then the row removed) — it is the one automatic path that already existed for this shape; <code>--repair-resurrected</code> is additive, for the stricter liveness-proven shape <code>foldArchivedRegistryRows</code>'s own gate does not cover. A new unconditional, report-only DETECT section (<code>checkResurrectedRows</code>, "check mode included") surfaces the candidate count and the exact <code>--repair-resurrected</code> command on every plain <code>doctor</code> run AND <code>doctor --check</code>, without writing anything. <code>update.js</code>'s <code>reRetireResurrectedPostUpdate</code> is REPORT-ONLY: it detects on every update (one-time per-version stamped, same shape as <code>cursorHygienePostUpdate</code> — the report itself prints once, not the removal) and tells the operator the candidate count plus the exact command (<code>doctor --repair-resurrected --apply</code>) to run; it never calls the write path itself.</li> <li><strong>Fixed (P1): the ingest daemon's launchd/systemd unit never pinned HOME, so a daemon installed under a non-default HOME wrote into the real <code>~/.anti-hall</code> store</strong> (defect d1c57e67998f, field-verified — a review agent's temp-HOME experiment, label <code>...r3repo-bare-i7ycii-cc6261</code>, registered a real launchd job whose live process then wrote into the operator's real store, because <code>devswarm-ingest.js</code> resolves its store root via <code>os.homedir()</code> and the scheduler hands its unit the SCHEDULER's own default HOME, not the installer's). <code>unitEnvFor</code> (<code>companion/ install-devswarm-ingest.js</code>) now unconditionally pins <code>HOME</code>/<code>USERPROFILE</code> (the installer's own resolved <code>os.homedir()</code>) into every unit shape — launchd's <code>EnvironmentVariables</code>, systemd's <code>Environment=</code>, and the cron fallback's assignment prefix — even when hivecontrol/exec cannot be resolved (previously PATH/HIVECONTROL alone gated whether ANY environment was baked at all). The installer also now REFUSES (forces dry-run, prints a loud stderr notice) when the resolved HOME is under <code>os.tmpdir()</code> OR under <code>/tmp</code>/<code>/private/tmp</code> (checked explicitly — on macOS <code>os.tmpdir()</code> resolves to the per-user <code>$TMPDIR</code> under <code>/var/folders/...</code>, not <code>/tmp</code>, so a HOME planted directly under <code>/tmp</code>/<code>/private/tmp</code>, e.g. a session scratchpad path, previously sailed past the guard entirely), unless <code>ANTIHALL_INGEST_ALLOW_TMP_HOME=1</code> — closing the class the existing <code>NODE_TEST_CONTEXT</code> guard cannot see (a non-<code>node --test</code> run, e.g. a manual or agent experiment, under a scratch HOME). <code>doctor</code> now WARNs (report-only, never auto-repairs) when an already-installed unit's plist/service lacks a <code>HOME</code> key, naming reinstall as the fix.</li> </ul> <h2 id="0990-2026-09-08">0.99.0 (2026-09-08)<a class="headerlink" href="#0990-2026-09-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed (P0): one cursor per row id was shared by every process reading under that id, so whichever instance acked first consumed the mail for all of them</strong> (defect 8b211241bbe9). A second instance's <code>read-primary</code> returned 0 while the cursor had already advanced past rows it was never shown; twin rows (a meshId row and its uuid twin) made this routine rather than exotic. Each INSTANCE now keeps its own cursor at <code>cursors/<id>#inst-<short6>.json</code>, keyed by the existing per-process <code>instanceNonce</code>. A reader's window is <code>max(baseline, own instance cursor)</code>, and the shared pair is raised only to the MIN across instances (a running max of that min — <code>ackTo</code> is monotonic, so it never rewinds) rather than to any one reader's position, so a lagging peer is never skipped. No liveness oracle is consulted anywhere — none is available, since <code>inbox tick</code> refreshes only the heartbeat file and a quiet-but-live reader would age out of any outbound-row test.</li> <li><strong>Added: a loss-free baseline</strong> at <code>cursors/<id>#base.json</code>, moved ONLY by writers whose advance is loss-free by construction (a fold, after rows are forwarded into the survivor; reap-orphans, after a verified archive) and seeded once from the pre-fix <code>max(cursors/<id>.json, store cursor)</code>. Every pre-0.99 installation therefore resumes exactly where it left off, and an instance that has never read starts from the baseline rather than from a peer's position.</li> <li><strong>Fixed: a cross-worktree caller could ack a cross-linked twin's partition.</strong> <code>siblingAckGate</code>'s SELF short-circuit compared only ids and sessionIds — no location component — so a caller standing in the parent's worktree while holding a child's meshId consumed that child's mail. SELF now additionally requires the caller's cwd to resolve to the partition row's own worktree, failing OPEN to the previous verdict when the row carries no worktree path. <code>--ack-as-owner</code> is unaffected.</li> <li><strong>Added: a cursor write journal</strong> at <code>cursor-log/<repoKey>.ndjson</code> (append-only, capped at 2000 records with tail-preserving rotation). Every partition-cursor mutation records id, partition, callerId, namespace, from, to, delivered, pid, instance nonce, gate, verb and cwd. <code>callerId !== partition</code> and an advance with <code>delivered:0</code> are the two signatures that name a cursor eater the moment it recurs — this defect was diagnosed twice from symptoms alone because no writer left a trace. Broadcast cursors, the migrate-time cursor merge, and <code>.seen-</code> watermark writes are deliberately OUT of scope.</li> <li><strong>Fixed: the migrate-time baseline raise could consume a declared instance's mail.</strong> The raise added for the migrate cursor merge was unbounded, and the value it passes is <code>max(dst.cursorValue, src.cursorValue)</code> — shared-pair numbers an older build's own-position ack can have written. Migrate copies rows between backends and makes nothing reachable for a 0.99 instance, so that raise is not loss-free. Reproduced live: a declared instance sitting at 0 received 0 rows instead of 3. Every non-fold, non-reap raise is now BOUNDED by the declared floor (the min across existing instance cursors); fold and reap keep the unbounded raise because they forward or archive the rows first.</li> <li><strong>Fixed: the retired-redirect override refused a survivor with no live owner.</strong> The guard used the fail-toward-live liveness predicate, so an undetermined verdict read as live and refused the override — breaking the one path it exists to serve. It now requires POSITIVE evidence (a fresh heartbeat for the survivor) before refusing. Note the direction this cuts: <code>--ack-as-owner</code> now acks a survivor whenever its heartbeat is stale or missing, so a live twin mid-long-turn (the heartbeat only refreshes once per prompt) can lose mail to an explicit human override. That is the intended behavior for a stuck-child override — a human invoking it is asserting the survivor is not actually consuming, and the override is not meant to defer to a heartbeat that simply hasn't ticked yet.</li> <li><strong>Fixed: the cursor journal reported <code>delivered: 0</code> on a store-only read.</strong> The count was gated on the union flag, which <code>--ack</code> sets even when no union runs, so no row was ever tagged and a read that delivered rows recorded zero — a false positive of the exact signature the journal exists to make trustworthy. Own rows are now counted by what they are, not by which flag was set.</li> <li><strong>Added: doctor names pre-release cursor files in the old dot shape</strong> (<code><id>.inst-<6hex>.json</code>, <code><id>.base.json</code>) and never deletes them — such a name can equally belong to a real workspace, so removing it could destroy a live read position.</li> <li><strong>Fixed: an older build in a mixed fleet could consume a 0.99 instance's mail.</strong> The baseline briefly re-adopted the shared cursor as a LIVE floor on every read. That is safe only within this version: a 0.98.3 session's ack writes the shared pair to its own position with no min-projection, so re-adopting it raised every 0.99 instance's floor. Reproduced against the real cached 0.98.3 build — the old build received 3 messages and a declared 0.99 instance then received 0. The baseline is now seeded ONCE (upgrade continuity) and never re-adopts the shared value; the writers whose advances are genuinely loss-free raise it at their own call sites instead, including the migrate-time cursor merge in <code>companion/devswarm-migrate.js</code>.</li> <li><strong>Fixed: a workspace id could collide with anti-hall's own cursor filenames.</strong> <code>isSafeId</code> permits dots, so a workspace legitimately named <code>w.base</code> had the legacy cursor path <code>cursors/w.base.json</code> — byte-identical to workspace <code>w</code>'s baseline path under the first cut of this feature. Acking that workspace's cursor to N made <code>w</code>'s baseline read N and silently skipped N rows of <code>w</code>'s mail (reproduced live). The three cursor namespaces introduced here now use <code>#</code> as their separator (<code><id>#base.json</code>, <code><id>#inst-<6hex>.json</code>, <code><id>#nd-<6hex>.json</code>), and <code>#</code> is a character <code>isSafeId</code> forbids — so the collision is impossible by construction rather than by a validator every future call site must remember. A FRESH registration also refuses an id carrying <code>#</code>, <code>.seen-</code>, <code>.inst-</code>, <code>.nd-</code> or a trailing <code>.base</code>; an install that already holds such a row keeps working.</li> <li><strong>Added: bounded hygiene for instance cursor files</strong>, shipped in BOTH <code>update.js</code> (one-time per version) and doctor (report-only unless repairing). Both cursor namespaces are swept (<code>#inst-</code> for store rows and <code>#nd-</code> for the descriptor's NDJSON lines, grouped separately because they count in different index spaces); missing the second left a dead instance pinning the descriptor cursor forever. Only a STALE file (mtime past <code>DEFAULT_INSTANCE_CURSOR_STALE_MS</code>, 7 days) is ever a candidate — a fresh file is a live reader's position. A stale file is deleted when removing it does not advance the floor past another instance, and otherwise evicted with a journaled <code>gc-evict</code> record. Note what deletion costs in the SOLE-file case: with no other instance file left, that id falls back to its baseline, so the next read replays everything since the baseline. That is redelivery, never loss, and it is journaled. Names this code could not have written are never touched: the parser requires a six-hex nonce.</li> <li><strong>Added: <code>doctor --repair-test-stores [--apply]</code></strong> (defect be2c6c9e81a1). Doctor already inventories store entries whose recorded repo path is under the system temp dir and no longer exists; this adds the explicit, opt-in removal path for exactly that subset. Dry-run by default (it prints the plan and deletes nothing); <code>--apply</code> executes it, re-verifying each entry's eligibility immediately before deleting. Never folded into a plain <code>doctor</code>, <code>--fix</code> or <code>--dry-run</code> pass — same narrow, human-invoked posture as <code>--repair-ingest-orphans</code>. Contributed alongside this release; the broader task-#10 inventory work ships separately.</li> <li><strong>Changed (behaviour): the devswarm store refuses to fall back to the real home while running under <code>node --test</code>.</strong> <code>openStore</code>, <code>computeSummary</code> and <code>deriveSummary</code> previously resolved <code>o.home || os.homedir()</code>, so a test that forgot to pass an explicit <code>home</code> silently wrote a fixture registry row into the developer's own <code>~/.anti-hall/devswarm/store/</code>. Under <code>NODE_TEST_CONTEXT</code> (set by <code>node --test</code> and inherited by spawned children) that fallback now throws instead, naming the missing <code>home</code>. Outside a test run nothing changes. This is a deliberate behaviour change: a leaky test now FAILS rather than quietly polluting the machine.</li> <li><strong>NOT fixed in this release, characterized only: report item 3</strong> — a sibling ack advancing past rows that were never delivered (<code>part.cursor + physicalConsumed</code>, where <code>physicalConsumed</code> can exceed the delivered count for hash-suppressed or unparseable rows). This is a SEPARATE mechanism from the shared-cursor defect above and per-instance cursors do NOT fix it. The behavioural fix (quarantine an undeliverable row before the cursor passes it) needs the suppressed ROWS, and the code currently carries only a COUNT, so it ships separately. This release pins the mechanism with characterization tests and makes a recurrence mechanically detectable: every cursor record carries <code>delivered</code>, so an advance with <code>delivered:0</code> is visible without re-derivation.</li> <li><strong>Fixed: the retired-sender hint offered a substitutable <code><id></code> placeholder.</strong> <code>devswarm-parent-inbox.js</code> printed <code>inbox ack <id> --ack-as-owner</code> while naming the retired sender only in the surrounding prose. Since <code>--ack-as-owner</code> is exempt from every ownership gate, an agent filling that placeholder with a live child's id would consume that child's mail. The exact retired id is now interpolated into the command, with an explicit warning never to ack a live child id.</li> </ul> <h2 id="0983-2026-09-08">0.98.3 (2026-09-08)<a class="headerlink" href="#0983-2026-09-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: a test file leaked real LaunchAgent registrations onto the host machine</strong> (defect ec33954162ef). <code>tests/scripts/devswarm-fleet-2e8653787945.test.js</code> built a <code>selfHeal</code> ctx with no <code>ctx.io.spawnInstaller</code> mock, so <code>selfHeal</code>'s stale-daemon branch fell through to the REAL <code>defaultSpawnInstaller</code> (<code>scripts/devswarm.js</code>), which spawned <code>install-devswarm-ingest.js</code> for real under a throwaway temp <code>HOME</code>. That subprocess registered a genuine <code>KeepAlive</code> LaunchAgent whose <code>WorkingDirectory</code> pointed into the temp HOME; teardown deleted the HOME but never unloaded the registration, so launchd retried it forever (exit 78, "program gone") — confirmed live on the maintainer machine as 50+ loaded <code>com.anti-hall.devswarm-ingest.*</code> labels against 6 real on-disk plists. Fixed with an <code>io.spawnInstaller</code> mock (the same pattern <code>tests/companion/ingest-health.test.js</code> already used) plus a belt-and-braces regression test proving <code>install-devswarm-ingest.js</code>'s existing <code>ANTIHALL_INGEST_DRY_RUN=1</code> seam makes even the REAL, unmocked spawn path a no-op.</li> <li><strong>Added: a structural test-context guard closes the CLASS of the LaunchAgent leak above, not just the one fixed instance</strong> (same defect, fix-wave R2 item 7). <code>install-devswarm-ingest.js</code> now ALSO forces its own dry-run seam whenever <code>process.env.NODE_TEST_CONTEXT</code> is present — Node sets this in every <code>node --test</code> worker, and a <code>spawnSync</code>'d child inherits it by ordinary env inheritance (verified live with a probe test before relying on it). So a FUTURE test that makes the identical mistake (forgets <code>ctx.io.spawnInstaller</code>/<code>ANTIHALL_INGEST_DRY_RUN=1</code>) is still safe, with no opt-out and no per-test convention to remember. Prints one stderr line naming the defect, but only at the moment a real write/rm/spawn call is actually intercepted (not at module load) — so merely <code>require()</code>-ing this module under <code>node --test</code> stays silent, and the notice appears only when a real mutation was genuinely prevented. New regression test spawns the real installer <code>main()</code> under an isolated HOME with neither <code>--dry-run</code> nor <code>ANTIHALL_INGEST_DRY_RUN</code> set and asserts zero plist/service files written.</li> <li><strong>Fixed: two installer paths still bypassed the structural test-context guard above</strong> (Critic R2, same defect class). <code>install-devswarm-ingest.js</code>'s <code>installCron</code>/<code>uninstallCron</code> called <code>spawnSync('crontab', ['-'], {input})</code> directly instead of through <code>planRun</code>, so on Linux neither <code>--dry-run</code> nor the <code>NODE_TEST_CONTEXT</code> guard protected the crontab (only the plist/service writes were covered) — now routed through <code>planRun</code>. <code>install-reaper.js</code> had no <code>NODE_TEST_CONTEXT</code> guard at all (<code>DRYRUN</code> was <code>args.includes('--dry-run')</code> only); it now applies the identical <code>EXPLICIT_DRYRUN || NODE_TEST_CONTEXT</code> guard, with the same once-per-process stderr notice, as <code>install-devswarm-ingest.js</code>. The structural fix now spans every installer path that can register a real launchd/systemd/cron job: <code>install-devswarm-ingest.js</code> (plist/service writes AND the crontab fallback) and <code>install-reaper.js</code> (plist/service writes).</li> <li><strong>Added: <code>doctor</code> detects and can repair orphaned launchd/systemd ingest registrations</strong> (same defect, ec33954162ef). A plain <code>doctor</code> run now always prints an "Orphaned launchd/systemd ingest registrations" table — enumerated from the SCHEDULER'S OWN registration list (<code>launchctl list</code> / <code>systemctl --user list-units</code>), not from disk, so it catches a loaded label with no matching plist/service file at all — exactly the class the existing <code>git worktree list</code>-driven reap (D9) is structurally blind to once the worktree is gone. Silent when everything classifies <code>healthy</code>. <code>doctor --repair-ingest-orphans</code> previews the exact unload plan (dry-run by default, and no longer also triggers doctor's unrelated full auto-repair pass — fix-wave R2 usability fix); <code>doctor --repair-ingest-orphans --apply</code> executes it (<code>launchctl bootout gui/$(id -u)/<label></code> / <code>systemctl --user stop <unit>.service</code> — never <code>kill -9</code>, never deletes a file). <strong>Eligibility is exactly ONE class: <code>orphan-no-plist</code> (no plist/service file on disk at all) AND no live heartbeat/lock for that project/worktree.</strong> <code>orphan-path-gone</code> and <code>duplicate-label-same-project</code> are ALWAYS report-only — a plist DOES exist on disk for both, so unloading either remains the EXISTING <code>reapLegacyUnitsForRepo</code>/<code>stopLegacyUnitEntry</code> job, never this new label-only path. (Fix-wave R2, same defect: the first pass had a P0 — the duplicate-detection cross-check could mark BOTH members of a plist-present, worktree-present group eligible, which would have booted out real registrations; caught by review before merge, fixed with a single point of eligibility assignment plus a fail-closed invariant check. Also fixed: legacy per-worktree liveness now ORs a fresh heartbeat with the lock check, matching the per-project branch.) New exports on <code>install-devswarm-ingest.js</code>: <code>listLoadedIngestLabels</code>, <code>classifyLoadedLabel</code>, <code>orphanReapPlan</code>, <code>bootoutLoadedLabel</code>, <code>stopLoadedUnit</code>; new <code>doctor-repair.js</code> function <code>runIngestOrphanRepair</code>. Documented in <code>docs/KB-devswarm-hivecontrol.md</code> §44 and both Claude/Codex DevSwarm <code>SKILL.md</code> files.</li> </ul> <h2 id="0982-2026-09-08">0.98.2 (2026-09-08)<a class="headerlink" href="#0982-2026-09-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: <code>roster</code>/<code>diagnose</code>/<code>healthcheck</code> failed open on an unreadable registry</strong> (defect 77d5a5bbf614). <code>computeDiagnosis</code> and <code>cmdRoster</code> called <code>listRegistry()</code>/<code>computeSummary()</code> directly with no <code>getReadError()</code> probe — devswarm-store.js's <code>readAll()</code> swallows a genuine registry.ndjson read error (EACCES on a chmod-000 store dir, etc.) to an empty array, so a registry.ndjson outage read back as "0 workspaces, healthy" instead of surfacing the failure. All three commands now probe <code>getReadError()</code> right after their own registry-read call and report <code>known:false</code>, <code>storeUnavailable:true</code>, <code>storeUnavailableReason:<code></code>; <code>diagnose</code> and <code>healthcheck</code> also flip <code>degraded:true</code>/<code>status:'store-unavailable'</code> so an unreadable store can never look like a clean report.</li> <li><strong>Partial: <code>devswarm-parent-inbox.js</code> (the per-turn workspace table) skips the expensive transcript-mtime liveness read for an already-archived row</strong> (defect bf965e5729c5, ruled <code>partial</code>). An archived row's dormant/idle-alive result is never consulted by its <code>archived</code>/ <code>not-draining</code> display branch, so <code>readActivityTs</code>/<code>rowLivenessState</code> are now skipped for it; the cheap heartbeat read itself stays unconditional so the table's "last" column never goes stale for an archived row that still emits heartbeats. A cross-turn disk cache was tried and then <strong>removed</strong> on review (net cost for this hook's actual per-turn cadence, plus two correctness bugs — see the defect's ruling for detail). The field-reported 1.07s/turn-under-load latency this defect describes was <strong>not reproduced locally</strong> and remains open.</li> <li><strong>Added: a <code>git-stash-guard</code> command-guard branch blocks a mutating <code>git stash</code> (push/pop/drop/clear/apply/save, or the bare <code>git stash</code> == push shorthand, including flag-only forms like <code>-u</code>/<code>-k</code>/<code>-m</code>/<code>-p</code>/<code>-q</code>/<code>-a</code> and invocations using git's own global options like <code>-C <path></code>/ <code>--git-dir=X</code>)</strong> (defect b08b26566b92). The guard only fires once ARMED — a <code>.anti-hall/protected-stashes</code> marker at the repo's git toplevel, or <code>ANTIHALL_STASH_GUARD=1</code> — in both subagent and coordinator context (no unconditional default block). <code>git stash list</code> is unaffected, and a command that merely <em>mentions</em> "git stash" in a quoted argument (a grep pattern, a commit message) is never misclassified as an invocation. The guard's own skip name is in skip-guard.js's <code>DESTRUCTIVE</code> set (same protection level as <code>git-guard</code> — a blanket <code>"all"</code> skip cannot silence it).</li> </ul> <h2 id="0981-2026-09-08">0.98.1 (2026-09-08)<a class="headerlink" href="#0981-2026-09-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: <code>devswarm-child-gate.js</code>'s Stop hook re-fired every turn even though the child had followed the gate's own instructed heartbeat</strong> (defect a55d6b71a76f). Three compounding root causes:</li> <li>When <code>heartbeat --summary</code> was benignly refused (an unresolvable/ unregistered caller identity, or an ownership mismatch — <code>BENIGN_MESH_BROADCAST_REASONS</code>), the broadcast never reached the shared store's <code>recent[]</code> projection, so the gate's <code>alreadyReportedThisEpisode()</code> (which reads only <code>recent[]</code>) could never see that the child DID attempt to report — the same failing heartbeat command was re-prescribed forever. <code>cmdHeartbeat</code> now appends a local, bounded (last 50 lines, trimmed only once a file passes 100) row to a PER-WRITER-ID attempt file — <code>devswarm/summary-attempts/<repoKey>/<writerId>.ndjson</code>, never one file shared by every writer for a repoKey (that shape was a read-modify- write-rename race: two concurrent sibling writers could silently drop each other's row) — stamped with the WRITING process's own <code>instanceNonce</code> and a <code>sessionId</code> derived from the cwd-verified process-tree walk, with <code>CLAUDE_CODE_SESSION_ID</code> accepted only when the walk corroborates it (omitted otherwise), NEVER the caller-supplied <code>--session</code> flag (a prior shape trusted the flag directly and was provably forgeable: <code>heartbeat <victim-id> --summary x --session <victim's own sessionId></code> produced a record the victim's own gate accepted). A record is accepted iff it is nonce-authenticated, or matched to the workspace's registered session across its id forms (its own <code>workspaces/<id>.json</code> descriptor, or — when that descriptor is absent, e.g. a child heartbeating under its meshId while its env id is a separately-unregistered UUID — a descriptor provably the SAME identity on the SAME physical worktree: a uuid-prefix re-registration of the env id, or one whose worktree resolves to the same canonical meshId — never any same-worktree descriptor unconditionally), matched <strong>regardless of the record's own <code>id</code> field</strong>; an unauthenticated record does not satisfy. When a record for this exact id exists but authenticates against neither check (e.g. a genuine record from a PRIOR OS process — the nonce fallback changes on every restart), or when the gate's own nonce cannot be derived at all, it now logs ONE stderr diagnostic per session instead of silently re-blocking with no trail. If a block still fires anyway (e.g. a known unread-inbox backlog), the block text names the drop reason from a fixed whitelist only (never the raw stored string, which is not trusted input) and its remedy, instead of re-prescribing the exact command that just failed. <code>warnIdMismatch</code> (the separate <code>heartbeat</code>/<code>inbox tick</code> id-mismatch stderr warning) now gates on <code>isChildWorkspaceCorroborated</code>, not the bare env-only <code>isChildWorkspace</code>, so a Primary with a leaked <code>DEVSWARM_SOURCE_BRANCH</code> is never told to switch ids.</li> <li>The per-window forced-ack cap (<code>MAX_BLOCKS=2</code>) fully reset every <code>RESET_MS</code> (5 min), so it could re-arm indefinitely across a long session. A new, never-reset <code>MAX_BLOCKS_PER_SESSION=6</code> lifetime bound stops all further blocking for the rest of the session once reached (logged once to stderr).</li> <li> <p><code>isChildWorkspace()</code> trusted <code>DEVSWARM_SOURCE_BRANCH</code> alone, so a Primary that inherited a leaked env var could be gated as a child. A new <code>isChildWorkspaceCorroborated()</code> (<code>hooks/lib/devswarm-role.js</code>) additionally requires on-disk evidence — a registered <code>workspaces/<id>.json</code> descriptor, or cwd under the real DevSwarm worktree layout (<code>~/.devswarm/repos/...</code>) — before <code>devswarm-child- gate.js</code> treats a session as gate-eligible; no corroboration is a silent no-op. Codex needs no separate fix — <code>devswarm-child-gate.js</code> and <code>devswarm-role.js</code> are shared files, registered unmodified in <code>codex/hooks/hooks.json</code>.</p> </li> <li> <p><strong>Fixed: a child substituted the WRONG id (its own meshId) into wake/heartbeat/tick instructions</strong> (defect 735b179362e8) — every emitted wake/tick/heartbeat/read-primary instruction (<code>hooks/lib/devswarm-wake.js</code>'s <code>drainCmd</code>/<code>wakeDirective</code>/<code>wakeReassert</code>, and <code>hooks/devswarm-child-turn.js</code>'s <code>REMINDER</code>/<code>RECEIVE_NUDGE</code>) previously embedded the literal <code><DEVSWARM_BUILDER_ID></code> placeholder unconditionally, even though the real id is available in <code>env</code> at every call site — a child then had nothing to substitute and addressed the wrong mesh partition. Both files now substitute the REAL <code>DEVSWARM_BUILDER_ID</code> (validated against the same safe-id charset every other workspace id in this codebase is checked against) whenever it is present; the placeholder is kept, byte-identical to before, when the env var is absent or fails validation — never a bad/unsafe value is interpolated. Additionally, <code>cmdInboxTick</code>/<code>cmdHeartbeat</code> (<code>scripts/devswarm.js</code>) now warn (stderr, once per call, fail-open — never refuse) and set <code>idMismatch:true</code> in their JSON result when a CHILD workspace addresses an id other than its own real <code>DEVSWARM_BUILDER_ID</code>, naming both ids so the mismatch is visible without blocking a caller that has a legitimate reason to address a different id.</p> </li> <li> <p><strong>Added: a subagent inside a DevSwarm child workspace can no longer touch the shared mailbox</strong> (defect f0958b13fe2b, field-measured by a downstream project 2026-09-08 — 155 executions across 120 subagent transcripts in one workspace ran <code>devswarm.js inbox pull/ack</code> directly, each advancing the shared cursor and causing the workspace's own main thread to silently miss mail; brief-level prohibitions alone were proven non-mitigating). Hardened over three review rounds; the final blocked-verb set and mechanism:</p> </li> <li><code>hooks/command-guard.js</code>'s new <code>devswarm-subagent-mailbox-guard</code> block blocks any Bash invocation of <code>devswarm.js</code>'s <code>inbox pull|ack|read| read-primary|tick</code>, top-level <code>heartbeat</code>, top-level <code>reap-orphans</code> (writes cursors on reap), <code>inbox messages ... --ack</code>/<code>--ack-as-owner</code> (the documented, cursor-advancing expansion of <code>read-primary</code> — NOT the same as the safe non-acking <code>inbox messages</code>), <code>mesh read</code> without <code>--peek</code>/<code>--seq</code> (advances the broadcast cursor by default), <code>roster --ack</code> (an alias of <code>mesh read</code>, D23), and top-level <code>register</code>/<code>archive</code> (both advance cursors through <code>foldGroupIntoSurvivor</code> — a subagent never legitimately registers or archives a workspace; the separate <code>register-primary</code> and <code>archive-request</code>/<code>archive-ignore</code>/<code>archive-unignore</code>/<code>unarchive</code> verbs are unaffected). Read-only verbs stay allowed: <code>inbox count</code>, <code>inbox messages</code> (incl. <code>--tail</code>, without an ack flag), <code>inbox peek-primary</code>, <code>mesh read --peek</code>/<code>--seq N</code>, plain <code>roster</code>, <code>send</code>. Detection runs per shell segment with a flag-skip pattern that consumes an optional VALUE after each flag (a valued flag BEFORE the verb — <code>--session X inbox ack Y</code> — previously broke the match entirely and bypassed the guard). Two overrides: <code>~/.anti-hall/skip.json</code> under <code>devswarm-subagent-mailbox-guard</code>, or env <code>ANTIHALL_ALLOW_SUBAGENT_MAILBOX=1</code>.</li> <li><code>hooks/verify-first-subagent.js</code> appends one line to its SubagentStart injection when <code>isChildWorkspace(env)</code> is true, telling the subagent the main thread owns the mailbox and to report findings to its parent instead.</li> <li><strong>Root cause of the field incident, fixed directly:</strong> <code>hooks/devswarm-child-drain.js</code> (PostToolUse, matcher Bash, CHILD-ONLY) was the hook actually TELLING subagents to drain the mailbox — it gated only on <code>isDevswarmActive(env) && isChildWorkspace(env)</code>, both env-based and therefore true for a subagent's own tool calls too (the child's env is inherited), and its injected text literally read <code>Drain NOW via \</code>inbox pull ... && inbox ack ...`<code>. This is the exact command the three field-measured subagent runs executed. Now silently no-ops for subagent context instead of injecting anything — a subagent already gets the one-line rule from</code>verify-first-subagent.js<code>at spawn, and repeating it on every Bash call would be exactly the per-call noise this hook's own THROTTLE design exists to avoid. Swept every other PreToolUse/ PostToolUse-registered DevSwarm hook gated on child env alone for the same hole:</code>devswarm-child-turn.js<code>(UserPromptSubmit) and</code>devswarm-child-gate.js<code>(Stop) are NOT subagent-reachable at all (neither event fires for a Task-tool subagent —</code>SubagentStop<code>is a distinct, unregistered event); the two Primary-side hooks (</code>devswarm-parent-gate.js<code>,</code>devswarm-parent-reply-tracker.js<code>) return early for a child and are unaffected.</code>devswarm-child-drain.js` was the only live hole.</li> <li><strong>Payload-only subagent signal for both blocking gates above:</strong> both (1) and (3) now key off a new <code>isSubagentByPayload(payload)</code> (<code>hooks/coordinator-detect.js</code>) — <code>agent_id</code>/<code>agent_type</code> in the hook payload ONLY, no <code>CLAUDE_CODE_ENTRYPOINT=agent_tool</code> env fallback. The general-purpose <code>isSubagent()</code> (used by command-guard's normal coordinator-only gate, unchanged) legitimately uses that env fallback, but a child workspace's env is inherited by its entire process tree — using the same fallback for a BLOCKING gate could let a leaked <code>agent_tool</code> value (from how the child session itself was originally spawned) permanently misclassify that workspace's own main-thread cron tick / Monitor wake as a subagent, blocking it from its own mailbox.</li> <li><strong>Codex parity, precisely stated:</strong> both guards are registered via SHARED hook files in <code>codex/hooks/hooks.json</code> — no separate Codex code path. Their DENY behavior, however, depends on the harness actually supplying <code>agent_id</code>/<code>agent_type</code> in the hook payload; this has been <strong>verified on Claude Code only</strong>. A grep of <code>plugins/anti-hall/codex</code> for <code>agent_id</code>/<code>agent_type</code> returns 0 hits (<code>codex/README.md:48</code> confirms no such payload-marker mapping exists there), so whether Codex's harness populates these fields the same way is unverified — the guards are registered either way (fail-open if the markers are absent, same as any unmatched context), but blocking a Codex subagent specifically has not been demonstrated.</li> <li><strong>Known, harmless (deferred):</strong> a Wave R3 review pass flagged echo noise around this guard's deny path; triaged as known and harmless rather than fixed in this round — see KB §42 for the note. See <code>docs/KB-devswarm-hivecontrol.md</code> §42 for the full mechanism.</li> </ul> <h2 id="0980-2026-09-06">0.98.0 (2026-09-06)<a class="headerlink" href="#0980-2026-09-06" title="Permanent link">¶</a></h2> <ul> <li><strong>Changed: <code>inbox count</code>/<code>inbox read</code>'s <code>storeUnavailable</code> field is now a BOOLEAN, not an object</strong> — 0.97.1 and earlier returned <code>storeUnavailable</code> as a richer OBJECT (<code>{reason, error, registeredRepoKey, callerRepoKey, storeUnavailableReason}</code>) on <code>count</code>/<code>read</code>, while <code>read-primary</code>/ <code>peek-primary</code>/<code>messages</code>/<code>ack</code> already reported it as a bare boolean — two different shapes for the same field name depending on which verb you called. 0.98.0 unifies every read verb on the SAME shape: <code>storeUnavailable</code> is always a boolean (<code>true</code> only for a genuinely unreadable store — <code>store-unavailable</code>/<code>store-open-failed</code>; <code>false</code> for a more specific, non-genuine refusal like <code>project-context-mismatch</code>), the underlying fs error code is a sibling top-level <code>storeUnavailableReason</code> (string|null), and the full former object — <code>reason</code>/<code>error</code>/<code>registeredRepoKey</code>/ <code>callerRepoKey</code>/<code>storeUnavailableReason</code> — is preserved verbatim under a NEW <code>storeUnavailableDetail</code> key on <code>count</code>/<code>read</code>/<code>ack</code> (no information lost, just relocated). No in-repo consumer (hook, skill, or script) read the old object shape directly — this is a heads-up for any EXTERNAL JSON consumer parsing <code>inbox count</code>/<code>inbox read</code> output that this field's type changed.</li> <li><strong>Fixed: a fold-time retired tombstone no longer strands a re-registered id's mail</strong> — a read of an id whose only descriptor is a retired-redirect tombstone (73303d4c098b) now takes the one-hop redirect a caller with nothing live behind <code>id</code> needs, the read-side mirror of the existing send-side redirect fix; a caller with its own live row is never redirected.</li> <li><strong>Fixed: <code>register-primary</code> refuses a <code>live-primary-conflict</code> instead of silently double-registering</strong> — registering as Primary over a provably-live OTHER session for the same worktree now refuses with <code>live-primary-conflict</code> naming the live session; <code>--force</code> overrides it explicitly (7d0a948031cd).</li> <li><strong>Added: <code>mesh read --peek</code>/<code>--seq</code></strong> — a non-mutating peek at the shared mesh broadcast log and per-message sequence numbers for precise resume points (d68c561e1649).</li> <li><strong>Fixed: repoKey keying inside a git submodule now resolves to the superproject</strong> — a caller invoked from inside a submodule used to key identity/repoKey to the submodule instead of the superproject, flipping <code>registeredRepoKey</code> between a submodule invocation and a superproject one and failing closed as <code>project-context-mismatch</code> (d56bfaac2da0).</li> <li><strong>Added: a <code>forwarded</code> flag</strong> on a row relayed through a fold/redirect, so a reader can tell a forwarded row from an original one (e9e7c99ec924).</li> <li><strong>Added: <code>retryAfterMs</code> and <code>possiblyStaleRegistry</code></strong> on a transient registry-read failure, so a caller can distinguish "retry shortly" from a genuine refusal (2e8653787945).</li> <li><strong>Added: per-process <code>instanceNonce</code>, roster <code>instances</code>/<code>instance-split</code>, <code>diagnose instanceSplits</code>, and an <code>@short</code> sender tag on <code>read-primary</code></strong> — two live processes of the SAME session id (a <code>claude --resume</code> racing its own prior process, or a fork) no longer have their mesh rows/messages silently attributed to each other; roster/diagnose now surface the split so it is visible instead of silent (d3d571495bf6).</li> <li><strong>Fixed: <code>wake-watch</code> now takes over a stale-live lock, restamps its own liveness, and exits cleanly on a lost lock</strong> — the REFUSED stderr line now also names the CURRENT lock HOLDER's pid/age/version (not the refused watcher's own), so an operator sees who actually holds the lock (8143ced316d3).</li> <li><strong>Added: a shared <code>computeRowLive</code> helper and a <code>phantom</code> hint</strong> on a registry row that looks live but carries no verifiable liveness evidence (298b79969409).</li> <li><strong>Fixed: <code>isArchivedForRouting</code> now keys on <code>sessionId</code></strong> so an archived marker left by a PREVIOUS occupant of a reused id no longer shadows routing decisions for the CURRENT live occupant of that same id (d386d8a610b7).</li> <li><strong>Added: <code>send --answers</code></strong> — reply correlation so a direct reply to a blocking question can be matched back to the question it answers (93c41cc09ff6).</li> <li><strong>Fixed: an app-archived partition no longer stays listed as a STALE WORKSPACE in the parent-inbox notification</strong> — <code>devswarm-parent-inbox.js</code>'s <code>staleRegistryPartitions</code> table now also recognizes an app-level archive marker (the owner archiving the workspace in the DevSwarm app), not only anti-hall's own internal archive tombstone, before naming a row "STALE WORKSPACE" (a9ac2fc7e368). Scoped to that one table — <code>diagnose</code>/roster's separate <code>orphans[]</code> surface is NOT covered by this fix (see Known).</li> <li><strong>Fixed: read-side <code>known</code>/withheld-state fields and <code>repoKey</code>/<code>storePath</code> meta are now consistent across <code>count</code>/<code>read</code>/<code>messages</code>/<code>read-primary</code>/ <code>peek-primary</code>, and a meshId <code>send --to</code> can resolve is now accepted by those same read verbs (<code>resolvedFrom</code>)</strong> — <code>count</code>/<code>read</code> previously folded <code>meshGroupUnresolved</code> inconsistently into <code>known</code>, and <code>messages</code>/ <code>read-primary</code>/<code>peek-primary</code> carried none of <code>repoKey</code>/<code>storePath</code>/<code>known</code>/ <code>meshPartitionIds</code> at all, so a caller checking those alongside a refusal's <code>reason</code> saw <code>undefined</code>. A same-worktree twin-sibling row read now also returns a real, non-null <code>reason</code> (<code>ownership-mismatch</code>) with the same B1 meta instead of a bare refusal. <code>read-primary</code>'s ownership-gated ack path now resolves a bare meshId BEFORE opening its store (matching what <code>peek-primary</code>/<code>count</code> already did), so the two verbs can no longer diverge on the SAME meshId input (902d3c5e7531, 1932b53a3ace).</li> <li><strong>Fixed: the Codex installer's <code>ANTI_HALL_HOOKS</code> now lists every hook <code>codex/hooks/hooks.json</code> lists</strong> — still a manually-maintained list (not dynamically derived from <code>hooks.json</code>), but a new parity test now enforces the two stay in sync so a drift can't ship silently again; also bumped the doctor's minimum Node to >=22 to match CI, and the scan-throttle write path now carries an explicit flag (GPT-6 audit findings AH01/AH03/AH07).</li> <li><strong>Changed: the Stop-gate mailbox-wake reassert is now a short pointer, not a re-stated prompt</strong> — every Stop-gate firing used to re-inline the FULL CronCreate prompt paragraph (the same text SessionStart already delivered once); it now names the CronList/Monitor conditions in a MEASURED bound (<code>wakeReassert</code>'s own fixed text — everything except the one embedded CLI path — stays <code><= 360 chars</code> by test, verified against both an 86-char and a 160-char cli fixture, for child and Primary, so the contract holds regardless of how long a real install path happens to be; raised from an initial 320 — the measured fixed text peaked at 315 chars, only 5 chars of headroom for future wording changes and backslash-escaped characters, not because any widened clause had tripped it: <code>wakeReassert</code> carries no drain/refusal clause at all, and the measured fixed length is 315 both before and after this release's other wording changes) and, on a miss, points at the new <code>wake-directive <id></code> CLI verb (an on-demand reprint of the full SessionStart text) instead of repeating it inline on every block. The watcher script path is derived from the already-emitted <code>$CLI</code> (<code>WATCH="$(dirname "$CLI")/../companion/lib/devswarm-wake-watch.js"</code>) rather than re-embedded as a second long literal, for the same budget reason the CLI path itself is emitted once — and the <code>CLI=</code> assignment itself is now double-quoted (<code>CLI="<path>"</code>), not backtick-quoted (<code>CLI=\</code><path>``, which in an actual shell means command substitution and would have executed the path as a command instead of assigning it).</li> <li><strong>Changed: a read verb (<code>count</code>/<code>read</code>/<code>ack</code>/<code>messages</code>/<code>read-primary</code>/ <code>peek-primary</code>) now refuses a genuine id collision the SAME way <code>send</code> already does — EXCEPT an EXACT registered-id match, which always wins as a no-op and is never redirected or refused</strong> — <code>resolveReadArgToId</code> used to short-circuit to the exact-id row the MOMENT any row's id exactly matched the literal arg, never even checking whether that same arg ALSO collides with a different live row's derived meshId; <code>send --to</code> already refused that shape as <code>ambiguous-target</code>, so a read verb on the SAME arg silently used the shadowed row instead of refusing. A later pass made read delegate to the same resolution <code>send</code> uses unconditionally — but that reintroduced a DIFFERENT bug: a <code>register-primary</code> row and a same-worktree "twin" derive the identical meshId (<code>canonicalMeshId</code> IS <code>primaryWorkspaceId</code>), so an EXACT read of the Primary's own id could silently land on the twin's partition instead (1932b53a3ace). The rule is now: an exact registered-id match ALWAYS wins as a no-op on a read (never redirected, never reported ambiguous). A non-exact arg resolves the same way <code>send</code> resolves one — the freshest LIVE row among same-worktree twins (<code>resolveMeshTarget</code>) — but can never hit <code>send</code>'s own <code>ambiguous-target</code> refusal on this path: that refusal fires ONLY on an exact-id collision, and resolveReadArgToId already resolves any exact match as the no-op above before delegation ever runs.</li> <li><strong>Fixed: a store that genuinely EXISTS but cannot be READ (<code>EACCES</code>, <code>ENOTDIR</code>/<code>EISDIR</code>, a corrupt sqlite header) no longer reads back indistinguishable from an empty/never-written store, for ANY id — registered or not.</strong> Every read verb (<code>count</code>/<code>read</code>/<code>ack</code>/<code>messages</code>/<code>read-primary</code>/<code>peek-primary</code>) now reports <code>storeUnavailable</code> as a boolean with the underlying error code in the sibling <code>storeUnavailableReason</code> field (string|null), and <code>known:false</code>; <code>emitKnownWarning</code>'s stderr line names <code>store-unavailable (<code>)</code> instead of the tautological <code>storeUnavailable (store-unavailable)</code>. A never-written store dir (<code>ENOENT</code>) stays fail-open, unchanged (902d3c5e7531 extended). The two remaining gaps are now also closed: an UNREGISTERED id (no descriptor) sharing a repo whose <code>registry.ndjson</code> is unreadable while <code>messages.ndjson</code> stays readable now reports <code>store-unavailable</code> via <code>messages</code>/<code>read-primary</code>/<code>peek-primary</code> too, not just <code>count</code>/<code>read</code>/<code>ack</code> — <code>resolveWorkspaceStoreForRead</code>'s existence guard now re-probes <code>getReadError()</code> AFTER its own <code>listRegistry()</code> call, not just before it. And the <code>read-primary</code>/<code>inbox messages --ack</code> ownership check, which resolves the caller's OWN registry row via the same <code>listRegistry()</code>-backed <code>resolveMeshTarget</code>, no longer misreports a caller that genuinely owns <code>id</code> as <code>caller-not-registered</code> when that same registry read fails — it now reports <code>store-unavailable</code> too. Scoped to the six READ VERBS above — <code>roster</code>, <code>diagnose</code>, and <code>healthcheck</code> still read the registry fail-open (see Known).</li> <li><strong>Fixed: the retired-sender INFORMATIONAL ack hint now says <code>--ack-as-owner</code></strong> — the hint in both <code>devswarm-parent-gate.js</code> and <code>devswarm-parent-inbox.js</code> told the Primary to run plain <code>inbox ack <id></code>, which fails ownership for a genuinely retired sender and never actually clears the hint; both now match the sanctioned cross-workspace-ack override already used elsewhere in <code>devswarm-parent-gate.js</code>.</li> <li><strong>Fixed: <code>meshRowCopy</code> no longer stamps an explicit <code>undefined</code>-valued key for a field the source row never had</strong> — a verbatim/forward copy of an old row (no <code>origHash</code>/<code>instanceNonce</code>) now stays byte-identical to the source's own key set instead of gaining phantom <code>origHash: undefined</code>/ <code>instanceNonce: undefined</code> own-properties.</li> <li><strong>Known:</strong> per-row cursor shared across instances (8b211241bbe9), parent- inbox hook latency under load (bf965e5729c5), Stop-gate escalation on dead rows (9aaaf2c5e7b0), <code>diagnose</code>/roster's <code>orphans[]</code> surface still lacks the app-archived-marker recognition <code>staleRegistryPartitions[]</code> gained this release (a9ac2fc7e368 scope note above) — needs its own, separate fix; a0b7dfba1803/b48016f88bc1/af2b580f1b7b/084dee6b6e20 still need field captures before a fix — carried to the next release. Also carried: a 2-hop retired redirect refuses as <code>unregistered-workspace</code> naming the second hop rather than a dedicated dead-end reason; <code>count</code>/<code>read</code>/<code>ack</code> on an unregistered id still report a probe <code>storePath</code> instead of null; the retired-redirect ack ownership gate checks only the freshest live row on the survivor's worktree, so a legitimate caller on a twin-split worktree may see <code>retired-redirect-unresolvable-caller</code> (fail-closed; use <code>--ack-as-owner</code> as the hint says); <code>roster</code>/<code>diagnose</code>/<code>healthcheck</code> fail open on an unreadable registry (<code>healthcheck</code> reports <code>ok:true</code> on <code>EACCES</code>); when only <code>cursors.ndjson</code> is unreadable, <code>messages</code> reports <code>known:true</code> while <code>peek-primary</code> reports <code>EACCES</code> and <code>inbox read</code> returns the full backlog with <code>storeUnavailable:true</code> — the three verbs disagree about one store.</li> </ul> <h2 id="0971-2026-09-06">0.97.1 (2026-09-06)<a class="headerlink" href="#0971-2026-09-06" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: a store-unavailable inbox count no longer stops the drain loop or skips the heartbeat</strong> — <code>inbox count</code> / <code>inbox tick</code> report <code>known: false</code> when the mesh store could not be read, but still carry a numeric <code>unreadTotal</code> (often 0, the NDJSON side alone). The mailbox-wake stop condition treated that as "nothing to do", so a Primary or child whose store was briefly unreadable stopped draining and the child gate skipped the heartbeat. The injected stop condition now also requires <code>known</code> is not <code>false</code> (an absent field still counts as known, so older <code>count</code> shapes are unaffected), <code>inbox tick</code> records <code>known</code> in the wake-tick marker, and the child gate refuses to treat a <code>known: false</code> zero as a genuine no-op. Reported by the downstream fleet (c37ff1269685).</li> </ul> <h2 id="0970-2026-09-06">0.97.0 (2026-09-06)<a class="headerlink" href="#0970-2026-09-06" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: a reused-id archived marker no longer marks a live row archived</strong> — an archived marker left by a previous occupant of a reused workspace id (a different session id than the live descriptor) was outranking the live descriptor, so diagnose and the roster reported the row as archived (<code>live:false</code>/<code>archivedInApp:true</code>) even though its session was live; this affected a Primary's own anchor since 0.96.0 (reported on 0.96.1). Diagnose and roster now follow the live descriptor, and <code>doctor</code> lists superseded markers (report-only).</li> <li><strong>Changed: mailbox-wake defaults tuned for lower idle cost</strong> — the cron fallback now fires every 30 minutes instead of every 5 (<code>Monitor</code> remains the primary wake path; cron is never disarmed; <code>ANTIHALL_DEVSWARM_WAKE_CRON</code> is unchanged for anyone who has already overridden it). Each cron tick now runs a single <code>inbox tick</code> command that writes one wake-tick marker, so an empty tick costs one command and one line with no forced heartbeat. <code>doctor</code> now reports a cron-found-mail counter (ticks that found unread mail while a watcher lock was live), making the fallback's actual value measurable.</li> </ul> <h2 id="0962-2026-09-05">0.96.2 (2026-09-05)<a class="headerlink" href="#0962-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: app-side archive detection is now repo-scoped</strong> — hivecontrol's workspace listing is global, so the supervisor attributes each record to its own repo via its worktree path (same-repository siblings whose worktree cannot be resolved are kept, foreign or unattributable records are dropped and logged as <code>active-scope-drop</code> with per-tick <code>activeScope</code> counts) and stores only matching records per repo; records carry <code>repositoryId</code>/<code>label</code>/<code>branch</code> so the cross-repo archive guard can fire; probe and reconcile failures now log <code>error</code>/<code>status</code>/<code>signal</code>/<code>stderr</code> (200 chars) (<code>3cb559cb48d5</code>).</li> <li>Also from 0.96.1 but missing from its notes: defensive explicit <code>existsSync</code> guard when choosing a repo's representative worktree (<code>24a1ff3</code>).</li> <li><strong>Known:</strong> an archived-probe failure observed twice on one machine remains unreproduced (now instrumented; <code>3e000e49fe1b</code>); ack from an unresolvable caller still fails open; leaked fixture stores remain report-only.</li> </ul> <h2 id="0961-2026-09-05">0.96.1 (2026-09-05)<a class="headerlink" href="#0961-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: the supervisor no longer escalates a row whose harness session is alive</strong> — the idle-alive suppressor now consults the session pid directly, closing the 15-30 minute window where the dormancy gate (30 min) never reached the pid check while the stale gate (15 min) fired first (<code>90711586ecb1</code>). Rows already stuck in the terminal escalated state self-heal on the next supervisor pass when the session pid is alive (<code>recovery.log</code> reason <code>session-alive</code>); sticky escalated verdicts now carry measured <code>pending</code>/<code>notDraining</code>/<code>oldestUnreadAgeMs</code> instead of hardcoded <code>false</code>. <code>doctor</code> reports "escalated while session alive".</li> <li><strong>Added: the supervisor re-runs deferred post-update stages one per pass</strong> — <code>fold-all-stores</code>, <code>heal-orphan-partitions</code>, and <code>fold-archived-rows</code> are now driven by their resume markers, budgeted by <code>ANTIHALL_SUPERVISOR_SWEEP_BUDGET_MS</code> (default 20s), with a rotation cursor in <code>deferred-sweep-state.json</code>; the supervisor JSON line gains <code>deferredSweep</code> (closes the 0.96.0 Known item).</li> <li><strong>Docs:</strong> KB §26 gap sentence corrected.</li> <li><strong>Known:</strong> app-side archive detection may still lack a cache entry for a repo whose active probe fails (<code>3e000e49fe1b</code>, under investigation); ack from an unresolvable caller still fails open; leaked fixture stores remain report-only.</li> </ul> <h2 id="0960-2026-09-05">0.96.0 (2026-09-05)<a class="headerlink" href="#0960-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: routing sites select targets by heartbeat-aware liveness, not a bare non-empty session id (<code>f56dcc08f048</code>, routing half).</strong> <code>resolveMeshTarget</code> and <code>pickSurvivor</code> now require a row to actually be alive (heartbeat-checked) before treating its session id as a valid target; <code>rehomeMiskeyedRow</code> is intentionally unchanged.</li> <li><strong>Fixed: <code>callerOwnsRow</code> grants ownership of a lone same-worktree row only when it is unclaimed.</strong> A same-worktree row already claimed by another session is no longer silently treated as owned by the caller.</li> <li><strong>Fixed: watermark write and read+unlink now serialize under one lock</strong>, closing a race where a concurrent read and write could interleave.</li> <li><strong>Added: <code>send</code>/<code>heartbeat</code> results carry additive <code>identity:{id,kind}</code></strong> so callers can tell which kind of identity answered without guessing from shape.</li> <li><strong>Fixed: post-update sweeps honour deadlines mid-sweep</strong> (<code>foldMeshDuplicates</code>, <code>healOrphanPartitions</code>); fold-archived stages now get per-pair deadlines with resume markers, and <code>ANTIHALL_UPDATE_POSTPULL_BUDGET_MS</code> (default 90s) caps the whole devswarm stage sequence, deferring entire stages when the budget runs out (<code>e7307778b614</code>).</li> <li><strong>Fixed: reconcile requires a resolvable git root before per-row work</strong>, counted against the same budget; submodule/broken paths are now skipped and reported as <code>skippedNotGitRoot</code> (<code>6ef55fd42cc9</code>).</li> <li><strong>Fixed: <code>inbox ack</code> checks ownership before either cursor namespace moves</strong> and refuses a resolvable ownership mismatch; <code>read</code>'s refusal message now names the verb and points to <code>peek</code> (<code>66c7c4e9973e</code>).</li> <li><strong>Fixed: <code>diagnose</code> adds <code>archivedInApp</code></strong> and reports app-archived rows as not live (<code>07e01aee4f1f</code>); app-archive matching now requires <code>repositoryId</code> agreement and normalizes paths before comparison.</li> <li><strong>Docs:</strong> KB §36 ownership line added.</li> <li><strong>Known:</strong> the three deferred post-update stages (fold-all-stores, heal-orphan-partitions, fold-archived-rows) are not yet re-run by the supervisor when a budget defers them — they only run on the next update (tracked as a follow-up); <code>ack</code> from an unresolvable caller still fails open; leaked fixture stores remain report-only (see KB §38).</li> </ul> <h2 id="0950-2026-09-05">0.95.0 (2026-09-05)<a class="headerlink" href="#0950-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: <code>diagnose</code> resolves a row's sessionId through the descriptor when the registry is stale (<code>2c4ae6576fab</code>).</strong> When the registry row is absent or still carries the <code>unclaimed:</code> marker, <code>diagnose</code> falls back to the descriptor's session id and reports <code>descriptorSessionId</code> on disagreement, so a stale registry no longer silently misattributes a row.</li> <li><strong>Fixed: <code>unclaimed:</code> promotion derives the caller's session id from the harness process tree (<code>54a6539e2d69</code>).</strong> Promotion first tries <code>--session</code>, then the <code>CLAUDE_CODE_SESSION_ID</code> env var, and — only for rows still carrying the <code>unclaimed:</code> marker — walks the parent-pid chain to find the harness session file, verified with a cwd-in-worktree realpath check and a pid-liveness/pid-reuse guard. The fallback fails closed: a stale or reused pid rejects promotion rather than guessing a session id.</li> <li><strong>Fixed: descriptor/registry marker divergence is repaired both ways</strong> for the caller's own row once a session id is confirmed by either path above.</li> <li><strong>Fixed: registry write failures during promotion are surfaced, not swallowed.</strong> Inbox pull/read output now reports <code>promotion.registryWriteError</code> plus one stderr line on a failed write, and the promotion is retried on the next read instead of being lost silently.</li> <li><strong>Docs:</strong> KB §40 added; hooks KB documents the env-var launch-path fallback; SKILL docs updated to match.</li> <li><strong>Known:</strong> leaked fixture stores remain report-only (see KB §38 for manual cleanup); remaining open defects are tracked in the defect channel.</li> </ul> <h2 id="0941-2026-09-05">0.94.1 (2026-09-05)<a class="headerlink" href="#0941-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fixed: reconcile budget tests no longer race the wall clock</strong> — <code>cmdReconcile</code> samples an injectable clock (<code>ctx.reconcileNow</code>). v0.94.0's tag exists but was never published because both CI runs failed on that flaky test; v0.94.1 is the first published build of the 0.94 line.</li> </ul> <h2 id="0940-2026-09-05">0.94.0 (2026-09-05)<a class="headerlink" href="#0940-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: deterministic sender attribution for pending questions (<code>f3b8f326bfc3</code>).</strong> <code>companion/lib/devswarm-attribution.js</code> now picks a stable <code>pendingQuestions[].from</code> id across repeated summary passes instead of drifting with liveness churn: it prefers a real live-session row, then the row matching the worktree's branch slug, then falls back to the lexically smallest candidate id. Slug matching tolerates a trailing separator on either side, so <code>wave9-A</code> and <code>wave9-A-</code> resolve to the same attribution.</li> <li><strong>Fix: unbounded reconcile could run indefinitely on a large repo set (<code>f3c1bc827d89</code>).</strong> Reconcile now defaults to a 60-second wall-clock budget (<code>ANTIHALL_RECONCILE_BUDGET_MS</code>, <code>0</code> = unlimited); rows whose worktree is already missing are skipped before a subagent is spawned. Rows deferred by the budget are saved to <code>reconcile-resume.json</code> and are processed first on the next run. Output gains additive fields only — no existing field changes shape or meaning.</li> <li><strong>Improvement: <code>update.js</code> reports per-stage progress on stderr</strong> so a long update doesn't look hung; <code>ANTIHALL_UPDATE_QUIET=1</code> silences it. stdout output is unchanged, so scripts parsing it are unaffected.</li> <li><strong>Fix: <code>repoKey</code> git spawn had no timeout</strong> and could hang on an unresponsive filesystem/remote; it now times out after 10 seconds.</li> <li><strong>Feature: doctor §6l reports leaked test-fixture DevSwarm stores</strong> (report-only — no automatic cleanup). Surfaces fixture directories left behind by prior test runs so they can be cleaned up manually.</li> <li><strong>Tests:</strong> four tests fixed to isolate <code>HOME</code> so they no longer read or write the real user store; a new hygiene lint flags tests that spread or <code>Object.assign</code> a bare <code>process.env</code> into a child env instead of isolating it.</li> <li><strong>Docs:</strong> KB §38/§39 added; <code>RELEASING.md</code> step 4 now also runs <code>doctor.js --check</code> as part of the release gate.</li> <li><strong>Known:</strong> leaked fixture stores are currently report-only (see KB §38 for manual cleanup); <code>2c4ae6576fab</code> and <code>54a6539e2d69</code> remain open, targeted for 0.95.0.</li> </ul> <h2 id="0930-2026-09-05">0.93.0 (2026-09-05)<a class="headerlink" href="#0930-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Feature: app-side archive detection by absence, not by a field.</strong> Measured on hivecontrol 2.5.1 (maintainer machine): <code>workspace list all</code> exposes no archive field, so archive status can't be read directly. The supervisor sweep now caches the active set (<code>hivecontrol-active.json</code>) whenever the list call succeeds with at least one record; a registry row under the DevSwarm repos root that is absent from that cache by both id and worktree path, and older than the snapshot by a 10-minute grace, is treated as app-archived while the cache stays fresh (within 2x the reconcile cooldown). This is a liveness-axis signal only — a genuine unread question still gates regardless of archive status.</li> <li><strong>Fix: unanswered-question attribution never resolved to the recipient's own identity family (<code>f3b8f326bfc3</code>).</strong> The sender of a pending question is now resolved with the recipient and its cross-linked twins excluded first: the stored sender id wins, then a cross-linked row, then the shared freshest-live ranking. With no live sender-family row left, the raw sender id is kept rather than mis-attributed. A reply from the true sender's identity family now clears the question; a reply from the recipient itself or a twin does not. Forwarded copies keep <code>needsReply</code> set, since the copy is the only carrier once the source row retires. The per-turn notice and the Stop gate share the same identity-family-aware clearing path, including store-registry rows that carry no descriptor. <strong>Contract change:</strong> <code>pendingQuestions[].from</code> is now the true sender's identity-family id (or the raw sender id when no live family row exists) instead of the freshest live row on the worktree; consumers of <code>summary.json</code> that keyed on the old value should expect different ids for existing stores.</li> <li><strong>Fix: union counts migrate covered legacy lines exactly once.</strong> A new body-multiset tier in the union path covers NDJSON lines with no <code>_h</code> hash whose legacy hash is also absent, reusing the migration helpers already shared with <code>devswarm-unread.js</code>.</li> <li><strong>Docs:</strong> KB re-test recipe added for <code>0a668d81c0c6</code>; the identity-family attribution contract is now documented.</li> <li><strong>Known:</strong> a question from a retired sender whose identity family has no live row is listed under the raw sender id and cannot be cleared by a reply yet (under review). A watermark read-then-unlink residual window remains. The tail-cap refusal stays in place as defense in depth. An active-list snapshot smaller than half the previous one for a repo is refused (partial-list guard); a question whose sender has no registry row of any kind (removed) is informational and must be acked after inspection; archived-but-live senders still block.</li> </ul> <h2 id="0920-2026-09-05">0.92.0 (2026-09-05)<a class="headerlink" href="#0920-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Feature: session-sourced row liveness.</strong> anti-hall now classifies a DevSwarm row as <code>active</code>, <code>idle-alive</code>, or <code>dormant</code> instead of a single stale/alive split. An interactive Primary sitting at its prompt with a live harness session pid now reads <code>idle-alive</code> — it is never called dormant and never escalated (prompted by a field report; <code>699a236129c5</code>).</li> <li><strong>Fix: ghost same-path registry rows now age out.</strong> A row with a null session id and no descriptor, heartbeat, cursor, or ack, older than 72 hours (tunable via <code>ANTIHALL_DEVSWARM_GHOST_ROW_MAX_AGE_H</code>), now retires through the existing forward-then-tombstone fold — but only when exactly one non-ghost survivor exists for that path (<code>76891c157288</code>).</li> <li><strong>Fix: archived rows never alert.</strong> A row archived through anti-hall's own archive verb now carries an <code>archived</code> label and is excluded from the stale/escalated/dormant axis entirely. Limit: an archive performed only in the DevSwarm app writes no marker anti-hall can see yet, so that path is not detected (planned: a supervisor-cached hivecontrol archived list).</li> <li><strong>Fix: broadcast ownership recognizes identity families.</strong> <code>heartbeat --summary</code> now accepts same-worktree or cross-linked identity-family rows as owned instead of dropping them; a genuine mismatch still keeps <code>dropped</code>/<code>dropReason</code> (<code>ecd7ad60e4cc</code>).</li> <li><strong>Hardening (Round 15):</strong> same worktree alone no longer grants broadcast ownership — an identity link (cross-linked rows, the caller's real session id on the target, or a true placeholder) is required; <code>archived</code> never hides a <code>not-draining</code> backlog in the per-turn table; the watermark is deleted only when the ack covers its freshly re-read value; the supervisor sweep honors idle-alive; a reused pid (started after the session file) no longer reads alive.</li> <li><strong>Fix: live-to-dead sibling handoff no longer re-delivers the covered backlog.</strong> The watermark is now read unconditionally, the ack is anchored on the delivery frontier (not an unconditional read), and the watermark file is retired once a covering ack lands — closing the one-time re-delivery gap on that transition.</li> <li><strong>Hardening: watermark filenames.</strong> Ids containing <code>.seen-</code> are now refused outright; the filename parser is an exact inverse of the writer; doctor's sweep is existence-based instead of pattern-guessing.</li> <li><strong>Deferred:</strong> app-side archive detection (see limit above); a migrate-skipped-legacy-line double-count (P2, tracked not fixed); the tail-cap hard refusal is now unreachable in practice and kept only as defense in depth; <code>callerOwnsRow</code> clause 3 can promote a lone foreign <code>unclaimed:</code> row in the caller's worktree (capped by the promotion idempotence guard).</li> </ul> <h2 id="0910-2026-09-05">0.91.0 (2026-09-05)<a class="headerlink" href="#0910-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix (P0): recurring sibling re-delivery.</strong> <code>inbox read</code> (primary) and <code>inbox count</code> now size a sibling's delivery window from the max of its JSON ack file and its store cursor, instead of trusting either alone; fold and <code>reap-orphans</code> write both cursor namespaces in lockstep so the two never drift apart again. <code>reconcileOrphanCursor</code> no longer folds the NDJSON line cursor into the store-row minimum and never rewinds from an untrusted zero. A new caller-scoped seen-watermark stops perpetual re-delivery from a live (non-ackable) sibling without touching that sibling's own cursors. Fold-forward now skips rows the survivor's reader already consumed and stamps <code>origHash</code> on the forwarded copy, closing the same hole in the supervisor's periodic fold, <code>update.js</code>, and doctor-repair. A non-ackable sibling on an acking read now takes the structural (contiguous) prefix instead of an arbitrary cursor.</li> <li><strong>Feature: exact dedup of forwarded copies.</strong> An additive, nullable <code>orig_hash</code> column identifies a forwarded row's original message; legacy rows without the column are reconstructed for comparison. Weak-key duplicates are consumed (not left to re-accumulate) with a log line (<code>64861a623503</code>).</li> <li><strong>Fix: unread-count unification.</strong> <code>computeSummary</code> now uses the shared loss-free union instead of a second counting path; legacy-line hashing is unified, fixing a double-count of legacy lines (<code>8f2aec40e2ff</code>).</li> <li><strong>Feature: <code>unclaimed:</code> session-id promotion.</strong> A row still tagged <code>unclaimed:<id></code> is promoted to the real session id on read/pull, but only ever for the caller's own row, never the row being read. Ships with a forward migration in both <code>update.js</code> and doctor-repair (idempotent, fail-open, no-delete).</li> <li><strong>Feature: send receipts.</strong> <code>cmdSend</code> writes a <code>repoKey</code>-scoped receipt for each send; the reply tracker credits from a matching receipt before falling back to stdout parsing.</li> <li><strong>Clarified: <code>inbox read</code> is read-only by design.</strong> It now reports <code>cursorAdvanced: false</code> plus an <code>ackHint</code> naming the exact <code>inbox ack <id></code> command when something is outstanding; a plain read never advanced a cursor (<code>56ba248504d0</code>).</li> <li><strong>Feature: doctor repair mode</strong> now runs the drain-marker, reaped-log, receipt, and <code>unclaimed:</code> promotion sweeps in one pass. The authority-override log honors the configured home, and a cross-project override now warns about the orphan it leaves behind.</li> <li><strong>Fix:</strong> <code>computeSummary</code> skips the union entirely when the NDJSON tail has no unread, avoiding needless work on the common case.</li> </ul> <p>Deferred to v0.92.0: a session-sourced liveness axis (<code>699a236129c5</code>), ghost same-path row ageing (<code>76891c157288</code>), archived-workspace rows must not alert, heartbeat broadcast ownership by identity family (<code>ecd7ad60e4cc</code>), and migrate-skipped legacy lines still double-counted (P2).</p> <p><strong>Operator note:</strong> after updating, the first read-primary pass may deliver a sibling backlog once as cursors reconcile — this is expected, not a regression. Use <code>inbox messages <id> --tail N</code> to inspect a channel without moving any cursor.</p> <h2 id="0900-2026-09-05">0.90.0 (2026-09-05)<a class="headerlink" href="#0900-2026-09-05" title="Permanent link">¶</a></h2> <ul> <li><strong>Feature: <code>inbox messages</code> gains <code>--since <index|date></code> and <code>--tail N</code></strong> as a post-delivery projection window. Every ack-bearing verb refuses the window flags outright, and <code>--tail</code> under truncation refuses with <code>tail-under-truncation</code> (pointing callers at <code>--since</code>) instead of silently returning the last N rows of the oldest surviving batch as if they were the newest.</li> <li><strong>Fix: <code>cmdSend</code> now reports the UTF-8 byte length</strong> of the sent payload in its receipt, instead of a character count that undercounts multi-byte text.</li> <li><strong>Fix: <code>cmdSpawn</code> reports <code>launched: true | 'unknown'</code></strong> from positive post-spawn evidence within a bounded, clamped poll, rather than conflating "created" with "launched." A prior occupant's still-fresh heartbeat on a reused worktree is no longer misread as the new spawn's own launch signal.</li> <li><strong>Fix: fold anchor guard now cross-references liveness.</strong> The liveness half of the guard uses heartbeat- and reader-evidence together (<code>isSiblingPartitionLive</code>); a cross-referenced anchor keeps both heartbeat and reader-evidence protection so a live anchor row is never folded away as a candidate.</li> <li><strong>Fix: a same-worktree UUID twin is now classified as the Primary's own identity family</strong>, not a neglected child, closing a false-positive escalation path in the parent gate.</li> <li><strong>Feature: in-flight drain marker.</strong> A writer marks an ack-bearing read at entry and clears it in a <code>finally</code>; the parent-gate consumer now treats a downgrade from an active marker as non-escalation-spending, so a crashed drain can no longer silence the gate for a TTL. A doctor sweep removes stale (TTL-expired) markers.</li> <li><strong>Feature: <code>reap-orphans</code></strong> (dry-run by default; <code>--apply --max N</code> required for a live pass) archives a DevSwarm partition's unread rows to <code>reaped/<id>.ndjson</code> and verifies the archive before advancing its cursor — no message row is ever deleted. Gated to human-only invocation (refuses under automation env or non-TTY stdin without an explicit human flag).</li> <li><strong>Feature: <code>--force-cross-project <id></code></strong> on <code>archive</code>, accepted only when the value matches the target id exactly, audited to a dedicated log file. Deliberately archive-only — it does not reopen ack/gate cross-project data movement.</li> <li><strong>Feature: <code>reconcile-registry</code></strong>, a report-only verb that surfaces registry/hivecontrol drift in both directions plus <code>worktreePath</code> mismatches, failing soft (not crashing) on an unrecognized hivecontrol JSON shape.</li> <li><strong>Feature: <code>heartbeat --summary</code></strong> now reports <code>dropped</code>/<code>dropReason</code> when an ownership refusal discards a summary, instead of silently reporting <code>ok: true</code> with no indication anything was dropped.</li> <li><strong>Feature: a new PreToolUse hook additively throttles repeat-prone, repo-wide scan commands</strong> (e.g. a knowledge-graph update) by rewriting them to run at background priority (<code>taskpolicy -c utility nice -n 19</code> on macOS, <code>nice</code>/<code>ionice</code> on Linux) via <code>hookSpecificOutput.updatedInput</code>. Prefix-only and idempotent, extendable via an environment variable allowlist, with a kill switch, and it fails open when the platform's throttling tool isn't available.</li> <li><strong>Feature: a new SessionEnd hook sweeps orphaned MCP server processes left by prior crashes.</strong> Claude Code shuts down its own MCP children on a normal SessionEnd, and SessionEnd never runs on SIGKILL, so any leaked MCP process it finds is necessarily left over from an earlier crash. The sweep only matches processes already reparented to PID 1 (a genuine init), requires a minimum process age, caps how many it will touch per run, and is on by default with a kill switch; it skips processes under a launchd/systemd supervisor. (Corrects an earlier internal note: the SessionEnd payload's field is <code>reason</code>, not <code>end_reason</code>.)</li> <li><strong>Fix: a latent bug</strong> where <code>heartbeatTs</code> was never imported into the DevSwarm CLI module, left over from an earlier refactor.</li> </ul> <p>Deferred to v0.91.0: unread-count unification, <code>unclaimed:</code> session-id promotion + migration, delivery/send receipts, exact dedup of archived-then-forwarded message copies, and a retention sweep for the <code>reaped/</code> archive directory.</p> <h2 id="0890-2026-09-04">0.89.0 (2026-09-04)<a class="headerlink" href="#0890-2026-09-04" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix (P0): the parent-gate reply tracker silently dropped credited replies.</strong> <code>devswarm-parent-reply-tracker.js</code> parsed a Bash tool's ENTIRE stdout as one JSON value; a compound command (heredoc + <code>grep -c</code> + a <code>devswarm.js send</code>) printed a <code>0</code> line before the send's JSON, the parse threw, and the reply vanished with no log — defeating the 0.88.0 identity-family fix regardless of cause. <code>parseSendResponse</code> now scans stdout line-by-line and keeps the LAST line that parses to a JSON object and is itself send-shaped, tolerating a trailing JSON line from a chained command; a total parse failure now writes one bounded NDJSON diagnostic (<code>event: 'reply-parse-drop'</code>) to <code>~/.devswarm/parent-inbox.log</code>, gated behind the arming check so unarmed sessions log nothing. <code>hooks/lib/devswarm-detect.js</code> gained <code>hasOnDiskDevswarmState()</code> as a second arming path alongside the env-var gate.</li> <li><strong>Fix (P0): <code>foldOne</code> could fold a live Primary's own mesh row as a fold <em>candidate</em>, eating its mail every child turn.</strong> <code>devswarm.js</code>'s <code>foldGroupIntoSurvivor</code> protected the anchor only when it was the FOLD CALLER; a co-located child's self-register (<code>hooks/devswarm-child-turn.js</code> → <code>retireWorktreeDuplicates</code>) could still fold the anchor in as a candidate and advance its cursor. The guard now protects only an ATTENDED anchor (a live session id, or an on-disk descriptor) rather than blocking on bare identity, so unattended archive fixtures still retire correctly while a live anchor is protected — restoring both the field regression fix and all 14 previously-failing <code>devswarm-archive-group.test.js</code> cases. "Attended" is decided by three independent signals: a live session id, an on-disk descriptor, or READER EVIDENCE (store cursor > 0, or a primary cursor file exists) — the last one added after a field report showed a live Primary row carrying a synthetic <code>unclaimed:</code> session id, so its only protection had been the descriptor, which <code>archive</code> deletes without a liveness check. Rows the guard leaves are now reported as <code>mesh-anchor-attended</code> by the archive/fold sweeps instead of the misleading <code>raced-re-register</code>. Regression tests use the exact field row shape.</li> <li><strong>Change (owner-mandated): parent↔workspace <code>SendMessage</code> must go through the mesh, not <code>ListAgents</code>.</strong> New <code>hooks/devswarm-comms-guard.js</code> on <code>PreToolUse</code>/<code>SendMessage</code> resolves a target's cwd via <code>~/.claude/sessions/<pid>.json</code> and blocks any target whose cwd sits under a DevSwarm workspace root, closing a path where a workspace-backing session was addressable directly as a peer. Allows <code>main</code>, agentId-form targets, and any unresolved name (fails open); a name-match allow backstop was removed as a loophole once regression-tested. Active only while DevSwarm is detected as running.</li> <li><strong>Fix: three guard false positives.</strong> <code>command-guard.js</code> had no heredoc awareness, so heredoc body lines starting with a heavy verb (e.g. "make sure…") were matched as commands; it also matched heavy patterns against grep/sed/awk <em>search-pattern</em> operands instead of only the command shape. <code>edit-guard.js</code> had no exemption for the harness-assigned per-session scratchpad and, separately, compared paths without resolving <code>/tmp</code> vs <code>/private/tmp</code> on macOS, so the harness's own scratchpad path could fail closed.</li> <li><strong>Fix: union-read's never-read sibling cap could pin a non-ackable partition to its oldest 200 messages forever.</strong> The cap gated on the partition's own cursor being zero, which never advances for a partition the caller cannot ack. It now gates on backlog size instead, keeps the newest capped rows for a non-ackable partition, hard-refuses any ack against a capped read, and surfaces a <code>neverReadCapHint</code> pointing at <code>inbox messages <pid></code>.</li> <li><strong>Fix: the twin-aware sibling ack gate stranded a caller's own twin's cursor</strong>, causing perpetual re-delivery of messages already folded into the caller's own reads. The SELF-twin branch now advances the twin's cursor using the same ack-target arithmetic as the caller's own ack.</li> <li><strong>Fix: an escalated verdict could not be cleared for a DONE child</strong>, and the supervisor reported a fresh spawn as "idle 0m" before its first heartbeat. <code>devswarm-parent-gate.js</code> now reads <code>archive_ready</code> off the derived summary; <code>devswarm-supervisor.js</code> gives a 2-minute grace on descriptor mtime and excludes already-done children from the idle check.</li> <li><strong>Fix: <code>defect-store.js</code> could lose a record on write overflow</strong> instead of failing closed; overflow now spills as a distinct <code>t:'overflow'</code> line that every reader filters on explicitly.</li> <li><strong>Fix (doctor §6k): flag MCP server children orphaned under a live broker.</strong> The existing PPID==1 reaper cannot see children reparented under a still-running app-server broker rather than PID 1 — correct conservatism, but it leaves a real leak invisible. Doctor gained a read-only, warn-only check (<code>checkOrphanedMcpUnderBroker</code>) that never gates pass/fail and stays silent when <code>ps</code> is unavailable. Not anti-hall's own leak; reported so it can be tracked upstream.</li> <li><strong>Fix: the gone-worktree gate hint recommended a destructive override on a weak signal.</strong> <code>devswarm-parent-gate.js</code>'s hint suggested <code>inbox ack --ack-as-owner</code> whenever a worktree read as gone via a bare ENOENT stat — a MOVED worktree reads identically to a retired one, and <code>--ack-as-owner</code> has no liveness check on its target. The hint now says to inspect with <code>inbox read <id></code> first and warns explicitly that a moved worktree can read as gone before recommending the override.</li> <li><strong>Fix: the reply tracker could pick a trailing non-send JSON line</strong> (e.g. a chained <code>inbox count</code>) over the actual send response; it now requires the send's own shape. The arming-gate check was also reordered ahead of parse/diagnostic logging so an unarmed session never writes a drop-log line for stdout it wasn't tracking.</li> <li><strong>Perf: <code>canonicalMeshId</code> is now memoized per fold pass</strong> instead of spawning <code>git</code> once per fold candidate.</li> <li><strong>Docs:</strong> added <code>docs/KB-claude-code-hooks.md</code> (hook event reference, including the measured 10k-char injection cap and <code>FileChanged</code> behavior).</li> <li><strong>Field report:</strong> the P0 reply-tracker and P0 fold-anchor fixes above were both driven by an external field report of live mail loss in a DevSwarm-coordinated session; both are reproduced locally and covered by regression tests in this release.</li> </ul> <p><strong>Known / deferred to v0.90.0:</strong> - The comms guard's liveness predicate uses <code>isLiveSessionId</code> rather than <code>isSiblingPartitionLive</code>; these can diverge for a partition that is live but not the calling session's direct sibling. - <code>isStaleCrossReference</code>'s bypass for anchor rows needs a narrower definition than "any anchor." - Phantom-rescue can self-block under specific same-worktree spawn timing. - No in-flight drain marker exists between "message accepted" and "message durably persisted," leaving a narrow window unaddressed. - <code>cmdSpawn</code> can leave a row registered-but-never-launched past its normal cleanup path in one ordering; the 6-hour dead-classification deadline (shipped in 0.88.0) bounds but does not eliminate this.</p> <h2 id="0880-2026-09-03">0.88.0 (2026-09-03)<a class="headerlink" href="#0880-2026-09-03" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: <code>inbox ack</code> and <code>read-primary</code> sibling acks could each independently consume a live sibling child's mail on stale-looking liveness signals.</strong> Both surfaces now route through a single shared <code>siblingAckGate</code> that requires positive evidence of death before an ack is allowed to drain a sibling's messages, instead of two separately-maintained gates that could diverge. <strong>This gate controls whether an ack is allowed to happen — it is not a display-only signal.</strong></li> <li><strong>Fix: <code>inbox ack</code> re-counted the live file tail at ack time (both the NDJSON and store code paths), racing a concurrently-arriving message.</strong> A message that landed between the initial read and the ack call could be acked-away undelivered. <code>inbox ack</code> now acks only the snapshot captured at the initial read.</li> <li><strong>Fix: a registry row that registered but never launched could resurface its backlog forever.</strong> Such a row is now classified dead after a 6-hour deadline, allowing its partition to be drained instead of blocking indefinitely. Known limitation: the deadline reads a descriptor file mtime that several routine operations legitimately rewrite (<code>ensure</code> on every <code>inbox pull</code>, rehome/heal, archive/unarchive round-trips, <code>migrateOwnerKeys</code> via doctor-repair or the updater) — a repair, update migration, or archive round-trip can extend the 6h window. The fail direction is toward protection (no mail loss), never toward premature drain.</li> <li><strong>Change: child wake/drain instructions now use the cursor-advancing <code>read-primary</code> instead of the non-mutating <code>inbox read</code>.</strong> <code>inbox read</code> could never clear a withheld gap on its own.</li> <li><strong>Fix: the parent gate could report an already-answered question as unanswered.</strong> Replies are now matched across an agent's full identity family instead of by raw string equality against a single id.</li> <li><strong>Test infra: mutation tests now run against a scratch copy of the plugin tree.</strong> Previously a green <code>node --test</code> run could leave an injected mutant sitting in the working tree's real source.</li> </ul> <h2 id="0871-2026-08-29">0.87.1 (2026-08-29)<a class="headerlink" href="#0871-2026-08-29" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: the parent gate never consulted the mesh store on a live worktree's mid-teardown inbox ENOENT, treating it as neglect it could not clear.</strong> <code>devswarm-parent-gate.js</code>'s un-clearable-axis rule only recognized a dead descriptor when <code>worktreeIsGone()</code> fired on a definitive ENOENT stat of the worktree path itself. A worktree that still existed on disk but whose native <code>inbox.ndjson</code> was absent (mid-teardown, or before the first message) fell straight to <code>unreadUnknown = true</code> and never asked the store — the actual source of delivery truth (<code>storeSeq</code>) — even when the Primary's own messages had already been delivered there. Added a new branch, <code>reason === 'inbox-missing' && !foreignProject</code>, that opens the store and requires REAL, non-empty message history (<code>cursorValue</code> + <code>listMessages().length > 0</code>) before clearing the unknown axis; a bare store open is not sufficient evidence, since <code>openStore</code> auto-creates the partition on first touch and the real per-turn registration path already opens it for unrelated bookkeeping. EACCES and every other non-ENOENT reason, plus foreign-project descriptors, still fall through and fail closed exactly as before. Ships alongside 0.87.0's verdict-corroboration gate in the same file; both fixes coexist (110/110 tests in <code>devswarm-parent-gate.test.js</code> pass).</li> </ul> <h2 id="0870-2026-08-29">0.87.0 (2026-08-29)<a class="headerlink" href="#0870-2026-08-29" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a stale/escalated liveness verdict alone could hard-block the Primary indefinitely, even after a workspace had genuinely finished.</strong> <code>readVerdictStatus()</code> in <code>devswarm-parent-gate.js</code> discarded the verdict file's own <code>pending</code>/<code>notDraining</code> flags and returned only the bare <code>status</code> string. Because <code>escalated</code> is STICKY (<code>liveness.js</code>'s terminal short-circuit returns it unchanged until a fresh heartbeat a finished session will never emit again), a persisted verdict of <code>{"status":"escalated","pending":false,"notDraining":false}</code> — the verdict itself saying nothing was outstanding — force-blocked the Primary on ~20 consecutive turns. <code>readVerdict()</code> now threads <code>pending</code>/<code>notDraining</code> through to <code>main()</code>, and a bare <code>stale</code>/<code>escalated</code> status can no longer drive a hard block by itself: it now needs corroboration from at least one of four independent axes (the verdict's own <code>pending</code> flag, a real union-unread backlog, an unreadable unread axis — fail-open toward blocking, never toward silence — or an unanswered question from that family). An uncorroborated status degrades to a one-time stderr advisory instead of a hard block; a family with any other real signal still blocks normally. See <code>docs/KB-devswarm-hivecontrol.md</code> for the corroboration invariant this generalizes: <strong>a bare verdict label is not evidence.</strong></li> <li><strong>Fix: <code>liveness.js</code>'s own union-unread signal double-counted a caller's own outbound message as evidence the target was neglecting inbound work.</strong> A message the Primary itself just sent into a child's mailbox, still sitting unread pending the CHILD's own read, was being read as the child "not draining." <code>resolveSelfId()</code> resolves the caller's real Primary id (mirroring <code>recovery.js</code>'s addressee-hash fix — <code>primaryWorkspaceId()</code> is a pure hash of the path handed to it, and a linked worktree's own root hashes to that worktree's id, not the real Primary's, unless resolved through <code>resolveMainWorktree()</code> first). A new <code>pendingInbound</code> value excludes store-only rows sent by that resolved self id from the staleness gate, while the pre-existing <code>pending</code> value is unchanged and still reports full mailbox depth for drain/ack accounting elsewhere.</li> <li><strong>Fix: the DevSwarm wake instruction unconditionally told an agent to run the full mailbox drain+read sequence on every wake turn, which the agent's own cron prompt routinely delegated to a subagent even when the mailbox was empty.</strong> <code>drainCmd()</code> in <code>devswarm-wake.js</code> now runs the cheap, inline, non-mutating <code>inbox count</code> first and only pays for a drain/read (optionally delegated) when <code>unreadTotal > 0</code>. The child branch still pulls its native queue unconditionally (cheap and the only way a native-queue-only backlog becomes visible to a later count); only the read step is gated on the count.</li> <li><strong>Fix: <code>ackTo()</code> (the durable-inbox cursor primitive) was callable unlocked from multiple sites, so two overlapping drains could race a cursor backward</strong> — a slow writer's lower <code>ackTo()</code> landing after a fast writer's higher one regressed the cursor, causing re-delivery. <code>ackTo()</code> is now monotonic by default (raises the write target to at least the current on-disk cursor). <code>reconcileOrphanCursor()</code> in <code>scripts/devswarm.js</code> is the one proven legitimate exception — a MIN-only reconciliation across three cursor namespaces that must be able to lower a namespace stuck above the others — and opts in explicitly via <code>{ allowRewind: true }</code>.</li> <li><strong>Fix: a fold pass never advanced a folded-away candidate's own cursor</strong>, even after every one of its unread rows had fully forwarded to the survivor, so an already-forwarded backlog kept rendering as "N unread / not draining" on a <code>left</code> (never-tombstoned) candidate indefinitely. The cursor now advances once the fold's forward loop completes with no exception; a partial/failed forward still leaves the cursor untouched so the next pass safely re-forwards idempotently (hash dedupe) instead of silently dropping rows off the read frontier. The supervisor's reconcile sweep is unaffected by the same stale-cursor condition as a result.</li> <li><strong>Fix: the orphan-entry classifier treated a forwarded entry as orphaned without checking whether it had also been drained by the consumer.</strong> A forwarded entry that is later drained is not an orphan; the classifier and the store's forwarded/drained bookkeeping now agree on entry state.</li> <li><strong>Docs:</strong> every hard-coded remediation/usage string that told an agent to run <code>send --to <id> --message "..."</code> (the command-guard block reason, the child role/turn per-turn reminders, the parent-inbox unread/urgent/ unanswered nudges, and the gate's unanswered-question segment) now points at <code>--message-file <path></code> (or <code>--message-stdin</code>) instead — a shell-quoted <code>--message</code> body with embedded newlines/quotes, exactly the shape a structured question/reply tends to have, is prone to shell-quoting mangling. Also documents the <code>seq</code> (durable, store-wide, comparable across calls) vs <code>index</code> (page-local positional ordinal, the unit <code>--to N</code>/the ack cursor advance in) distinction on every returned message row, in both <code>SKILL.md</code> docs, to prevent a future caller from comparing <code>index</code> values across separate calls.</li> </ul> <h2 id="0860-2026-08-28">0.86.0 (2026-08-28)<a class="headerlink" href="#0860-2026-08-28" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: the ingest/supervisor daemon units emitted a <code>PATH</code> that could not resolve <code>node</code>, so every <code>hivecontrol</code> grandchild died exit 127 and reconciliation silently healed nothing.</strong> The installers bake an ABSOLUTE node path as the unit's interpreter (<code>process.execPath</code> at install time — commonly a version-manager directory such as <code>~/.nvm/versions/node/vX/bin</code>, which is on no scheduler's default <code>PATH</code>), but built the unit's <code>PATH</code> from the resolved <code>hivecontrol</code> directory plus a minimal fallback only. The node bin dir was in neither. The daemon itself therefore always started and passed every "is it running" check (absolute <code>argv[0]</code>); the failure lived one process lower, because <code>hivecontrol</code> is a SCRIPT whose shebang re-resolves <code>node</code> THROUGH <code>PATH</code>. Measured on a live install before the fix: <strong>23,928 <code>env: node: No such file or directory</code> failures across 1,757 supervisor sweeps spanning three repoKeys, with <code>healed:0</code> on every single sweep</strong> — reconciliation had never once succeeded, for any scope, for as long as the units had existed. Fixed at the single chokepoint all six plist/service/cron emitters across both installers derive from (<code>install-devswarm-supervisor.js</code> imports this very function): <code>unitEnvFor</code> now prepends <code>dirname(execPath)</code> and takes <code>execPath</code> as a REQUIRED argument, with each emitter passing the very <code>exec</code> it writes — so a unit's <code>PATH</code> structurally cannot disagree with the interpreter baked into that same unit. This completes the v0.65.0/v0.66.0 <code>hivecontrol</code>-path fixes, which addressed finding the CLI but not running it. See <code>docs/KB-devswarm-hivecontrol.md</code> §31 for the generalized invariant.</li> <li><strong>Also closed in the same pass:</strong> the unit environment used to be suppressed ENTIRELY when <code>hivecontrol</code> could not be resolved, coupling two independent facts — a <code>PATH</code> that resolves <code>node</code> is worth emitting even when the CLI path cannot be pinned, so an environment is now omitted only when NEITHER input is usable. And install refuses outright if the node binary at <code>EXEC</code> is not a real file, rather than baking a permanently unstartable unit whose only symptom is a line in a scheduler log.</li> <li><strong>Fix: the ingest auto-heal was gated so it could never fire when it was needed.</strong> <code>runUpdate</code> attempted <code>healIngestDaemon</code> only when that run had synced new bytes into the version cache. But a daemon's baked script path goes stale with NO version bump — the plugin manager relocating or <code>.bak</code>-ing the version-pinned cache dir the unit was built from, which is precisely the case the heal function's own header describes. In that steady state the installed version already equals latest, <code>syncCache</code> no-ops, and <code>classifyIngestUnit</code> — the one thing that would notice the dangling <code>scriptPath</code> — was never reached. The heal now ALSO fires when the installed unit fails to classify <code>ok</code>, via the extracted <code>inspectInstalledIngest</code> so the heal DECISION and the heal ACTION read the same unit through the same lookup and cannot drift. On a no-sync run the added arm is read-only (a unit enumeration plus a few <code>statSync</code>s — no spawn, no writes) and fail-open. <code>absent</code> is deliberately NOT a trigger: first-installing an opt-in daemon is the update skill's own documented step, so treating it as "needs heal" would spawn an installer on every no-op update for every user who never enabled the daemon.</li> <li><strong>Fix: <code>devswarm-store-leak-report.js</code> advertised itself as read-only but wrote a file on every run.</strong> The script's own banner promises it "deletes, moves, renames, and truncates NOTHING", yet <code>--out</code> defaulted to a timestamped path under <code>.anti-hall/reports/</code>, so merely asking a question about the store deposited a JSON file in the tree. The write is now OPT-IN — no <code>--out</code>, no file — and the summary says plainly that nothing was written and how to ask for the artifact. Every <code>--out</code> safety property is unchanged and still applies whenever a path is named (realpath containment against the store root, the <code>.json</code> requirement, the <code>O_EXCL</code>/<code>O_NOFOLLOW</code> create, and the report-marker check before any overwrite). The audit/classification logic is untouched.</li> </ul> <h2 id="0850-2026-08-25">0.85.0 (2026-08-25)<a class="headerlink" href="#0850-2026-08-25" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: archiving a workspace now retires its whole identity family, so the Primary's Stop gate can no longer be blocked forever by an inbox that cannot exist.</strong> <code>archive</code> tombstoned by <code><id></code> only, but a descriptor's identity family can be cross-linked by <code>sessionId</code> instead (one row's <code>sessionId</code> IS the other row's <code>id</code>), so the twin stayed live in <code>workspaces/</code> after its sibling was archived. <code>devswarm-parent-gate.js</code> then nagged every turn about the missing inbox file of a workspace that by design could never produce one — un-clearable without editing state by hand. <code>cmdArchive</code> now retires the whole family at archive time, and <code>foldArchivedFamilyDescriptors</code> is a forward migration for the descriptor sets already split by the bug (wired into BOTH <code>skills/update/scripts/update.js</code> and <code>doctor</code>'s AUTO-SAFE <code>fold-archived-family-descriptors</code> repair, per this repo's persisted-shape rule). It is the descriptor-file counterpart of v0.70.0's <code>foldArchivedRegistryRows</code>, which only covered the registry half.</li> <li><strong>The parent gate distinguishes a dead descriptor from neglect.</strong> The rule used to be that <code>known:false</code> on the unread read ALWAYS blocked, unconditionally, including an absent inbox file. That absolute was the defect: an <code>inbox-missing</code> (ENOENT, and only ENOENT) on a descriptor whose <code>worktreePath</code> is ALSO provably gone from disk is not neglect, it is a dead descriptor, and no action the Primary can take would ever clear it. That one conjunction no longer raises the unknown axis. <strong>Nothing is hidden:</strong> a gone worktree with store-side unread still blocks on <code>unionUnread</code>, a stale/escalated verdict still blocks, and every other unreadable reason (<code>inbox-unreadable</code>/EACCES/EISDIR, <code>cursor-*</code>, <code>no-inbox-path</code>, <code>read-threw</code>) still blocks regardless of the worktree. "Gone" requires a definitive ENOENT <code>lstat</code> on an ABSOLUTE path — a missing/empty <code>worktreePath</code>, a relative path, a dangling symlink, or a stat failing for any other reason is NOT provably gone and therefore still blocks.</li> <li><strong>A descriptor is retired only against PROVEN write authority.</strong> Adversarial review of the first pass found three ways the retire path could delete a LIVE descriptor instead of the intended tombstone twin: a classification made before the lock and acted on after it, a race with the non-locking child-turn descriptor writer, and a reused id whose old tombstone was accepted as authority over a brand-new unrelated descriptor. All three are closed by proving authority before any write — a coherent inode+bytes generation fingerprint (<code>descriptorFileGeneration</code>/<code>sameDescriptorGeneration</code>) re-read INSIDE the per-id lock and compared against the scan-time snapshot, a per-id lock now taken by <code>devswarm-child-turn.js</code> around its own descriptor rename (bounded ~1s, fail-open, no nested acquisition), and <code>worktreeIsProvablyGone</code>'s fail-closed gate on the migration path. A generation mismatch or unproven gone-ness REFUSES the retire rather than guessing. Grouping uses the id/<code>sessionId</code> cross-link only, never bare worktree equality, so two legitimately-live tabs on one worktree are never retired.</li> <li><strong>Safety refusals are reported, not rendered as a clean no-op.</strong> A pass that declines to retire a twin (tombstone bytes differ, lock busy, descriptor changed since the scan, worktree still present) now surfaces those in <code>left[]</code> — through <code>update</code>'s summary line and through a doctor <code>notice</code> rendered on both the pending and the not-pending path — and a run that raised reports <code>ok:false</code>. Reporting only the retire count made every refusal look identical to "nothing to migrate".</li> <li><strong>Worktree paths are persisted absolute, and legacy relative paths fail closed.</strong> A relative <code>worktreePath</code> is only meaningful against the cwd it was registered from, which the descriptor does not record; resolving it from the Primary's cwd answers a different question. <code>scripts/devswarm.js</code> now writes absolute paths, and both readers treat a non-absolute path as "not provably gone" — i.e. keep blocking, keep the descriptor.</li> <li><strong>Fix: <code>npm test</code> no longer writes into the developer's real HOME.</strong> <code>tests/hooks/flutter-debug.test.js</code> spawned the REAL <code>hooks/doctor.js</code> with a wholesale <code>process.env</code> spread and no HOME override, and <code>doctor.js</code> defaults to repairs ON with <code>dryRun:false</code>. None of the 12 <code>migrationFix</code> passes in <code>hooks/lib/doctor-repair.js</code> sit inside the <code>gateOpen</code> block, so running the suite folded store DBs, rewrote registry rows, migrated owner keys, and touched recovery-intent markers against the real <code>~/.anti-hall/</code> and <code>~/.claude/settings.json</code>. CI was never affected (a fresh runner has no state), so this was a developer-machine-only hazard, and it predates this release. The test now runs doctor with <code>--check</code> plus an isolated disposable HOME, and <code>tests/helpers/spawn-hook.js</code>'s <code>testHook</code>/<code>testHookRaw</code> fall back to a per-process <code>mkdtemp</code> dir instead of the real machine home. The regression guard that should have caught it was a hardcoded 4-file allowlist that missed this fifth spawner entirely; it is now a scan over the whole <code>tests/</code> tree for anything that spawns <code>doctor.js</code> as a child process, with a self-check so a broken scan fails loudly instead of silently protecting nothing.</li> </ul> <h2 id="0840-2026-08-23">0.84.0 (2026-08-23)<a class="headerlink" href="#0840-2026-08-23" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a mesh partition now resolves from the workspace's own registered project, not from the directory the command happened to run in.</strong> <code>inbox read-primary</code>/<code>inbox count</code> derived the store partition from the caller's working directory, so a Primary could be told — by anti-hall's own Stop gate — to drain mail it structurally could not see, and running the prescribed command from the wrong directory risked writing a read cursor into an unrelated project's partition. Resolution now comes from the workspace's registered <code>repoKey</code>, via one shared helper (<code>companion/lib/devswarm-repokey.js</code> <code>registeredRepoKey</code>, precedence: fresh key → recorded <code>repoKey</code> → a non-hash <code>ownerKey</code>) that the CLI and the Stop hook both call, so the two can no longer disagree about which workspaces a session owns. Closes the gate/verb scope mismatch recorded as a known open issue in the DevSwarm KB.</li> <li><strong>Fix: cross-project commands no longer move a foreign project's workspace before deciding they are not allowed to touch it.</strong> <code>gate</code>, <code>ensure</code>, and <code>archive</code> re-homed a workspace registered to another project — copying messages and registry rows and rewriting its <code>ownerKey</code> — <em>before</em> their own ownership guard ran, so even an invocation that ended in <code>ok:false</code> had already mutated another project's state; <code>archive</code> additionally removed the live descriptor. The ownership check now runs first, and a refused call writes nothing.</li> <li><strong>Fix: <code>inbox ack</code> on a workspace the caller does not own no longer skips that workspace's mail.</strong> The cursor advanced even after the resolver had already refused the read, permanently stepping over messages nobody had seen.</li> <li><strong>Fix: the Primary Stop gate no longer hides genuinely drainable mail.</strong> Two paths could make real unread invisible to the gate: a persisted <code>repoKey</code> that had gone permanently stale, and a descriptor carrying only an <code>ownerKey</code>. Both now resolve through the same shared helper the CLI uses.</li> <li><strong><code>inbox count</code>/<code>inbox read</code> fail honestly instead of returning a silent zero.</strong> When the caller's project does not match the workspace's, the result now carries <code>known:false</code> plus the named <code>registeredRepoKey</code> and <code>callerRepoKey</code>, so "I cannot read this from here" can no longer be misread as "there is no mail".</li> <li><strong>Fix: defect fields are no longer silently truncated.</strong> The shared clamp cut over-length values at their cap and reported plain success, with nothing in the result or the stored record to show text had been lost — a survey of the live store found 41% of ruling notes and 27% of <code>observed</code> values sitting exactly at the cap, i.e. amputated. Truncation is now named in <code>scripts/defect.js</code>'s JSON result and on stderr, and marked inside the persisted value (<code>[truncated from N chars]</code>). The caps for <code>note</code> (300 → 1200), <code>claimed</code>, and <code>observed</code> were raised, bounded so a maximum-length field still cannot push a record past <code>MAX_LINE_BYTES</code>. The write itself never fails.</li> <li><strong>Fix: archived partitions nothing can ever read no longer warn every turn.</strong> The orphan detector counted an archived child's own outbox copy as unread mail with no reader, producing a permanent, unactionable per-turn warning (12 partitions in one real store). It now excludes exactly the set <code>healOrphanPartitions</code> classifies as <code>unhealable/archived-no-family</code>, by calling heal's own exported helpers rather than re-implementing the rule — with an equivalence test that fails CI if the two predicates ever drift. The count is preserved in a new quiet <code>archivedStranded</code> field rather than dropped, and the classifier fails open, so it can only ever quiet a warning it positively proved is unactionable.</li> </ul> <h2 id="0830-2026-08-23">0.83.0 (2026-08-23)<a class="headerlink" href="#0830-2026-08-23" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm child workspaces can no longer update the knowledge graph.</strong> A child workspace may still query the graph freely, but any command that writes or rebuilds it (<code>graphify update</code>, <code>--update</code>, <code>--obsidian</code>) is now refused there — maintaining the graph is the Primary's job. Previously every child workspace carried the same doctrine text and built its own duplicate copy of the graph, so many copies of the same data accumulated and none of them were ever read. Detection reuses the existing DevSwarm child-workspace signal, and the check fails open, so a Primary is never blocked.</li> <li><strong>The query-the-graph-first guard no longer blocks — it's advisory now.</strong> The guard already exempted subagents, and under anti-hall's own delegation-first doctrine essentially all code search happens inside subagents, so the block could never reach the work it was aimed at while still adding friction to the coordinator and occasionally blocking legitimate commands. The recommendation stays; the block is gone.</li> <li><strong>Write detection sees through more wrapping.</strong> <code>bash -c</code>, <code>eval</code>, and <code>$()</code>/backtick wrapping are now unwrapped up to a bounded depth before the write check runs, and a write flag takes precedence over the subcommand — <code>graphify query --update</code> counts as a write, not a query.</li> <li><strong>Fix: mesh routing is now deterministic when no workspace in a group is live.</strong> This closes the non-determinism noted as a known limitation in 0.82.0 — the send target used to fall back to whatever the registry happened to enumerate first, so successive sends from one session could land in different partitions minutes apart. The fallback now selects by most-recently-updated with a stable tiebreak, so the same set of rows always yields the same target. Live-workspace selection is unchanged.</li> <li><strong>Fix: a send now verifies the message is actually readable before reporting success.</strong> Earlier releases reported success straight from the write call; a claimed fix in an earlier version only added echo fields to the response and never verified anything. The send now re-reads the target partition and confirms the message is there. A verification <em>error</em> (couldn't check) is reported as unverified, not as failure — only a positive absence of the message is reported as a failed send.</li> <li><strong>Merged inbox reads are now bounded.</strong> A read that merges several registry partitions plus the file-based channel had no size limit. There is now a generous default limit with an override. A truncated read never advances a cursor past a withheld message, and the response says explicitly that it was truncated and how many messages were withheld.</li> <li><strong>Fix: archiving a workspace could be silently undone.</strong> Routine per-turn re-registration could recreate an archived workspace's descriptor and registry row, quietly reversing the archive. Re-registering an archived workspace id is now refused — unless the registration is for a genuinely different workspace that happens to reuse the same id, which is allowed and reported as such.</li> <li><strong><code>diagnose</code> now surfaces a partition that has exactly one live workspace.</strong> The condition was already detected internally but nothing in the output showed it, so a genuinely split mesh could still look healthy at a glance. Adds a degraded/warning field plus a human-readable line.</li> <li><strong>Fix: acknowledging mail now reports when the read cursor could not be saved.</strong> A store-side cursor write failure was previously swallowed behind a success result, so already-read messages could silently reappear as unread. The failure is now reported, naming the partition and channel that failed to persist. Delivery behavior is unchanged — messages are still acknowledged on the durable channel.</li> <li><strong>The defect CLI now rejects unknown flags instead of ignoring them.</strong> Previously an unrecognised flag was accepted, its value silently dropped, and success reported — which repeatedly produced defect records with empty fields. Also adds a <code>partial</code> ruling status for a fix that shipped in part, and documents that only <code>--sym-file</code>/<code>--repro-file</code> accept file paths (other value flags do not).</li> <li><strong>Fix: the status line no longer drops its second line under load.</strong> A too-short internal timeout made a healthy-but-slow render fail open into a misleading single-line status.</li> </ul> <p>Known limitations:</p> <ul> <li>Write detection does not see through <code>source <(...)</code> process substitution, shell aliases, or an absolute path straight to the <code>graphify</code> binary. This is a guardrail, not a security boundary — a determined bypass is still possible. Heredoc wrapping is now unwrapped for the common forms, but nested or multiple heredocs on one line, and <code><<<</code> here-strings, fall back to prior behavior, erring toward allowing.</li> <li>A merged inbox read's returned count may exceed the stated limit when truncation occurs — per-source boundaries are deliberately widened so that no cursor ever advances past a withheld message.</li> </ul> <h2 id="0820-2026-08-23">0.82.0 (2026-08-23)<a class="headerlink" href="#0820-2026-08-23" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a Primary could not see mail delivered to it.</strong> anti-hall addresses a workspace by a logical mesh id, but one logical Primary can legitimately have two registry rows — one keyed by the host tool's own workspace id, one by anti-hall's derived id. Sending resolved that group dynamically (picking the live row); the Primary's own inbox read used a fixed derived id and looked at only one partition. Mail delivered to the other row was invisible to the reader. The Primary's own inbox read verbs (<code>read-primary</code>, <code>peek-primary</code>, and any <code>--ack</code> read) now cover every partition in the mesh group, so a Primary using those verbs sees all of its mail regardless of which row a sender resolved to.</li> <li><strong>Fix: <code>peek-primary</code> and <code>read-primary</code> only read one of the two message channels.</strong> anti-hall keeps a file-based inbox and a store partition; the counting verbs already merged both, but the two verbs a Primary uses to actually read its mail did not — so the unread count and the visible mailbox could disagree substantially. Both now use the same merge.</li> <li><strong>Fix: cursor safety.</strong> A read cursor is only ever advanced past messages that were actually delivered to the caller — derived from what the read returned, never from a partition total. Prevents the failure where a cursor claims mail was read that was never delivered.</li> <li><strong>Fix: honest reporting when a read is incomplete.</strong> If a cursor cannot be persisted, the result now says so and names the partition and channel instead of reporting plain success. If the set of sibling partitions cannot be determined, the result marks the group unresolved and the totals partial, rather than silently reading a narrower set and presenting the total as complete.</li> <li><strong>By design: duplicate delivery over suppression.</strong> Where two messages cannot be proven to be the same message, both are delivered — a duplicate is visible and recoverable, while a dropped message is neither.</li> </ul> <p>Known limitations:</p> <ul> <li>The counting/reading verbs that take an explicit workspace id (<code>inbox count <id></code>, <code>inbox read <id></code>, <code>inbox messages <id></code>) still resolve a single partition and can under-report mail sitting in the sibling row, unlike the Primary's own read verbs above.</li> <li>When no row in a mesh group is currently live, which row a send resolves to is not deterministic — it falls back to registry enumeration order rather than a stable rule.</li> </ul> <h2 id="0810-2026-08-23">0.81.0 (2026-08-23)<a class="headerlink" href="#0810-2026-08-23" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a Primary's own inbox could be permanently unreachable.</strong> A <code>primary-*</code> descriptor created without an explicit inbox path kept <code>inboxPath: null</code> forever — child descriptors self-heal this every turn, but Primary rows had no equivalent, so reconcile failed the same way on every run and mail addressed to that mailbox was undeliverable. The register ensure-path now backfills <code>inboxPath</code>/<code>cursorPath</code> from the caller's defaults (only when empty, never overwriting an existing value), and the inbox pull now derives the standard default path instead of erroring — matching what three sibling call sites already did.</li> <li><strong>Fix: <code>diagnose</code> could report a dead workspace as live, and a live one as dead.</strong> Liveness came from a bare session id string with no expiry and no heartbeat correlation: a closed workspace's row stayed "live" forever, while a running workspace whose session id was never stamped read as dead. <code>diagnose</code>/<code>healthcheck</code> now derive the displayed liveness from the heartbeat instead. This is a <strong>display-only</strong> change — routing, fold/retire/adopt, and tombstone decisions still use the previous signal unchanged, deliberately, so no fold can newly happen as a result of this fix.</li> <li><strong>Fix: a workspace with a fresh heartbeat but an unclaimed session id showed as dead.</strong> A fresh heartbeat now takes precedence for the displayed liveness field; a row that is unclaimed with no (or a stale) heartbeat still shows not-live, and routing still treats "unclaimed" as never-live, unchanged.</li> <li><strong>Fix: small stores could starve behind large ones during update sweeps.</strong> Store processing order was raw directory order under a wall-clock budget, so one large store could consume the entire budget and leave trivial ones unprocessed run after run. Sweeps now process smallest-first.</li> <li><strong>Fix: unfixable orphan stores burned sweep budget and were invisible.</strong> Orphan stores with no descriptor — which anti-hall refuses to adopt, because adopting one would mean inventing ownership it can't verify — are now processed last and skipped cheaply once the budget is spent, and a capped sample of which ones and why is now reported instead of just a count. Nothing about an orphan is persisted: a store whose descriptor reappears is still adopted normally on the next pass.</li> <li><strong>Docs: <code>lastCompletedHash</code> in the sweep state file is observability-only, not a resume cursor</strong> — <code>pendingHashes</code> is the authoritative resume list. Documented to prevent a future reader from treating it as one.</li> </ul> <h2 id="0800-2026-08-22">0.80.0 (2026-08-22)<a class="headerlink" href="#0800-2026-08-22" title="Permanent link">¶</a></h2> <ul> <li><strong>New: a <code>/anti-hall:defects</code> skill on both ports</strong> — the durable defect channel shipped in v0.78.0 worked but had no discoverable entry point: no skill described it, and the runtime nudge that told an operator to run it pointed at a skill that did not exist, on both the Claude and Codex ports. It was only documented in files an agent doesn't load. The <code>devswarm</code> skill now also points at it, since that's the skill a DevSwarm session actually loads. A test now pins every <code>/anti-hall:<name></code> reference in the plugin to a skill that actually exists, on both ports, with no allowlist.</li> <li><strong>New: <code>--sym-file</code>/<code>--repro-file</code> on <code>defect.js report</code>.</strong> A report body can now be passed as a file instead of a shell argument, so a caller reporting a bug about shell quoting doesn't have to fight shell quoting to file it.</li> <li><strong>New: regression detection, derived from the reporter's installed version.</strong> A report matching a defect the maintainer already ruled <code>fixed</code> is now classified automatically: on a build at or past the version that claimed the fix, it's a genuine <code>regressed</code> reappearance; on an older build, it stays <code>fixed</code> and is flagged <code>staleBuild</code> instead — a report from someone who just needs to update, not a new regression. A regressed defect never auto-archives as resolved.</li> <li><strong>Fix: reporter identity no longer depends on the current directory.</strong> It used to come from the cwd basename, so a report filed from a scratch directory got an identity nothing else could match — silently breaking <code>--mine</code>, the only way a reporter sees rulings on their own reports. It now prefers an explicit flag, then an environment variable, then the project's repo key, matching old cwd-basename reports too so nothing already filed stops matching.</li> <li><strong>Fix: a Primary that explains itself is no longer escalated like one ignoring the gate.</strong> The parent gate invited a Primary to say a block was intentional but never used the answer — a stated reason accumulated toward escalation at the same rate as silence. A stated intent now suppresses escalation while the condition that triggered it is unchanged; the first block always still fires, and escalation resumes the moment the condition actually changes.</li> <li><strong>Fix: the wake watcher no longer declines in silence.</strong> Its three decline paths (not a DevSwarm session, an identity it couldn't resolve, a lock another watcher already holds) wrote to stderr and exited zero — indistinguishable from a healthy, quiet watcher to a Monitor caller who only sees stdout. Each refusal now prints one line on stdout naming the reason, so a caller who thinks they have wake coverage can tell they don't.</li> <li><strong>Fix: a partition with exactly one live row is now visible.</strong> The mesh split check recognized two-or-more-live and zero-live groups but missed the shape where one row of two is live — sends could land on a row nobody was draining while the diagnostic reported healthy. <code>send</code> and <code>diagnose</code> now also agree on which rows belong to a group (previously grouped by different identities), and <code>send</code> reports how many candidates it had and which row it chose.</li> <li><strong>Fix: "inbox unreadable" now says which file failed and why.</strong> A descriptor with no inbox path, an absent file, and an unreadable file all used to collapse into one unhelpful message, and a family member's failure could be misattributed to the Primary's own id. The gate now names the actual cause (missing field, absent file, or a read/parse failure with path and errno) and the actual workspace it happened to.</li> <li><strong>Fix: the updater no longer reports "unknown error" for a reconcile failure it already had the real cause for</strong> — it now surfaces the per-target error, matching a fix the doctor already had.</li> </ul> <h2 id="0790-2026-08-22">0.79.0 (2026-08-22)<a class="headerlink" href="#0790-2026-08-22" title="Permanent link">¶</a></h2> <ul> <li><strong>New: anti-hall now notices when its own knowledge goes stale.</strong> It already probed the installed DevSwarm version against its baseline; it now also compares the installed Claude Code CLI version against the version its harness KB was audited on, and compares the hook/skill counts <code>docs/KB.md</code> CLAIMS against what is actually on disk. All three probes are advisory, SessionStart only, deduped so they cannot nag, and fail open and silent.</li> <li><strong>Why the model check is a date, not a network probe:</strong> model facts (the current lineup, pricing) are not discoverable from the local machine, so a probe that cannot verify its claim would either invent an answer or fail constantly. Instead it records when the model KBs were last audited and advises past 60 days — an honest "not checked in N days" instead of a guess.</li> <li><strong>Both fired on their first run:</strong> the repo self-drift probe caught <code>docs/KB.md</code> claiming 49 hooks when there were actually 53, and the model KB staleness clock read 85 days.</li> <li><strong>Model KBs re-audited:</strong> Opus 4.8 is deprecated (superseded by Opus 5, June 2026); Fable 5 is the current flagship; Sonnet 5 held its introductory price; cache multipliers and current per-model pricing recorded. MODEL-POLICY routes by tier token and resolves to the newest family member at runtime, so executable routing was never affected — only the prose was stale.</li> <li><strong>New: downshift guidance.</strong> Moving off the flagship to conserve limits still needs 1M context, so Sonnet 5 is the target; Haiku 4.5 is disqualified at 200k despite being cheaper (still right for trivial leaf work).</li> </ul> <h2 id="0780-2026-08-22">0.78.0 (2026-08-22)<a class="headerlink" href="#0780-2026-08-22" title="Permanent link">¶</a></h2> <ul> <li><strong>New: a durable defect channel between agents and the maintainer.</strong> Any agent running anti-hall in any repo can now file a structured defect report via a CLI (<code>scripts/defect.js report</code>), and the maintainer session drains reports and writes back rulings (<code>scripts/defect.js rule</code>). Reports live in <code>~/.anti-hall/defects/</code>, one append-only NDJSON file per defect — cross-repo, survives session death, and never enters a git worktree.</li> <li><strong>Why not the mesh:</strong> bug reports about anti-hall's own messaging layer should not travel through that layer.</li> <li><strong>No stored state:</strong> status, occurrence count, and first/last-seen are derived from a defect file's own lines on every read, so two fields can never disagree — there is only one field. Status is the last ruling in append order (not by timestamp), so clock skew between writers cannot flip it.</li> <li><strong>Writes are verified:</strong> every append is re-read and matched byte-exact against the file; a write that cannot be confirmed reports <code>write-unverified</code> and exits non-zero instead of claiming success.</li> <li><strong>Nothing is deleted:</strong> acking and resolving both append a ruling line; rotation renames ruled-and-stale files into an archive directory. An open defect never moves regardless of age.</li> <li><strong>Bounded from the start:</strong> at most 200 open defects, 20 reports per defect, 64 KiB per file, 4 KiB per line, 1000 archived files — every cap refuses with a distinct outcome instead of silently dropping a report.</li> <li><strong>Discovery is quiet:</strong> one SessionStart line, at most once per day, never a Stop hook and never a forced acknowledgement — and it carries no reporter-supplied text, only counts and ages.</li> <li>New CLI verbs on <code>scripts/defect.js</code>: <code>report</code>, <code>list [--mine|--open]</code>, <code>show <fp></code>, <code>rule <fp> --status ack|fixed|wontfix|notabug|dup</code>, <code>archive</code>.</li> </ul> <h2 id="0771-2026-08-22">0.77.1 (2026-08-22)<a class="headerlink" href="#0771-2026-08-22" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a partition nobody was draining could have its mail moved the wrong way.</strong> When every row in a mesh identity group failed the liveness check, the fold fell through to whichever row sorted first in the registry — which for a real case would have forwarded a drained row's backlog into the row nobody reads. The fold now refuses to fold such a group and records it as needing attention. Refusing is deliberate: with no live row, a cursor-based tiebreak cannot tell an actively-drained row from one advanced once and abandoned. An unfolded group self-heals once a row goes live; a wrong-direction forward does not.</li> <li><strong>Fix: the check meant to catch that reported it as healthy.</strong> A split required two or more LIVE rows, so the dangerous shape — two rows, nobody draining either — never appeared, while the benign case of two live tabs on one worktree was flagged degraded. Diagnose now reports <code>liveSplit</code> and <code>deadSplit</code> separately; healthcheck counts dead splits, treats them as degraded, and prints a distinct warning.</li> </ul> <h2 id="0770-2026-08-22">0.77.0 (2026-08-22)<a class="headerlink" href="#0770-2026-08-22" title="Permanent link">¶</a></h2> <ul> <li><strong>Fix: a DevSwarm workspace could be counted twice in "needs attention."</strong> Two descriptor files sharing one worktree — a builder-id descriptor and a slug descriptor — each produced their own row, including duplicating the Primary's own row; a live session saw 3 where the app showed 2. Reads now collapse rows by identity family at read time. Nothing is retired or deleted — the store layer's refusal to retire a descriptor-backed row is deliberate, protecting two live tabs open on the same worktree.</li> <li><strong>Fix: unread mail stranded in an unregistered mesh partition.</strong> A Primary drained one row, saw "0 unread," and was technically right — four real messages sat unread in a different, unregistered partition for hours. The heal now adopts an unregistered partition and folds it into its identity family. Cursor reconciliation takes the MINIMUM across namespaces and only ever lowers a cursor, never raises one — taking the maximum would mark messages read that nobody had actually seen.</li> <li><strong>Fix: archived mail was unreachable.</strong> Descriptors move to <code>archived/</code> on workspace retirement, but the orphan lookup only ever read <code>workspaces/</code> — 105 of 110 reported "no descriptor" orphans actually had one, just in the wrong directory. Archived orphans are now forwarded (never adopted, never deleted), carrying provenance and capped at 30 days old.</li> <li><strong>Fix: mesh read/ack values meant different things in different places.</strong> <code>seq</code> was a positional ordinal in one verb and the physical mesh sequence number in its sibling; <code>unread</code> was a boolean in one place, a count in another, and a two-channel sum in a third; <code>read</code>/<code>ack</code> returned <code>ok: true</code> after acking nothing on an unregistered id; <code>send</code> silently discarded the hash it had already computed. Every value is now named explicitly, with the old keys kept as aliases so nothing that reads them today breaks.</li> <li><strong>New: <code>--message-file</code> and <code>--message-stdin</code> for <code>devswarm</code> message sends</strong>, so a body with newlines, quotes, or shell-special characters survives verbatim instead of being mangled by argv escaping. <code>send</code> now also echoes the <code>bytes</code> and <code>hash</code> it computed for the sent message.</li> <li><strong>New: counters name their own scope.</strong> <code>orphansWithUnread</code> (this repo's orphans with unread mail) is now reported separately from <code>orphanPartitions</code> (all stores, any repo), and the healthcheck output states which <code>repoKey</code> it scoped its counts to.</li> <li><strong>New: <code>ANTIHALL_DEVSWARM_ARCHIVE_FORWARD_MAX_AGE_DAYS</code></strong> tunes the archived-orphan forwarding age cap (default 30 days).</li> </ul> <h2 id="0760-2026-08-21">0.76.0 (2026-08-21)<a class="headerlink" href="#0760-2026-08-21" title="Permanent link">¶</a></h2> <ul> <li><strong>Security/data-loss fix: the DevSwarm parent gate could mark a child's question read without the child ever seeing it, with no way to undo it.</strong> <code>devswarm-parent-gate.js</code> counted the Primary's OWN just-sent outbound messages toward a workspace's "neglect" score, so a Primary that had just messaged a child could still trip the neglect check — and the remediation it then suggested, <code>inbox read-primary <child-id></code>, advances that child's read cursor. Following the guard's own advice would silently mark the child's real, unread question as read, with no un-ack and no recovery. Fixed by excluding own-sender rows from the neglect count and pointing the remediation at the read-only <code>inbox peek-primary <id></code> instead, which inspects the inbox without moving the cursor. A Primary genuinely reading its own inbox for the first time still correctly uses <code>read-primary</code>.</li> <li><strong>Fix: doctor's <code>[reconcile]</code> check collapsed every failure into <code>unknown error</code>, discarding the real per-target error.</strong> Root cause: one failing target had a <code>worktreePath</code> that no longer existed on disk, and Node reports a missing <code>cwd</code> passed to a spawned child as an <code>ENOENT</code> on the <em>executable</em>, not the directory — a generic message that gave no hint which target or why. <code>doctor-repair.js</code> now surfaces the real per-target errors (bounded, with a "+N more" summary for large batches) instead of swallowing them. A missing <code>worktreePath</code> is now a recognized benign skip (<code>worktreeMissing</code>), alongside the existing <code>locked</code> / <code>hivecontrolMissing</code> skips, and a descriptor with an existing archived counterpart is now detected and reported as <code>archivedDuplicate</code> — detection only; nothing is ever deleted.</li> <li><strong>Security: DevSwarm destructive-verb blocks bypassed via the <code>devswarm</code> alias.</strong> <code>command-guard</code> anchored its four DevSwarm destructive-command blocks on the literal verb <code>hivecontrol</code>, but <code>hivecontrol</code> is a thin shim that execs <code>devswarm</code> — the primary command name, equally on PATH. Running <code>devswarm workspace monitor</code>, <code>read-messages</code>, <code>message-child</code>, or <code>message-parent</code> directly therefore bypassed every block. Fixed with a shared verb set and an alternation that matches either name; both the Claude and Codex ports share the same hook file, so both are covered. Latent since the blocks were added, not a new regression.</li> <li><strong>New: DevSwarm version drift detection.</strong> anti-hall had no visibility into which DevSwarm version was installed, so its integration silently drifted from a 2.3.5 baseline to 2.5.1 without anyone noticing — a real risk because <code>command-guard</code> matches DevSwarm subcommands by literal string, so a renamed verb in a future release would make a block silently stop matching. A new SessionStart hook (<code>devswarm-version.js</code>, backed by a detached background refresh so session start is never blocked) probes the installed DevSwarm version, compares it against the shared baseline, and advises on a major/minor drift (patch-only stays silent; a downgrade is worded accordingly). The advisory dedupes on (installed, baseline) so it never nags twice for the same drift. Absent DevSwarm or unparseable output fails open and silent. Registered once, shared by both ports; also wired into the doctor health check.</li> <li><strong>Fix: <code>graphify-guard</code> recommended a wiki index file that doesn't exist.</strong> The guard's block message pointed agents at <code><graph>/wiki/index.md</code>, but <code>graphify update</code> never produces a <code>wiki/</code> directory — only <code>graph.json</code>, <code>GRAPH_REPORT.md</code>, and <code>manifest.json</code> — so every block sent the agent to a dead path. Fixed with an existence-checked fallback cascade (wiki index → <code>GRAPH_REPORT.md</code> → <code>manifest.json</code> → the bare <code>/graphify query</code> command) so the guard never names a path that isn't there. The block itself is unchanged, only the recommendation text.</li> <li><strong>Fix: unbounded per-session state growth under <code>~/.anti-hall</code>.</strong> Every per-session state file (task-tracker, speculation-guard, tasklist-guard, codex-nudge) was kept forever and never read back once its session ended. On a heavy multi-session machine this reached 47,000+ files across 71 days — and anti-hall's own <code>doctor.js</code> self-tests made it worse, orphaning one file per hook on every run. A shared <code>pruneStale()</code> helper is now wired into each hook's existing write path (no new hook, no new event): it removes same-prefix files older than 7 days, throttled to once per 6 hours so the cleanup scan never sits on the hot path, never touches the current session's own file, and fails open if pruning itself errors. This run pruned 42,236 orphaned files (<code>~/.anti-hall</code> 58,803 → 583 entries).</li> <li><strong>Docs.</strong> Recorded DevSwarm 2.5.x findings in the knowledge base: the unauthenticated local HTTP API surface, Claude Code harness messaging semantics, the v2.5.0 chat surface (a cloud-relayed, non-persisted feature distinct from workspace messaging), and the <code>transcriptByteOffset</code> agent-liveness signal.</li> </ul> <h2 id="0751-2026-08-08">0.75.1 (2026-08-08)<a class="headerlink" href="#0751-2026-08-08" title="Permanent link">¶</a></h2> <ul> <li><strong>Updater performance + UX.</strong> The <code>update</code> skill's post-update DevSwarm sweeps (<code>fold-all-stores</code>, <code>heal-registry-rows</code>) could turn a routine update into a minutes-long stall on a heavily-used machine with hundreds of per-project stores, since each sweep re-enumerated and re-walked every store in full on EVERY run regardless of whether anything had changed. Fixed with a run-once-per-version stamp (<code>~/.anti-hall/update-sweep-state.json</code>) so a re-run at the same version skips a completed sweep entirely (idempotent, fail-open, no-delete); a single shared store-hash enumeration reused across both post-update sweeps instead of each doing its own full directory listing; and per-store throttling via a bounded time budget (<code>ANTIHALL_UPDATE_SWEEP_BUDGET_MS</code>, default 20s) with a persisted resume list so a sweep that hits the budget stops cleanly mid-list and picks up where it left off next run, never re-walking already-processed stores. The one-time <code>ownerKey</code> migration and the ingest-daemon heal are similarly stamped per-version to skip a redundant re-run. Separately, git calls inside the updater (<code>gitState</code>, <code>gitPullFfOnly</code>) now time out after 20s instead of 60s, so a hung or slow remote fails fast into the existing fail-open path (report + continue with local state) rather than blocking the update for up to a full minute per call.</li> </ul> <h2 id="0750-2026-08-08">0.75.0 (2026-08-08)<a class="headerlink" href="#0750-2026-08-08" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm: partition-split identity family fix — evidence-based row-liveness ranking + unified freshest-live primitive.</strong> Two or more workspace rows could describe the same partition (mesh ID + workspace ID + repo path), creating silent state divergence when a parent restarted, drained, or the workspace was accessed across sessions. Root causes: (1) row-liveness ranking used only recency (<code>lastHeartbeatMs</code>), conflating activity with live-ness and demoting proof of active drains or session-authored heartbeats; (2) no unified selection primitive, causing three callers to independently pick rows with drifting tie-break rules. Fixed with a three-tier evidence-based ranking (<code>companion/lib/devswarm-liveness.js</code>, <code>rankRowLiveness</code>): session-reference integrity (does the row's <code>sessionId</code> exist?), drain-activity proof (explicit <code>draining</code> heartbeat + wall-time staleness within 20s), session-authored heartbeats (a heartbeat from the exact session listed in the row) — recency is now a final tie-break only. <code>devswarm-store.js</code>'s <code>selectFreshestLiveRow</code> centralizes that ranking (replaces three independent pickers in fold, read paths, and the parent table); all callers thread through it, eliminating drift. <code>fold</code> operation now merges realpath-proven duplicate rows (forward-before-tombstone, zero data loss): reads the full partition set, scores each row, merges older rows into the freshest-live row via <code>mergeHeartbeatRows</code>, and prunes obsolete copies. Parent inbox now applies a non-acking <code>peek-primary</code> read variant (reads the primary durable NDJSON inbox first, then the store-only backlog) and re-runs liveness against the union before rendering. Parent-gate carries a two-axis escalation ceiling (bounded plain-backlog nag count AND a separate bounded unanswered-question count) so a persistently-neglected workspace or an unanswered question can never nag forever without a human being told to look. Titles-not-ids instruction added to primary-facing injections (parent table + prose guidance) so end-users see <code>Workspace: "Fix the parser gate…"</code> not <code>Workspace: "mesh-abc123def"</code>; mesh IDs remain in CLI operations. Two new self-heal migrations (both platforms, idempotent, fail-open, no-delete): <code>update.js</code>'s and <code>doctor-repair.js</code>'s <code>migrateOwnerKeys</code> backfill the <code>ownerKey</code> descriptor field and re-home an active hash-bucket split across both active and archived descriptors. Heartbeat callers instrumented (<code>heartbeatCallersLogPath</code>/<code>appendHeartbeatCallerLog</code>) so an unidentified caller invoking <code>heartbeat</code> without <code>--session</code> is logged for attribution.</li> </ul> <p><strong>[CORRECTION 2026-08-22]</strong> An audit (<code>git log -S</code> against every named symbol, cross-checked against the actual shipped files) found this entry mis-describes what <code>db5822b</code> actually shipped. Corrected, without deleting the original text above: - <code>companion/lib/devswarm-liveness.js</code> never existed under that name — the file is <code>companion/lib/devswarm-liveness-select.js</code>. - <code>rankRowLiveness</code> never existed under that name — the exported function is <code>pickFreshestLive</code>. - <code>selectFreshestLiveRow</code> never existed under that name — the store-side function is <code>resolveSenderRegistryId</code> (<code>companion/lib/devswarm-store.js</code>), which delegates to <code>pickFreshestLive</code>. - <strong><code>mergeHeartbeatRows</code> was never built.</strong> It exists in no commit and no file — grep of the full tree and <code>git log -S</code> both return zero hits outside this CHANGELOG's own prose. This was announced and never shipped; flagged as the most serious item per this repo's false-completion stance. - "all callers thread through it, eliminating drift" is <strong>false as written</strong>: <code>hooks/devswarm-parent-gate.js</code> (verified via <code>grep</code>) has zero references to <code>devswarm-liveness-select.js</code>, <code>pickFreshestLive</code>, or <code>resolveSenderRegistryId</code>. The 0.75.0 work landed in the store/registry-row layer only — the descriptor layer that renders the parent gate table never adopted it. - Signal (b), described above as "drain-activity proof (explicit <code>draining</code> heartbeat + 20s wall-time staleness)", did not ship as described. The actual shipped signal (verified in <code>devswarm-liveness-select.js</code>) is comparative cursor-row evidence (<code>hasCursorRow</code>/<code>cursorValue</code>/ <code>messageCount</code>) — there is no <code>draining</code> field and no 20-second window anywhere in the file. - Signal (a), described above as "session-reference integrity (does the row's <code>sessionId</code> exist?)", is mis-described. The shipped check is alias detection: <code>sidStr !== String(d.id) && groupIds.has(sidStr)</code>. - Accurate as originally written (no correction needed): signal (c) session-authored heartbeats, <code>peek-primary</code>, <code>migrateOwnerKeys</code>, and <code>heartbeatCallersLogPath</code>/<code>appendHeartbeatCallerLog</code>. - The descriptor-layer gap this entry implied was already fixed is now actually fixed, in <code>d1c8625</code> ("fix(devswarm): collapse identity families so one workspace is counted once" — reader-side identity-family collapse).</p> <ul> <li><strong>Agent-reliability rails.</strong> Three small, independently fail-open additions hardening background/teammate-agent and session-resume reliability:</li> <li><code>verify-first-subagent.js</code> (SubagentStart) now tells every spawned agent to <code>SendMessage</code> its final report to the coordinator BEFORE finishing (a bare turn-end silently loses it) and to never end a turn waiting on a background task (its completion notification routes to the main session, not to the waiting agent — run long commands foreground or poll the output file).</li> <li><code>tasklist-guard.js</code>'s "tracked NO tasks" nudge now points at a PRIOR session's <code>.anti-hall/handovers/<today>/*/state.md</code> snapshot (if one exists) so a session resuming with an empty task list recreates it from the snapshot instead of inventing one from scratch.</li> <li>Resume-verification enforcement: <code>handover-resume.js</code> now records a small per-session marker whenever it injects a guided-resume pointer, and its injected instructions ask the agent to append a <code>resume-verified: <ISO timestamp> -- <summary></code> line to the HANDOVER file after running the checklist. <code>tasklist-guard.js</code> backs this up mechanically — if a resume injection happened this session, file-changing work then occurs, and no <code>resume-verified:</code> marker ever lands in the referenced file, ONE capped Stop-hook block fires naming it. The handover skill (both platforms) documents the new marker line and its enforcement.</li> </ul> <h2 id="0740-2026-08-08">0.74.0 (2026-08-08)<a class="headerlink" href="#0740-2026-08-08" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm: unpushed/no-upstream risk surfacing + git-verified merged gate (report-only).</strong> A child could self-declare <code>merged</code> while its work sat unpushed and un-reviewable — nothing detected the single-copy-on-disk state. Two independent, fail-open git ground-truth probes land in <code>companion/lib/devswarm-git-truth.js</code> (<code>gitPushState</code>, <code>gitMergedInto</code>), both using the same argv-array <code>spawnSync</code> convention as <code>liveness.js</code>'s <code>defaultGitCommitTs</code>: never shell-interpolated, a 4s timeout, and <code>null</code> (never a fabricated fact) on any probe failure. <code>devswarm-child-turn.js</code>'s <code>writeHeartbeat</code> attaches one <code>gitPushState</code> probe per turn (<code>noUpstream</code>/<code>unpushed</code>, omitted entirely when unresolved); <code>devswarm-store.js</code>'s <code>computeSummary</code> threads that into each workspace projection along with a new <code>merged_verified</code> gate row (never a new persisted-shape column); <code>devswarm-parent-inbox.js</code> renders a <code>⚠ no upstream</code> / <code>⚠ N unpushed</code> / <code>merged (unverified)</code> marker next to the workspace title in the roster table. <code>scripts/devswarm.js</code>'s <code>cmdGate</code> runs the ancestry check whenever <code>--set merged</code> fires and persists <code>merged_verified</code> alongside <code>merged</code>, warning on stderr when the check resolves false (a squash/rebase merge legitimately breaks ancestry even though the work IS merged) — the gate is set either way. Strictly REPORT-ONLY throughout: nothing here blocks, kills, or archives; it only makes an existing self-declared state visible before a human acts on it.</li> </ul> <h2 id="0730-2026-08-07">0.73.0 (2026-08-07)<a class="headerlink" href="#0730-2026-08-07" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm child inbox-neglect, fixed (a downstream project field incident).</strong> Root cause had three parts: (1) <code>send --to</code> writes ONLY to the store — never the durable NDJSON inbox — so any reader checking NDJSON alone (liveness's unread backlog, the child/parent Stop gates) was blind to a mesh-direct backlog that could climb 14→15 while every gate read 0; (2) a child has no mid-turn re-entry point — <code>devswarm-child-turn.js</code> fires once per UserPromptSubmit, never during a long autonomous task, so a poke never re-surfaced; (3) the existing remedy path silently mutated state instead of reporting. Fixed with a shared LOSS-FREE UNION unread primitive (<code>companion/lib/devswarm-unread.js</code>, NDJSON ∪ store-only, hash-deduped — <code>scripts/devswarm.js</code>'s CLI now delegates to it too instead of a third drifting copy), wired into liveness and both Stop gates so <code>notDraining</code> and <code>oldestUnreadAgeMs</code> are computed against the true union; and a new throttled <code>devswarm-child-drain.js</code> PostToolUse/Bash hook (child-only, mirrors the Primary-only reply-tracker) that gives a child a re-entry point on every tool call, re-injecting only when the unread count changes or a 10-minute window elapses. The child/parent Stop-gate messages now split "CHILD NOT DRAINING" from "YOUR INBOX" and name the workspace by title, and the heartbeat verdict path is unified with the union read. Registered on both platforms (<code>hooks/hooks.json</code> and <code>codex/hooks/hooks.json</code>).</li> <li><strong>Handover skill hardening (7 threads, owner amendments 2026-08-07).</strong> Self-write mandate: the handover skill now states explicitly that the agent holding session context must write it — never delegate to a subagent, which loses decision/trial fidelity — and documents this as the explicit exception to delegation-first (edit-guard already allows direct writes under <code>.anti-hall/handovers/**</code>). <code>model-routing-guard.js</code> adds a capped-once-per-session advisory that fires independently of model-tier routing when a spawn looks like it's being asked to write/prepare a handover. <code>state.md</code>'s template gains a Task list snapshot table, and both <code>HANDOVER.md</code>'s Open items and the next-session usage steps point at it so a fresh session reconciles its task list before working. <code>edit-guard.js</code> now redirects a NEW handover-named <code>.md</code> write outside <code>.anti-hall/handovers/**</code> back to the correct location instead of silently allowing it anywhere under cwd (existing files and ambiguous cwd are unaffected). <code>handover-resume.js</code> emits a one-line negative report on a clear/compact source when no handover is found at all, naming the wrong-location write as the likely cause. <code>tasklist-guard.js</code> adds two capped, ride-along (never a new block on their own) Stop advisories: boundary-surfacing when a block is about to fire with no handover dir yet, and a staleness rail when file-changing work happened after the newest <code>HANDOVER*.md</code>'s mtime. The skill also adds an explicit quiesce gate (enumerate/await/park every running background item, <code>⏳</code>/<code>✅</code> status format) and a terminal declaration rule — the <code>✅ SAFE TO COMPACT</code> line is the last act of the turn; any work after it makes the handover stale and requires refresh + re-declare.</li> <li><strong>Fix:</strong> <code>devswarm-supervisor.js</code>'s <code>sweepOnce</code> built the object literal passed to <code>computeLiveness</code> without <code>env</code> (it was computed locally but never threaded through), so liveness always fell back to <code>process.env</code> instead of honoring an injected/test env.</li> </ul> <h2 id="0720-2026-08-07">0.72.0 (2026-08-07)<a class="headerlink" href="#0720-2026-08-07" title="Permanent link">¶</a></h2> <ul> <li><strong>New <code>handover</code> skill (Claude + Codex).</strong> Writes a comprehensive, multi-file session handover under <code>.anti-hall/handovers/</code> — task state, decisions, open threads, and verification status — with a global index and sequence chaining across handovers, so a fresh session can resume without re-deriving or guessing anything. Mirrored for Codex under <code>plugins/anti-hall/codex/skills/anti-hall-handover/</code>.</li> <li><strong>New <code>handover-resume</code> SessionStart hook.</strong> On session start (including after <code>/clear</code> or compaction), surfaces the latest handover and guides a structured resume, superseding the lossy default compact summary. Registered on both the Claude <code>hooks.json</code> and the Codex hook subset (<code>install-codex.js</code>, Codex <code>hooks.json</code>).</li> <li><strong>New KB doc:</strong> <code>docs/KB-session-handover.md</code> (24+ sources) documenting the handover design and resume flow.</li> </ul> <h2 id="0714-2026-08-05">0.71.4 (2026-08-05)<a class="headerlink" href="#0714-2026-08-05" title="Permanent link">¶</a></h2> <ul> <li><strong>New always-apply <code>autonomous-execution</code> discipline (both platforms).</strong> Once the user authorizes a scope ("do all" / "yes" / "go", a task list, or a named process), the agent now runs the WHOLE scope to done — driving each item build → review → fix → deploy → verify and acting on background results as they land — instead of pausing to re-confirm steps the authorization already covered. Naming/wording, running an already-requested process, shipping already-reviewed work, and picking between roughly-equivalent options are no longer check-in points; the agent picks the better one, notes it, and proceeds, reporting ONE consolidated end result. Injected via <code>verify-first-full.js</code> (SessionStart) and <code>verify-first-subagent.js</code> (SubagentStart, phrased for the worker framing: the assigned task IS the authorization), and mirrored for Codex as a new "Autonomous execution (always apply)" section in <code>AGENTS.md</code>.</li> <li><strong>The discipline lowers no existing bar.</strong> It explicitly cross-references, rather than overrides, the stop-points that already existed: a credential/secret the agent cannot supply, any destructive or irreversible action, deletions still requiring explicit confirmation, <code>DONE</code> still meaning VERIFIED (Positive Rule 6), and scope expansion past what was authorized still requiring confirmation (SCOPE & FIDELITY).</li> <li><strong>Output-style guidance strengthened (rule K + its mirrors).</strong> The rich, scannable presentation (tables for status/comparisons, <strong>bold</strong> verdicts, <code>code</code> for flags/paths, a single leading status glyph as SIGNAL) is now stated as the DEFAULT for every user-facing report rather than an occasional flourish, and sliding back to bare plain text over a long session is named as DRIFT to correct. Applied to <code>verify-first-orch.js</code> rule K, the <code>verify-first-subagent.js</code> scannability line, the <code>AGENTS.md</code> scannability bullet, and the Codex <code>anti-hall-orchestration</code> skill mirror. Emoji-as-signal-never- decoration is unchanged.</li> </ul> <h2 id="0713-2026-08-04">0.71.3 (2026-08-04)<a class="headerlink" href="#0713-2026-08-04" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm ingest installer resolves hivecontrol via a robust tiered chain.</strong> <code>install-devswarm-ingest.js</code> now resolves hivecontrol via <code>ANTIHALL_DEVSWARM_HIVECONTROL</code> → a persisted last-known-good cache at <code>~/.anti-hall/devswarm/hivecontrol-path.json</code> → a login-shell lookup → known install locations, and caches every success, so a reinstall from a minimal-env caller (hook, doctor repair, bare subagent shell) no longer silently bakes a PATH-less daemon that ENOENTs and stops ingesting; on total miss it now prints a loud stderr warning instead of failing silently, while still installing fail-open.</li> </ul> <h2 id="0712-2026-08-04">0.71.2 (2026-08-04)<a class="headerlink" href="#0712-2026-08-04" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm installer rejects unknown/mistyped flags.</strong> <code>install-devswarm-ingest.js</code> now rejects unknown/mistyped flags and handles <code>--help</code>/<code>-h</code> by printing usage and exiting, instead of silently running a full daemon install on an unrecognized flag.</li> </ul> <h2 id="0711-2026-08-04">0.71.1 (2026-08-04)<a class="headerlink" href="#0711-2026-08-04" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm ingest daemon busy-spin fix (elapsed-aware pacing, <code>7242575</code>).</strong> The ingest daemon's success-path loop spun as fast as the OS would schedule it instead of respecting <code>intervalSec</code>, causing a <code>data.kalloc.1024</code> macOS kernel-allocator leak and ~11% idle CPU. The loop now sleeps for the remaining time in the interval (elapsed-aware, ~1 iteration per <code>intervalSec</code>), clamped to a finite <code>MAX_PACE_MS</code> ceiling so a misconfigured interval can't overflow back into busy-spinning.</li> <li><strong>Delivery to running installs (<code>c09cf9f</code>, <code>3d0d03e</code>).</strong> The daemon now stamps its plugin <code>codeVersion</code> into the heartbeat; <code>doctor-repair</code> restarts an alive daemon found running stale code (self-clearing — no restart-bounce once it's current), and the updater force-restarts the ingest daemon after <code>update.js</code> runs so an already-running unbounded-loop daemon actually re-execs onto the paced code instead of persisting until its next natural restart.</li> <li>Measured in a controlled harness: ~362x fewer fork/exec spawns per second (busy-spin ~384/s -> paced ~1/s).</li> </ul> <h2 id="0710-2026-08-04">0.71.0 (2026-08-04)<a class="headerlink" href="#0710-2026-08-04" title="Permanent link">¶</a></h2> <ul> <li><strong>Append-only reply-state redesign (<code>recordReply</code>, merge <code>adfd61c</code>).</strong> DevSwarm parent decide-gate reply-state moved from a lockfile read-modify-write of a merged JSON object to an append-only JSONL log — <code>recordReply</code> is now one <code>O_APPEND</code> write with no lock, <code>readReplyState</code> folds the log on read, a fail-closed newline separator guards partial records, each record is capped at 480 bytes, and the fold accumulator uses <code>Object.create(null)</code> so a <code>__proto__</code>-named sender survives the fold instead of polluting the prototype. Ships with a loss-safe forward migration (<code>migrateReplyState</code>) wired into both <code>update.js</code> and <code>doctor-repair</code>, with an accepted, documented residual in the final write window. Structurally eliminates the disclosed steal-branch TOCTOU rather than patching around it.</li> <li><strong>Emoji-as-signal rule propagated to subagents + Codex (<code>e6f6e3f</code>).</strong> Rule K (status glyph as SIGNAL, never decoration) is now also injected at <code>SubagentStart</code> and in the Codex orchestration skill, not just the orchestrator's <code>SessionStart</code>.</li> <li><strong>Test-store-leak hardening + read-only leaked-bucket audit (merge <code>9b89fdd</code>).</strong> Fixed 4 doctor tests' <code>HOME</code>-default landmine, where an unset test override silently fell back to the real home directory. Added a new READ-ONLY store audit/classifier (REAL/GARBAGE/UNKNOWN) and a leak-report CLI whose <code>--out</code> is guarded (realpath canonicalization against the store root, <code>O_EXCL</code>/<code>O_NOFOLLOW</code> write, a distinctive report marker) so it can never overwrite a production <code>devswarm.db</code>; ambiguous or unreadable evidence degrades to UNKNOWN, never GARBAGE. Detection-only — it never deletes.</li> <li><strong><code>register-primary</code> records the real Claude <code>session_id</code> (merge <code>dfad611</code>).</strong> <code>--session</code> now defaults to <code>CLAUDE_CODE_SESSION_ID</code> (previously the workspace hash), so Primary registry rows resolve their transcript for liveness reads. Also corrected a KB doc's broken <code>$CLAUDE_SESSION_ID</code> reference.</li> </ul> <h2 id="0701-2026-08-04">0.70.1 (2026-08-04)<a class="headerlink" href="#0701-2026-08-04" title="Permanent link">¶</a></h2> <ul> <li><strong>Liveness-aware roster read — dormant-tier demotion (<code>7f253cd</code>).</strong> A DevSwarm mesh/ registry row outlives its workspace — closing a workspace in the DevSwarm app deletes nothing (registry row, worktree, descriptor, <code>hivecontrol workspace list</code> entry all survive) — so a row whose newest known activity signal (heartbeat timestamp, transcript mtime, or supervisor verdict) is at least <code>ANTIHALL_DEVSWARM_DORMANT_MS</code> old (default 30 min) is now labeled <code>dormant</code> (rank 5, sorts last, below <code>active</code>) instead of reading as still-active forever. Demotes, never hides — a dormant row still renders with its unread count, only ranked last; never overrides <code>escalated</code>/<code>stale</code>/<code>archive-ready</code>; fail-open on any read error. <code>scripts/devswarm.js</code> <code>roster</code> carries the identical hint so the two surfaces can't disagree.</li> <li><strong>Edit-guard: coordinator handover/compact-prep doc exclusion (<code>f6e2931</code>).</strong> The Primary write-block now excludes coordinator handover / compact-prep <code>.md</code> docs from the block, closing a false-positive that stopped a Primary from writing its own session handoff. Bypass-safe: gated by basename + <code>.md</code> extension, cwd-containment checked, hardlinks rejected.</li> <li><strong>DevSwarm <code>archive</code>: shortId/prefix resolve (<code>78425ae</code>).</strong> <code>archive <id></code> now resolves the workspace id by an unambiguous shortId/prefix, so the id shown in the roster/injection table is directly archivable without pasting the full id. An ambiguous prefix (matches more than one row) archives nothing and lists the candidates instead; exact full-id behavior is unchanged; <code>isSafeId</code> still gates (no <code>/</code>).</li> </ul> <h2 id="0700">0.70.0<a class="headerlink" href="#0700" title="Permanent link">¶</a></h2> <ul> <li><strong>DevSwarm mesh/store hardening — message-loss fix (merge <code>5856eb9</code>).</strong> <code>archive</code> used to tombstone exactly one registry row per archive, so a worktree that had been archived-and-reregistered across multiple generations could still hold LIVE sibling rows sharing that worktree — and a live row is what makes a message get forwarded there instead of to a genuinely-live partition. <code>foldArchivedRegistryRows</code> (new, <code>scripts/devswarm.js</code>) now folds ALL same-worktree registry rows for an archived id and picks the forward survivor by LIVENESS (<code>pickArchiveForwardSurvivor</code>), fixing a P0 where a real unanswered question could be forwarded into a dead partition nothing drains. Ships as a dual-path persisted-shape migration — wired into BOTH <code>update</code> (<code>skills/update/scripts/update.js</code>) and <code>doctor --fix</code> (<code>hooks/lib/doctor-repair.js</code>'s <code>migrationFix('fold-archived-rows', ...)</code>) — and is idempotent (a retired row is gone, a second run is a no-op), fail-open-honestly (never silently reports success on a raised error), and no-delete (message rows are never deleted; only registry rows are tombstoned after their unread is forwarded). Also in this merge: decide-gate follow-ups — durable per-project reply-state, cap escalation, and a <code>recordReply</code> lock closing a lost-update race under concurrent writers. Hardened via a multi-round deadly-loop.</li> <li><strong>Injection token / archived-row read-filter fixes (merge <code>58c307d</code>, incl. <code>f4d26ef</code>).</strong> A NEW <code>archivedOnlyIds</code>/read-side filter in <code>companion/lib/devswarm-store.js</code> excludes a genuinely archived workspace (<code>archived/<id>.json</code> present, <code>workspaces/<id>.json</code> absent) from the LIVE per-turn projection immediately — without needing a <code>doctor</code> run first, closing the window where a stale ACTIVE row inflated every DevSwarm injection. An archived workspace with real unread still surfaces as an <code>orphans[]</code> entry instead of going dark (no lost signal); the predicate structurally cannot hide a live row (a live workspace has its own descriptor by definition) and fails open to an empty set on any read error. <code>cmdArchive</code> (<code>scripts/devswarm.js</code>) gained a descriptor-conflict self-heal (<code>archivedTombstoneIsOrphaned</code>, decided by inode not by registry state, fail-closed on any incomplete scan) that unblocks re-archiving an id whose <code>archived/<id>.json</code> was a stale leftover from a prior archive generation. The per-turn parent-inbox STOP imperative (<code>hooks/devswarm-parent-inbox.js</code>) is softened to advisory wording for the normal tier; the loud (urgent/high) tier is untouched.</li> <li><strong>Test-flake hardening (<code>fe0d901</code>, <code>b99eafb</code>).</strong> 5 doctor/limit test files hardened against contention-killed subprocesses (<code>spawnSync</code> timeout + signal tolerance); no assertion was weakened.</li> <li><strong>Emoji-as-signal orchestration guidance (<code>873467d</code>).</strong> Rule K in <code>verify-first-orch.js</code> now names the exact signal glyphs (✅/❌/⚠️) instead of a vague "emoji = signal, not decoration" — same intent, less ambiguity for the model to over-apply.</li> </ul> <h2 id="0690">0.69.0<a class="headerlink" href="#0690" title="Permanent link">¶</a></h2> <ul> <li><strong>Harness Phase-1 hooks: <code>output-verify-guard</code> + <code>failure-root-cause-nudge</code>, plus a sandbox doc section.</strong> Adopted from the Fable-reviewed harness-feature adoption plan (<code>docs/superpowers/specs/2026-08-01-harness-feature-adoption.md</code>). <code>output-verify-guard.js</code> (PostToolUse, matcher Bash) scans a Bash tool call's own output for a passing signal (e.g. "8 passed", "PASS") alongside a failing signal (e.g. "2 failed", a confirmed non-zero exit) in the SAME run — the shape of a partial-pass summary that is easy to mis-report as a clean "tests pass" — and annotates (never blocks). <code>failure-root-cause-nudge.js</code> (PostToolUseFailure, matcher Bash) adds one short reminder pointing at <code>/anti-hall:root-cause</code> when a Bash command exits non-zero, deliberately terse since OMC already injects its own root-cause nudges in this harness. Both are advisory-only and fail-open. <code>AGENTS.md</code> gained a SANDBOXING sub-bullet recommending the harness's <code>/sandbox</code> mode for autonomous build-heavy sessions (doc-only, no mechanical enforcement).</li> <li><strong>DevSwarm parent decide+reply gate.</strong> The Stop-gate (<code>devswarm-parent-gate.js</code>) could previously be satisfied by a Primary merely <em>reading</em> a child's blocking <code>--question</code> — it now requires an OBSERVED reply: a new <code>devswarm-parent-reply-tracker.js</code> (PostToolUse, matcher Bash, Primary only) watches for a successful <code>devswarm.js send --to <id> --question</code>-style direct send and records it via a new durable, per-project reply-state store (<code>companion/lib/devswarm-reply-state.js</code>, keyed by a repo-scoped <code>repoKey</code> so it survives a read/ack and new Claude sessions, not a short-lived <code>session_id</code>). The forced-ack cap (<code>MAX_BLOCKS</code>) can no longer silence an unanswered question forever — once exhausted it escalates once with distinct wording instead of going quiet. Every Primary turn also re-asserts the decide+reply obligation. <code>recordReply</code>'s read-modify-write is now lock-protected (an O_EXCL lock with retry/backoff) after a reproduced race lost entries under concurrent writers. New persisted shape: a <code>needs_reply</code> column on mesh rows (sqlite: additive <code>ALTER TABLE ADD COLUMN</code>, idempotent + fail-open + no-delete, applied on every store open; journal backend: absent field reads as <code>false</code>) plus the new reply-state JSON file (fails open to <code>{}</code> when absent — no prior shape to migrate from). Hardened via a 6-round deadly-loop (Reviewer/Auditor/Critic) that found and fixed a read-vs-reply lifetime mismatch, an identity-space mismatch, 3 row-copy paths dropping the new flag, a Codex upgrade-detection gap, and the concurrent-write race above.</li> <li><strong>DevSwarm child self-continue directive.</strong> A child in a multi-round autonomous task (deadly-loop, iterative fix waves) previously had no mechanical push to keep working between rounds — ending its turn to "check in" cost a full wake-cycle (supervisor cron) before it resumed. A new per-turn directive in <code>devswarm-child-turn.js</code> tells the child to keep issuing tool calls across rounds within the same turn, reserving <code>Stop</code> for a genuine block, final completion, or an unrecoverable error. Shared verbatim with the Codex port (registered unmodified from <code>hooks/devswarm-child-turn.js</code>).</li> <li><strong>Windows support dropped.</strong> Removed <code>windows-latest</code> from the CI matrix (ubuntu-latest + macos-latest remain) and corrected support-claim language across <code>plugin.json</code>, <code>README.md</code>, <code>llms.txt</code>, and <code>AGENTS.md</code> from "runs on Windows/macOS/Linux" to macOS + Linux supported, Windows untested and not officially supported. Pure-Node cross-platform code is unchanged — Windows may still work, it's just no longer tested or claimed as supported.</li> </ul> <h2 id="0682">0.68.2<a class="headerlink" href="#0682" title="Permanent link">¶</a></h2> <ul> <li><strong>Doctor's install-divergence check now walks the full shipped tree instead of a 2-file sample.</strong> 0.68.1 added the installed-vs-source comparison (below) but only checked two files (the wake watcher and <code>monitors.json</code>), so real drift elsewhere — e.g. a stale <code>hooks/api-guard.js</code> — went unreported: exactly the silently-frozen-cache failure the check exists to catch. It now sha256-hashes every shipped file under the plugin root and reports content mismatches plus files present in only one of the two roots (installed vs marketplace source), bounded by per-file size (5 MB) and total-diff (20) caps so a pathological tree can't run away. Regression test covers <code>api-guard.js</code> drift, asymmetric file presence, and a clean tree.</li> </ul> <h2 id="0681">0.68.1<a class="headerlink" href="#0681" title="Permanent link">¶</a></h2> <p>This release exists primarily so the 0.68.0 gate fix can actually reach installed users.</p> <ul> <li><strong>Direct messages sent with <code>send --to</code> are readable by their recipient.</strong> They were written to the store while <code>inbox read</code>/<code>count</code> read a per-workspace NDJSON that only the native queue populates; with native messaging unavailable, nothing wrote it, so messages were unreadable while the projection counted them as unread — a workspace was told to commit finished work, never saw it, and the work was abandoned. Both channels are now unioned and deduped by content hash; <code>ack</code> advances the store cursor under the existing ownership check; <code>read</code> stays non-mutating.</li> <li><strong>Doctor detects an installed copy that diverged from its source.</strong> The harness runs a version-pinned cache dir, and the updater never overwrites an existing one — so a cache populated mid-release freezes pre-release code under the released version number and no later update can refresh it. That happened to 0.68.0: the executed copy carried an already-fixed gate bug and nothing reported it. Doctor now compares installed against source at the same version and names the differing files and the remedy. It also reports whether monitors.json is present in the installed root, since its absence removes the only mechanical arming path.</li> <li><strong>Note for anyone whose 0.68.0 install predates this release:</strong> a version-pinned cache cannot be refreshed in place by the updater, which is why this is a version bump rather than a silent re-sync.</li> </ul> <p>Known limitations:</p> <ul> <li>The 0.68.0 limitations still apply (see below).</li> <li>If a native message reaches the NDJSON but the best-effort store parity write fails, the union reports more unread than the store-only projection.</li> </ul> <h2 id="0680">0.68.0<a class="headerlink" href="#0680" title="Permanent link">¶</a></h2> <p>DevSwarm idle-wake gains a Monitor-tool path that wakes an idle orchestrator the moment new direct mesh mail lands, instead of waiting up to 5 minutes for the cron fallback. The cron fallback is never removed — the Monitor path is strictly additive, so nothing regresses if it is unavailable.</p> <ul> <li><strong>Monitor-based wake watcher.</strong> New <code>companion/lib/devswarm-wake-watch.js</code> plus a <code>monitors/monitors.json</code> plugin manifest. Pure <code>tick(state, snapshot)</code> core split from the IO runner so the logic is unit-testable without spawning a process. Edge-triggers on a monotonic per-workspace message counter, never on file mtime — an mtime trigger fires on every heartbeat write and would exhaust the ~19-20 event notification budget.</li> <li><strong>Mechanical single-instance lock.</strong> A duplicate watcher self-refuses: zero stdout (stdout is what wakes the agent), a one-line stderr notice, exit 0.</li> <li><strong>Four projection and silent-failure fixes</strong> merged alongside it: devswarm-pull no longer ignores a subprocess exit status (and no longer broadcasts "merge completed" on a FAILED merge); duplicate heartbeat broadcasts no longer saturate the projection; a corrupt truncated archived descriptor is surfaced instead of silently dropped; and the test suite no longer registers real launchd jobs on the host.</li> <li><strong>The wake watcher is gated on an active DevSwarm session.</strong> <code>monitors.json</code> declares <code>"when": "always"</code>, so without a gate the watcher would start for every personal-scope install — including users who never touched DevSwarm — costing them a polling process, an unsolicited transcript line, and state directories. It now exits quietly unless the session is genuinely DevSwarm-active, and fails closed if the check itself errors.</li> </ul> <p>Known limitations:</p> <ul> <li>Monitor-based wake is verified by construction, not by a live end-to-end runtime test. The monitors manifest test asserts schema and shape only and says so explicitly.</li> <li>Project-scope (<code>@skills-dir</code>) plugin installs do not load background monitors at all. Those users keep the cron wake path and get no Monitor path.</li> <li>Watcher identity is derived from the git worktree, so two orchestrators in the same repository resolve to the same identity; the second watcher refuses its lock and only the lock winner gets low-latency wakes. The loser still has cron.</li> <li>If a summary file is missing or unreadable, the watcher's counter reads as absent rather than as an error, while a PID-based liveness check still reports the process as live. A wedged watcher can therefore look healthy. Cron remains the backstop.</li> <li>Wake notifications are subject to the same ~19-20 event budget as any monitor; past that, low-latency wakes stop and only cron remains.</li> </ul> <h2 id="0671">0.67.1<a class="headerlink" href="#0671" title="Permanent link">¶</a></h2> <p>Fixes the DevSwarm supervisor escalation path for a Primary that has self-registered from the true main worktree — it had never delivered end-to-end. A Primary that has not self-registered that way still has its escalation land in the orphans list: the informational parent-inbox hook surfaces orphans, but the blocking parent-gate hook does not, so such a Primary can stop unblocked on an escalation the informational hook would have shown it. That gap is not fixed in this release.</p> <ul> <li><strong>Supervisor escalations now actually reach the parent.</strong> Four stacked defects, any one of which alone would have silently swallowed an escalation: the projection was never refreshed after an escalation was appended, so the parent-facing view stayed stale; the store was opened without a hash and wrote to a legacy bucket instead of the repoKey store; the parent id was derived from the child's own worktree path rather than resolved to the true parent, so escalations landed in the child's own bucket; and two fold/rehome paths skipped their projection refresh entirely on specific branches.</li> <li><strong>Hardened the roster's native fold against cross-repo env hijack.</strong> <code>hivecontrol workspace list children</code> resolves its scope entirely from <code>DEVSWARM_REPO_ID</code>/<code>DEVSWARM_BUILDER_ID</code>, never from cwd — a process holding a foreign repo's env gets that repo's children back with exit 0 and valid JSON. Each record's <code>repositoryId</code> is now cross-checked against a separate cwd-anchored, env-stripped <code>list all</code> lookup; mismatches are dropped and logged. Fails open unfiltered (older hivecontrol without <code>repositoryId</code>, no ground truth, spawn error) — the fold is never hard-failed.</li> <li><strong>Added <code>devswarm.js skip <guard> [--ttl <minutes>]</code></strong>, the CLI entry point edit-guard's own block message already pointed agents at but which did not exist — agents were hand-rolling <code>node -e</code> scripts to write <code>~/.anti-hall/skip.json</code> directly. Also fixed edit-guard's block message itself, which titled its reason "DEVSWARM EDIT-DELEGATION RULE" while the skip key it actually checks is <code>edit-guard</code>, so following the message's own wording wrote a useless key and got blocked anyway.</li> <li><strong>Guarded the <code>devswarm-repokey.js</code> require</strong> in <code>recovery.js</code> — it is loaded top-level by both <code>devswarm-supervisor.js</code>'s sweep loop and <code>devswarm.js</code>'s CLI entry, neither of which wraps its own top-level requires in the fail-open guarantees that cover their call sites. A throwing require (corrupt or deleted <code>devswarm-repokey.js</code>) would have crashed both consumers before any fail-open path could engage.</li> </ul> <h2 id="0670">0.67.0<a class="headerlink" href="#0670" title="Permanent link">¶</a></h2> <p>DevSwarm workspaces get human-readable names, a lost review seat can no longer report a clean gate, and model routing goes version-agnostic.</p> <ul> <li><strong>DevSwarm workspaces are now shown by human-readable name instead of raw UUID.</strong> The parent orchestrator previously surfaced workspaces as bare UUIDs (e.g. "archive 13531615…"), which meant nothing to a human reader. hivecontrol already exposed a free-text <code>label</code>, but <code>parseChildrenList</code> dropped it, mapping only <code>{branch, id, path}</code>. <code>devswarm.js</code> now sets a title after a successful <code>hivecontrol workspace create</code>, via a SEPARATE best-effort <code>hivecontrol workspace update-title -b <branch> "<title>"</code> call — derived from the <code>-p</code> brief (first non-empty line, one leading markdown marker stripped, whitespace collapsed, 60-char word-boundary truncation). A caller-supplied <code>-t/--title</code> is never overridden. New shared module <code>plugins/anti-hall/companion/lib/devswarm-names.js</code> owns an fs name cache (atomic tmp+rename; reads fail open). The parent-inbox hook renders <code>name (shortid)</code>, reading ONLY the fs projection — it never spawns hivecontrol, preserving its per-turn hot-path contract. <code>spawn</code> remains a strict thin pass-through: the argv forwarded to <code>hivecontrol workspace create</code> is untouched, so future hivecontrol flags keep working without anti-hall changes. <code>reconcile</code> caches whatever label hivecontrol actually has for pre-existing workspaces — it deliberately does NOT invent titles for workspaces with no brief on record, since an ungrounded label would be written into a real user-facing field.</li> <li><strong>Review-seat integrity: a lost seat can no longer report a clean gate.</strong> <code>ship-it</code>'s per-phase gate previously ignored dead seats entirely, so FEWER live seats produced FEWER findings and therefore <code>converged: true</code> — missing review coverage actively produced a PASSING gate. The gate now carries <code>totalSeats</code>, <code>liveSeats</code>, <code>deadSeats</code>, <code>degraded</code>, and <code>seatReports</code>, and <code>converged</code> requires <code>deadSeats === 0</code>. <strong>Behavior change:</strong> a phase that loses a seat will now correctly fail to converge where it previously passed silently. <code>ship-it</code> now also honors <code>args.codexAvailable</code> (mirroring <code>deadly-loop</code>, including the Opus adversarial-persona fallback). <code>codex-availability.js</code> now instructs the coordinator to thread its result into <code>ship-it</code>/<code>deadly-loop</code> Workflow invocations, the same way <code>fable-availability.js</code> already does — workflow scripts have no filesystem access and cannot read the JSON themselves. Corrected three docs that claimed an "enforced codexUp probe" gating the Critic seat: no probe existed, it is a caller-supplied flag that fail-opens to true, and a null Codex spawn remains the real backstop.</li> <li><strong>Model routing is now version-agnostic.</strong> Removed pinned model versions ("Sonnet 5", "Fable 5", "Opus 4.8") across skills, hooks, statusline docs, READMEs, and the Codex port; routing goes by tier token (<code>opus</code>/<code>sonnet</code>/<code>haiku</code>/<code>fable</code>), which the harness resolves to the newest model in each family. Added a standing MODEL-POLICY rule: never pin a model version. The statusline doc's model-id table now shows family globs (<code>claude-opus-*</code> etc.) rather than pinned ids. One deliberate exception remains: <code>speculation-judge.js</code> calls the Anthropic Messages API directly, which requires an exact model id (there is no "-latest" alias) — it stays overridable via <code>ANTIHALL_JUDGE_MODEL</code>.</li> <li>Replaced a raw NUL byte in <code>scripts/devswarm.js</code> (offset 81252) with the <code>\x00</code> escape. It was a deliberate collision-proof sentinel key, but as a literal byte it made <code>grep</code> treat the 245KB file as binary. Runtime string unchanged.</li> <li>Fixed a <code>hasFlag</code> redeclaration collision in <code>devswarm.js</code> where a new helper silently shadowed the pre-existing one and broke <code>--yes</code>/<code>--confirm</code> detection across <code>reconcile-active</code> and <code>reap-stale</code>.</li> </ul> <h2 id="0661">0.66.1<a class="headerlink" href="#0661" title="Permanent link">¶</a></h2> <p>CI-green patch for v0.66.0 — a genuine Windows bug in the reconcile spawn path, no behavior change on POSIX.</p> <ul> <li><strong>Fixed a real Windows bug in <code>devswarm reconcile</code>'s spawned pull.</strong> The reconcile sweep spawns <code>inbox pull</code> as a subprocess and threads the caller's <code>home</code> through by setting only the <code>HOME</code> env var. Node's <code>os.homedir()</code> does not read <code>HOME</code> on win32 (it reads <code>USERPROFILE</code>), so on Windows that subprocess silently fell back to the real OS home directory instead of the one the caller intended — breaking the guarantee that the spawned pull observes the same devswarm root (including the same per-id lock) as its caller whenever the two differ. Now sets <code>USERPROFILE</code> alongside <code>HOME</code>, matching the precedent already used elsewhere in this codebase for the identical reason.</li> <li>Two supervisor reconcile-sweep tests spawn the real entry script with an overridden <code>HOME</code> to isolate themselves; they had the same one-sided <code>HOME</code>-only gap, so on Windows they silently operated against the real OS home directory instead of their own isolated fixture — observing stale state (or none) rather than what the test set up. Fixed the same way, in the tests themselves.</li> <li>Swept the rest of v0.66.0's new/changed tests for the same class (binary spawns, cwd-based project resolution, path-separator assumptions); no other instances found.</li> </ul> <h2 id="0660">0.66.0<a class="headerlink" href="#0660" title="Permanent link">¶</a></h2> <p>Closes a family of failure classes found by design review rather than by tripping over them: code that reported success it had not observed, and recovery that only ran if something else happened to trigger it.</p> <ul> <li><strong>A destructive read can no longer lose messages silently.</strong> The native queue is consume-on-read, so a monitor batch that arrived but failed to parse — a shape change, stderr contamination, a timeout truncating the JSON mid-print — was gone with nothing logged. Such a batch is now logged and quarantined to disk. A well-formed empty result is still treated as normal, so an idle poll does not create an error storm.</li> <li><strong>Health is no longer asserted from a weaker second definition.</strong> Doctor carried its own daemon-health check that omitted the pid guard and the monitor-fault check, and used it to reap a unit as "confirmed running and healthy" — which meant a daemon that was alive but ingesting nothing could authorize reaping the only real drainer. There is now one definition, used by every consumer.</li> <li><strong>Project identity resolves properly, or refuses.</strong> Identity was derived from the working directory, so from inside a submodule it keyed off the submodule, and from a non-git directory it fell back to a legacy store and reported live workspaces as unregistered — a confident wrong answer with no error. Submodules now resolve via the superproject, and an unresolvable context refuses instead of quietly reading somewhere else.</li> <li><strong>Success no longer hides a nested failure.</strong> <code>heartbeat</code> returned ok while its mesh broadcast had failed; reconcile's aggregate reported success while individual targets had crashed or timed out; several paths returned ok from a caught exception. These now reflect what actually happened, with a genuinely absent hivecontrol treated as a known-benign skip rather than a failure.</li> <li><strong>Recovery runs on its own.</strong> Stranded messages previously sat until an update or an explicit repair happened to invoke reconcile. A cooldown-gated sweep now runs on the existing supervisor, using the same single-consumer lock as the drains so it cannot race a live one.</li> <li><strong>Smaller repairs.</strong> <code>devswarm logs</code> and <code>doctor --logs</code> now read rotated history, so the highest-volume period of an incident is no longer the part that is missing; the log rotation lock records its owner instead of being stolen on age alone; the Primary's own unreadable inbox is surfaced instead of silently counting zero; the singleton supervisor unit now carries the same resolved binary path as the per-project units, which is why its subprocesses kept failing after those were fixed.</li> </ul> <p>2382 tests, 0 fail. Verified with the tooling binary both present and absent, since several of these paths behave differently when it cannot be found.</p> <h2 id="0650">0.65.0<a class="headerlink" href="#0650" title="Permanent link">¶</a></h2> <p>DevSwarm daemon reliability — the ingest daemon now recovers itself, reports honestly when it cannot, and children never block the swarm on a question.</p> <ul> <li><strong>Root cause fixed: the ingest daemon was alive but ingesting nothing.</strong> It spawned <code>hivecontrol</code> by bare name while its service manager supplied only a minimal PATH, so every monitor cycle failed with ENOENT — invisibly, because the error was swallowed. The binary is now discovered at install time and baked into the generated launchd/systemd/cron unit (never a hardcoded path), and the daemon resolves it from an explicit option, the <code>ANTIHALL_DEVSWARM_HIVECONTROL</code> env var, or PATH. Found by the error logging shipped in v0.64.0.</li> <li><strong>Permanent faults no longer storm.</strong> ENOENT/EACCES/ENOTDIR are configuration faults, not transient ones: they now escalate through a capped backoff and log on state transitions plus a periodic rollup instead of once per retry. The backoff is sliced so the heartbeat keeps being written — a backing-off daemon is never mistaken for a dead one.</li> <li><strong>Stale locks self-heal.</strong> Orphaned ingest locks are swept on daemon start, a recycled pid no longer blocks restart forever, and a zombie holder is reclaimed without signalling anything. Every removal requires positive OS confirmation of death, reuse, or defunct state; an inconclusive read never authorizes removal, and unknown holder states block by default rather than falling through.</li> <li><strong><code>doctor --reclaim-ingest-lock</code></strong> — an explicit, opt-in path to sweep orphaned locks and reclaim a contended ingest lock, then reinstall. Never runs automatically.</li> <li><strong>Health stops lying.</strong> The heartbeat now records the monitor outcome, so a daemon that is alive but failing every cycle is reported as a FAILURE by <code>doctor</code> and surfaced in-session by a one-line banner, instead of reading as RUNNING. Heartbeats written by older daemons lack these fields and are treated as unknown, never as a fault.</li> <li><strong>Reaper detection.</strong> Installing now detects a memory-guard/reaper script that would kill the daemon (a service-managed daemon is parented to init, so orphan sweeps target it) and reports the file and the exact allowlist entry to add. Detect-and-report only — it never edits anything outside the repo.</li> <li><strong>Children never block the swarm on a question.</strong> A child now forwards a decision to its parent with its options, its recommendation and the default it will take, keeps working every other item, and proceeds on that stated default if unanswered — flagging the assumption loudly. Only an unauthorized destructive action is a hard stop. The ladder is child -> parent -> human, never child -> human. Injected into every workspace session, and documented in the devswarm skill on both ports.</li> </ul> <p>2329 tests, 0 fail. Minor bump — daemon self-heal, honest health reporting, one new opt-in doctor flag; all new paths are fail-open and no path removes registry or message data.</p> <h2 id="0640">0.64.0<a class="headerlink" href="#0640" title="Permanent link">¶</a></h2> <p>DevSwarm self-heal reliability + observability, plus an edit-guard plan-mode fix.</p> <ul> <li><strong>Reconcile self-heal now works on the first pass.</strong> <code>rehomeMiskeyedRow</code> normalizes a genuinely mis-keyed row's stored <code>worktree_path</code> to the descriptor's verified current path before rehoming, so a row stranded in a legacy bare-hash store no longer no-ops on the first <code>healRegistry</code> pass. No-delete, idempotent, fail-open; regression-tested for single-pass rehome + zero message loss.</li> <li><strong>Doctor daemon-liveness gate.</strong> <code>doctor --fix</code> no longer reports the ingest daemon "healthy" from install-shape alone — it checks the two-signal liveness primitive (fresh heartbeat + live-pid lock) and, when install-ok-but-dead, takes the reinstall path with a distinct dead-daemon reason instead of masking it. Win32's documented no-op is respected.</li> <li><strong>Structured error logging wired in.</strong> The ingest/lock/send/reconcile/inbox/register error paths (including a previously-swallowed top-level catch) now emit to the central JSONL logger, and two read-only surfaces expose it: <code>devswarm.js logs</code> (filter by --repo/--component/--min-level/--since/--limit) and <code>doctor --logs</code> — so a Primary can analyze a child project's recent failures from one place.</li> <li><strong><code>inbox messages --ack-as-owner</code> UX guard.</strong> Passing <code>--ack-as-owner</code> without <code>--ack</code> now warns clearly that it did NOT ack (and points to <code>read-primary … --ack-as-owner</code>) instead of silently staying read-only.</li> <li><strong>Child ack-path e2e coverage.</strong> Added the missing test: a child with a distinct <code>DEVSWARM_BUILDER_ID</code> self-registers via <code>inbox pull</code> then acks via <code>read-primary --ack-as-owner</code>.</li> <li><strong>edit-guard plan-mode false-positive fixed.</strong> The coordinator edit-guard no longer blocks the orchestrator's own plan/scratch/handoff/memory writes (added <code>plan.md</code>, <code>*.continue-here.md</code>, and the out-of-cwd Claude memory store to the allowlist) and exempts plan mode — while STILL blocking undelegated source-file edits in any mode (symlink-honesty preserved).</li> </ul> <p>2224 tests, 0 fail. Minor bump — self-heal reliability + observability + a guard-UX fix; all new paths are fail-open and control-flow-neutral.</p> <h2 id="0631">0.63.1<a class="headerlink" href="#0631" title="Permanent link">¶</a></h2> <p>CI-green patch for v0.63.0 — two test-only fixes; no production behavior change.</p> <ul> <li><strong>Hermetic migration-gate test.</strong> <code>update-skill.test.js</code>'s "healRegistry runs inside a DevSwarm session" test no longer depends on the ambient machine having real <code>~/.anti-hall/devswarm/store/</code> entries — it now seeds an isolated temp <code>home</code> (matching its sibling tests), so it exercises the gate deterministically on a clean checkout instead of passing by accident of local state.</li> <li><strong>Windows skip for the daemonHealth-healthy H4 case.</strong> On Windows the ingest daemon is a documented no-op (<code>daemonHealth()</code> returns <code>unsupported</code>), so the "healthy daemon → don't refresh summary" assertion is structurally unreachable there; the test now carries the same <code>{ skip: win32 }</code> guard its sibling daemonHealth tests already use. Windows behavior is intentionally unchanged (the per-turn <code>deriveSummary</code> fallback keeps the Windows parent inbox fresh).</li> </ul> <p>Verified under a clean isolated HOME with DEVSWARM_* unset (mimicking a fresh CI checkout): 2176 pass, 0 fail, 2 skip.</p> <h2 id="0630">0.63.0<a class="headerlink" href="#0630" title="Permanent link">¶</a></h2> <p>DevSwarm mesh usability + self-healing hardening — addresses real parent/child coordination footguns found running a Primary + child workspace, plus new adaptive/self-healing infrastructure.</p> <ul> <li><strong>Send addressing: <code>send --to</code> now accepts the roster <code>id</code>, not only the internal meshId.</strong> <code>resolveSendTarget</code> tries the worktree-derived meshId first (full back-compat) then falls back to an exact registry-<code>id</code> match, with an <code>ambiguous-recipient</code> guard that fails closed rather than silently picking. <code>roster</code> now surfaces each row's <code>meshId</code>.</li> <li><strong>Reconcile self-heal: mis-keyed registry rows are healed, not silently rejected.</strong> A row whose stored path drifted from its descriptor's real worktreePath is corrected in place; a row physically in the wrong store is rehomed (no-delete, message-preserving, idempotent) via a <code>healRegistry</code> pre-pass in <code>cmdReconcile</code>. The aggregate <code>ok</code> now requires <code>rejected===0</code>, so a per-row ownership rejection can no longer be masked as success.</li> <li><strong>Self-healing migration in doctor + update.</strong> <code>doctor --repair</code> and the updater run the registry heal sweep (idempotent, fail-open, no-delete, skips malformed stores) so a broken store is repaired on update/health-check.</li> <li><strong>Ingest daemon self-heal.</strong> A wedged-but-alive ingest daemon (blocked event loop, SIGTERM undeliverable) whose own liveness heartbeat is confirmed stale is SIGKILLed and its lock reclaimed — never a fresh-heartbeat daemon (fails toward never-kill on any inconclusive read). The Primary's roster projection refreshes itself (<code>deriveSummary</code>) when the daemon is unhealthy, with a non-destructive empty-store guard.</li> <li><strong>Structured error-logging foundation.</strong> New central JSONL logger (<code>companion/lib/anti-hall-log.js</code>) — fail-open, size-bounded/rotating — wired into ingest/lock/parent-inbox error paths; previously-swallowed catches now log. (Cross-component wiring + an analyze-from-here <code>--logs</code> CLI land in a follow-up.)</li> <li><strong>Parent/child SKILL clarity.</strong> The devswarm skill (Claude + Codex mirror) now spells out addressing, identity (<code>inbox read-primary</code> vs <code>$DEVSWARM_BUILDER_ID</code>), the mutating-ack path, and an edge-case→remedy table.</li> </ul> <p>2176 tests (up from 2130 in v0.62.2), 0 fail. Minor bump — new self-healing capability + logging foundation; no destructive paths (all heals are no-delete, message-preserving, idempotent, fail-open).</p> <h2 id="0622">0.62.2<a class="headerlink" href="#0622" title="Permanent link">¶</a></h2> <p>DevSwarm store deep hardening — closes a third unlocked registry writer plus consistency/robustness gaps found across multiple review rounds.</p> <ul> <li><strong>Concurrency: third unlocked registry writer now locked.</strong> <code>registerStoreDescriptor</code> now runs under the same per-id lock as <code>upsertRegistry</code>/<code>rekeySubdirRegistryRows</code>, closing the last unserialized write path into the registry.</li> <li><strong>Consistency: <code>upsertRegistry</code> path-change guard.</strong> A same-id write that would silently change <code>worktree_path</code> (sqlite and journal backends alike) is now rejected by default with an explicit <code>true</code>/<code>false</code> return so callers can detect a skip, instead of silently clobbering a different worktree's registration; opt-in via <code>allowPathChange</code>.</li> <li><strong>Robustness: truncated/malformed <code>DEVSWARM_BUILDER_ID</code> defense.</strong> <code>registerChildDescriptor</code> no longer writes a phantom row under a bad id — it recovers the full id when possible, and falls back to non-destructive same-worktree phantom retirement (truncation-signal only) that forwards any unread direct messages to the survivor before archiving.</li> <li><strong>Caller fixes for legitimate path changes.</strong> <code>cmdRegister</code> keeps descriptor and registry consistent on owner re-register; <code>rehomeCore</code> now compares <code>worktree_path</code> + <code>sessionId</code> before tombstoning the source, preventing an orphaned row; migrate verification no longer reports success when the underlying registry write was silently skipped by the new path-change guard.</li> </ul> <p>2130 tests (up from 2117 in v0.62.1). Behavioral fix/hardening release (patch bump) — no new hooks, skills, or disciplines.</p> <h2 id="0621">0.62.1<a class="headerlink" href="#0621" title="Permanent link">¶</a></h2> <p>Follow-up hardening: <code>rekeySubdirRegistryRows</code> now serialized under the per-id lock with an in-lock re-read (closes a lost-update race vs concurrent register/ensure/heartbeat/re-home in doctor's fold path); <code>cmdSpawn</code> confirmed race-free (documented). Docs: devswarm verb tables (KB §8.8, READMEs, llms.txt) updated with diagnose/healthcheck/unarchive/migrate-owner-keys/reap-stale/reconcile-active.</p> <h2 id="0620">0.62.0<a class="headerlink" href="#0620" title="Permanent link">¶</a></h2> <p>DevSwarm Primary-orchestrator lifecycle: heals the split-brain "no primary set" failure mode, adds parent-driven archiving of abandoned/reconciled workspaces, and decouples heartbeat-alive from stale/escalated liveness verdicts.</p> <ul> <li><strong>Split-brain "no primary set" healed via store re-home.</strong> A Primary could register its descriptor under a stale hash-keyed store bucket while the fresh <code>repoKey</code> resolution pointed elsewhere — the roster then read as having no Primary at all, even though one was live. <code>rehomeCore</code> migrates the registry row + backlog + cursor into the correct <code>repoKey</code>-keyed bucket; it's invoked from the read path (<code>inbox messages</code>/<code>read-primary</code>), <code>register</code>/<code>ensure</code>, and the new <code>migrate-owner-keys</code> forward-migration below, so a stranded Primary self-heals on its own next read or register rather than staying invisible.</li> <li><strong>Parent-driven archiving: <code>reap-stale</code> and <code>reconcile-active</code> (new CLI verbs).</strong> <code>reap-stale [--yes|--confirm]</code> scopes to this project's descriptors verdicted <code>stale</code>/<code>escalated</code> and archives the survivors after two hard safety gates (a fresh heartbeat, or recent worktree git activity, both mean never-reap) — dry-run by default. <code>reconcile-active [--active id,...] [--allow-empty] [--stdin] [--yes|--confirm]</code> archives every current workspace NOT named in an explicit "still active" set (e.g. from a roster screenshot); a match always spares a workspace, and an empty active set is refused unless <code>--allow-empty</code> is passed explicitly. Both reuse <code>cmdArchive</code>'s own pre-archive revalidation so a workspace that heartbeats or commits between listing and archiving is skipped, never wrong-archived.</li> <li><strong><code>unarchive <id></code> (new CLI verb).</strong> Reverses <code>archive</code> — restores a descriptor from <code>archived/</code> back to active (crash-safe hardlink-then-unlink) and revives the store registry row. Rejects if the archived descriptor's ownerKey doesn't match the current project (cross-project reject) or a conflicting active descriptor already exists.</li> <li><strong>Heartbeat→active reactivation.</strong> <code>devswarm-parent-gate</code> (the Stop-hook neglect gate) now suppresses a <code>stale</code>/<code>escalated</code> verdict from gating when the workspace has a FRESH heartbeat — proof the environment is alive, emitted only by the workspace's own live session — while leaving the real-unread coordination axis untouched (a live, heartbeating workspace with genuine unread backlog still gates; that's neglect, not staleness).</li> <li><strong>Hardening: per-id lock now fail-closed.</strong> The shared per-id lock guarding register/archive/unarchive/reap/re-home against concurrent interleaving now fails CLOSED (refuses the operation) rather than open, on the class of workspace-mutation verbs this release adds.</li> <li><strong>Hardening: crash-safe archive recovery-intent with fingerprint verification.</strong> <code>cmdArchive</code>'s recovery-intent marker (left when a prior archive tombstoned the registry row but its rollback/clear didn't complete) is now fingerprint-checked before a doctor/next-run companion pass acts on it, closing a window where a stale marker could misfire against an unrelated descriptor.</li> <li><strong>Hardening: cross-project reject.</strong> <code>unarchive</code> and the re-home path both refuse to act across project boundaries — an archived/stranded descriptor whose ownerKey doesn't match the current project's repoKey is left alone instead of silently adopted.</li> <li><strong>Hardening: future-heartbeat guard.</strong> A heartbeat timestamped in the future (clock skew or a forged write) is no longer treated as fresh/alive.</li> <li><strong><code>ownerKey</code> forward-migration (<code>migrate-owner-keys</code>, new CLI verb) in both <code>update.js</code> and <code>doctor</code>.</strong> Idempotent, fail-open, no-delete: scans every active + archived descriptor once, backfills a missing <code>ownerKey</code>, and re-homes an ACTIVE descriptor still stranded under a stale hash-keyed bucket into its fresh <code>repoKey</code>-keyed one (archived rows are never re-homed — only active ones can silently black-hole reads). Wired into both <code>update.js</code> (post-update) and <code>doctor</code>'s auto-safe repair, so most stores self-heal without an operator ever calling the verb by hand.</li> <li><strong>Read-only roster fallback.</strong> The roster projection falls back gracefully when a project's summary can't be resolved, instead of surfacing an empty/misleading table.</li> </ul> <p>2117 tests (up from 2037 in v0.61.0). This is a fix-and-capability release (minor bump) — no new hooks, skills, or disciplines; four new CLI verbs (<code>unarchive</code>, <code>migrate-owner-keys</code>, <code>reap-stale</code>, <code>reconcile-active</code>) on the existing <code>scripts/devswarm.js</code> DevSwarm CLI, shared identically by the Claude and Codex ports (<code>devswarm-parent-gate.js</code> is registered on both).</p> <h2 id="0611">0.61.1<a class="headerlink" href="#0611" title="Permanent link">¶</a></h2> <p>DevSwarm neglect-gate over-nag fix + registration inbox hardening.</p> <ul> <li><strong>Parent-gate (Stop hook) now counts only REAL unread.</strong> The DEVSWARM NEGLECT gate previously blocked on the raw unread line-count, so a "ghost" workspace whose entire backlog was the Primary's own <code>[Primary poke]</code> message mirrored back nagged on every Stop — a closed feedback loop (the nag prompts a poke, the poke refreshes the backlog). It now excludes system-generated poke/mirror noise via a shared classifier (<code>companion/lib/devswarm-noise.js</code> <code>isNoiseText</code>) and blocks iff realUnread > 0 or the workspace is stale/escalated. The message-age/freshness axis was dropped — a ghost's unread is <em>fresh</em> poke traffic, so freshness never excluded it. Fail-open: an unreadable inbox or an unparseable row still blocks.</li> <li><strong>Registration precreates an empty durable inbox.</strong> A freshly-registered child now reads as known/empty (0 unread, no nag) instead of absent (<code>known:false</code>), closing a false-silence hole where a genuinely-neglected child with an absent inbox was indistinguishable from a fresh one. The inbox is created with a truncation-proof append-mode open (<code>{flag:'a'}</code>) — it can never clobber a concurrent <code>devswarm-pull</code> drain on any filesystem, avoiding the O_EXCL/NFS unreliability of an earlier <code>wx</code> approach. Applied at both the per-turn hook registration path and the CLI register path; the descriptor is now published after the inbox/cursor init to close a transient discovery window.</li> </ul> <p>Behavioral fix to existing DevSwarm coordination — no new hooks, skills, or disciplines.</p> <h2 id="0610">0.61.0<a class="headerlink" href="#0610" title="Permanent link">¶</a></h2> <p><strong>DevSwarm mesh self-heal — instructions now reliably reach child workspaces; the substrate detects, routes around, migrates, and surfaces the mis-registration/misroute failure modes that could silently strand messages.</strong></p> <ul> <li><strong>Drain-aware routing.</strong> A <code>send</code>/mesh delivery now resolves to the partition a child is actually draining (not just the first matching registry row), plus a <strong>phantom-only rescue</strong>: on a child's very first mechanical self-register (<code>SessionStart</code>), any backlog that landed in a stranded pre-registration phantom partition is forwarded into the child's real partition — closing the class of bug where a message existed but no live session ever drained it.</li> <li><strong>Pure mesh-health projection.</strong> <code>computeSummary</code> (the store's pure summary projection) now derives <code>orphans[]</code> (message partitions with real unread backlog and no live workspace attached) and <code>staleRegistryPartitions[]</code> (a registry row whose worktree no longer exists on disk) — surface-only, never auto-deleted.</li> <li><strong>New read verbs.</strong> <code>diagnose</code> (read-only mesh-health detail: split/duplicate detection, orphans, stale partitions) and <code>healthcheck [--json]</code> (pass/fail over the same data, exit 0/2 — for monitors, CI, and the ingest daemon) join the CLI. <code>healthcheck</code> without <code>--json</code> prints one compact human-readable line.</li> <li><strong>Register-time dedup with a noise filter.</strong> A child re-registering from the same worktree now folds onto its existing partition instead of forking a duplicate, forwarding real backlog via a new <code>isForwardable</code> filter that skips poke/hash-mirror junk and forwards only genuine directs.</li> <li><strong>Parent-inbox surfacing.</strong> <code>devswarm-parent-inbox</code> now surfaces orphans and stale registry partitions to the Primary alongside the existing per-workspace status table — capped, read-only, never a delete or gate change.</li> <li><strong>Migration: <code>foldMeshDuplicates</code>.</strong> Folds every prior store shape (phantom rows, dual/legacy pairs, subdir-split registrations, stale entries) onto one canonical survivor per worktree, keyed by <strong>git-toplevel canonical identity</strong> (a child registered from a subdirectory now resolves to the same mesh identity as its toplevel). Idempotent, non-destructive (forward-before-tombstone; message rows are never deleted), fail-open. Wired into both <code>update.js</code> (runs post-update) and <code>doctor</code>'s auto-safe repair (dry-run detect doubles as a read-only mesh-shape check under <code>--check</code>, then applies).</li> <li><strong>Pure reads.</strong> <code>roster</code>, <code>workspaces list</code>, and <code>diagnose</code> no longer write <code>summary.json</code> as a side effect — they are now genuinely read-only.</li> <li><strong>Docs.</strong> The DevSwarm <code>SKILL.md</code> (Claude and Codex mirrors) gained a full "daemon + CLI ops reference" section — every CLI verb, the daemons, how to read mesh health without hand-reading the store, and self-heal behavior.</li> </ul> <p>This is a fix-and-capability release (minor bump).</p> <h2 id="0600">0.60.0<a class="headerlink" href="#0600" title="Permanent link">¶</a></h2> <p>Orchestration doctrine actually reaches the model now — it was silently spilling to a file — plus DevSwarm status idle-demotion.</p> <ul> <li><strong>The orchestration ruleset was reaching NO session inline.</strong> Hook <code>additionalContext</code> is capped at ~10,000 chars per hook command; the cap's failure mode is <strong>spill-to-file, not truncation</strong> (first-party documented) — a payload over the cap delivers only the first ~2,000 chars inline plus a file path the model must choose to open, so the tail effectively never lands. <code>verify-first-full.js</code> had grown to ~15,323 chars, so rules A–N (the orchestration/delegation ruleset) and the DevSwarm-Primary workspace-tier rule W reached no session inline.</li> <li><strong>Split into two SessionStart hooks.</strong> <code>verify-first-full.js</code> now carries the core Iron-Law + rationalization-table + Positive Rules + Scope & Fidelity + the disciplines/skills index (~7.7k chars); the new <code>verify-first-orch.js</code> carries the orchestration ruleset (rules A–N + the Primary-gated rule W, ~7.7-8.1k chars). Both clear the cap; both now land 100% inline; zero content dropped. Registered on both the Claude plugin and the Codex port.</li> <li>Reverted two prior cap-driven micro-optimizations that turned out to be cargo-cult, since they optimized position within a truncation window that never actually existed: rule L restored to its natural alphabetical slot, rule W restored to its full content (workspace = top fan-out tier, the spawn command, the choice rule, the failure mode).</li> <li>Added <code>tests/hooks/injection-cap.test.js</code> — runs every context-injecting hook and asserts each emitted payload is <code><=10,000</code> chars, so a future re-spill regression fails CI instead of silently degrading doctrine delivery.</li> <li><strong>DevSwarm status table: idle demotion.</strong> A workspace idle beyond <code>ANTIHALL_DEVSWARM_IDLE_MS</code> (default 6h) is now relabeled <code>active</code>→<code>idle</code> in the parent-inbox live status table — view-only (no delete, no gate change, no row removal), so a workspace that finished hours ago stops reading as "active" forever. Never overrides <code>escalated</code>/<code>stale</code>/<code>archive-ready</code>.</li> <li>Corrected <code>docs/KB.md</code>/<code>docs/KB-claude-codex.md</code> (the injection-cap behavior is spill-to-file with a first-party citation, replacing an earlier inaccurate plain-truncation claim and a circular self-citation) and fixed stale doc references to rule W now that it lives in <code>verify-first-orch.js</code>.</li> <li><strong>Synced the Codex-port manifest.</strong> <code>plugins/anti-hall/.codex-plugin/plugin.json</code> had silently frozen at <code>0.52.0</code> for 10+ releases because it was missing from the release checklist — bumped to match, and <code>RELEASING.md</code> step 1 now bumps both manifests together so this can't recur.</li> </ul> <h2 id="0590">0.59.0<a class="headerlink" href="#0590" title="Permanent link">¶</a></h2> <p>DevSwarm workspace-tier orchestration doctrine, idle self-wake via <code>CronCreate</code>, and a P0 edit-guard symlink-bypass fix.</p> <ul> <li><strong>DevSwarm workspace-tier doctrine.</strong> A DevSwarm Primary is now proactively directed that its top fan-out tier is a CHILD WORKSPACE (<code>devswarm.js spawn <branch> -p "<brief>"</code>), not a subagent; <code>edit-guard</code>/<code>command-guard</code> redirect the Primary there instead of naming "spawn a subagent" at the exact point it's blocked from working. Injected doctrine only — no mechanical scale classifier. A DevSwarm CHILD and any non-DevSwarm session are byte-identical to before.</li> <li><strong>DevSwarm idle-wake (<code>CronCreate</code>).</strong> A SessionStart directive tells a workspace to self-schedule a recurring mailbox-drain via <code>CronCreate</code> (the only primitive that fires while the REPL is idle); default <code>*/5 * * * *</code>, tunable via <code>ANTIHALL_DEVSWARM_WAKE_CRON</code>. A bounded Stop-gate re-verify on both <code>devswarm-child-gate</code> and <code>devswarm-parent-gate</code> handles the 7-day cron auto-expiry. Claude-only (gated on <code>DEVSWARM_AI_AGENT=claude</code>; Codex is never told to call <code>CronCreate</code>).</li> <li><strong>SECURITY (P0): edit-guard symlink bypass fixed.</strong> The coordinator-artifact allowlist (<code>PLAN.md</code>/<code>STATE.json</code>, and new <code>CONTINUE-HERE.md</code>) matched by path string and could be pointed at a symlink to write through to an arbitrary file. Now rejects any symlinked target or traversed symlinked directory (realpath cross-check, win32-aware); a non-existent file (first <code>Write</code>) is still allowed. This closes a pre-existing hole that also affected <code>PLAN.md</code>/<code>STATE.json</code>.</li> <li><strong>edit-guard: <code>CONTINUE-HERE.md</code> allowlisted.</strong> The session handover is coordinator-authored by design; added to the same root-anchored allowlist (traversal-safe).</li> <li><strong>Hardening.</strong> <code>ANTIHALL_DEVSWARM_WAKE_CRON</code> is now validated per-field against a strict cron charset (blocks prompt-injection via the env var into model-visible directive text); the new devswarm-wake lib is lazy-required + try/caught in all three hooks so a packaging failure fails open instead of crashing SessionStart/Stop.</li> <li><strong>Docs/KB.</strong> Hivecontrol reference KB updated to v2.3.5 (adds <code>jira</code> and <code>team</code> command groups, the hidden <code>workspace search</code> verb, <code>DEVSWARM_NO_AUTO_AUTH</code>); corrected stale "unreleased / still 0.56.0" version claims across llms.txt/README/skills that contradicted the actual shipped version; replaced real captured UUIDs/ports/PIDs in the KB with placeholders (public-repo agnostic rule).</li> </ul> <h2 id="0582">0.58.2<a class="headerlink" href="#0582" title="Permanent link">¶</a></h2> <p>DevSwarm hook parity with Codex, and a documentation correction.</p> <ul> <li>The five DevSwarm hooks (<code>devswarm-child-role</code> on SessionStart, <code>devswarm-parent-inbox</code> + <code>devswarm-child-turn</code> on UserPromptSubmit, <code>devswarm-parent-gate</code> + <code>devswarm-child-gate</code> on Stop) are now registered for the Codex port, using the same shared hook files Claude uses — zero changes to the hooks themselves, pure registration.</li> <li>Corrects a false claim that had propagated through the docs unverified: these hooks were described as Claude-only because their <code>DEVSWARM_*</code> env gate supposedly only applied to <code>claude</code> sessions. It does not — hivecontrol sets <code>DEVSWARM_REPO_ID</code>/<code>DEVSWARM_SOURCE_BRANCH</code>/<code>DEVSWARM_BUILDER_ID</code> per workspace regardless of which agent runs there (<code>DEVSWARM_AI_AGENT</code> names the agent). The gap was wiring, not capability. A test that asserted the gap has been flipped into a parity test.</li> <li>Still Claude-only, and structurally so: the liveness supervisor and the on-demand <code>devswarm-recover</code> CLI, which identity-bind to <code>claude --resume</code> processes by argv.</li> </ul> <h2 id="0581">0.58.1<a class="headerlink" href="#0581" title="Permanent link">¶</a></h2> <p>Fixes a real-world ingest-daemon outage, a message-loss gap in its own fix, a slow lock recovery path, dishonest reconcile reporting, and removes a manual step. Raises the Node floor to 22.</p> <ul> <li>Ingest daemon could wedge indefinitely: <code>main()</code> never passed a hard timeout, so the <code>spawnSync</code> running <code>hivecontrol workspace monitor</code> had no OS-level kill timeout — only the cooperative <code>-t</code> flag handed to hivecontrol. A child that ignores its own <code>-t</code> blocks the daemon forever, including its heartbeat; <code>isAlive()</code> then correctly refuses to steal a lock whose holder is genuinely alive-but-wedged, so the daemon can never recover (observed: 4,319 consecutive lock refusals over ~15h). Now a hard timeout (cooperative timeout + 10s) backstops the spawn and surfaces as a retryable failure through the existing backoff path.</li> <li>Ingest daemon leaked its lock on signal: no SIGTERM/SIGINT handlers existed, so an OS stop (e.g. launchd restarting the unit) bypassed the <code>finally</code> that releases the lock, stranding it for the full stale window. Signals now release the lock before exit.</li> <li>Message loss fixed: the hard timeout above (and any other retryable failure) discarded the killed monitor's stdout outright — but <code>hivecontrol workspace monitor</code> is DESTRUCTIVE, popping messages off the native queue as it prints them, so whatever it had already drained before being killed was gone for good. <code>spawnSync</code> preserves stdout written before a kill; that output is now ingested unconditionally whenever non-empty, even when the attempt is otherwise treated as a retryable failure.</li> <li>Ingest lock: a lock whose holder pid is confirmed dead is now reclaimed immediately instead of waiting out the full 15-minute stale window. Signal handlers (previous bullet) are not what guarantees this recovery path — Node cannot dispatch a JS signal handler while the event loop is blocked inside <code>spawnSync</code>, which is where this daemon spends its entire risky window, so a daemon killed mid-spawn never gets a chance to run its own handler. Dead-holder-immediate-reclaim in the lock acquisition path is what actually closes the leaked-lock window for every other starter.</li> <li><code>reconcile</code> no longer reports success while losing messages: a per-worktree <code>lost</code> count (from a real native-queue shortfall, distinct from a benign <code>locked</code> contention skip) now propagates all the way up — including across the <code>inbox pull</code> subprocess boundary that previously dropped it — and the aggregate <code>reconcile</code> result is no longer <code>ok:true</code> when any worktree lost messages. <code>doctor</code> now reports <code>failed</code> (not <code>fixed</code>) with the loss count, and <code>update</code> reports a not-success detail with the loss count, instead of both silently reading a lossy reconcile as clean.</li> <li><code>reconcile</code> (drains stranded per-worktree native queues into the shared store) now runs AUTOMATICALLY — as a gated <code>doctor</code> repair and as a post-update step inside a DevSwarm session — instead of requiring a manual command. It remains idempotent, skips (never races) a worktree a live child is already draining, and honors <code>--dry-run</code>/<code>--check</code>. The manual verb still works.</li> <li><strong>Node 22 is now the minimum</strong> (was 18). Node 18 and 20 are past EOL, and Claude Code's own npm package already requires Node 22+. CI matrix is now Node 22/24 across ubuntu/macos/windows.</li> </ul> <h2 id="0580">0.58.0<a class="headerlink" href="#0580" title="Permanent link">¶</a></h2> <p><strong>DevSwarm mesh-only messaging: a per-project mesh store becomes the sole agent-initiated messaging transport, mechanically enforced. hivecontrol's per-worktree native messaging is replaced (per-worktree queues, no from/to addressing, no broadcast); every other hivecontrol feature (create/list/check-merge/merge) is kept and thinly wrapped. DevSwarm coordination remains entirely OPTIONAL — dormant with zero behavioral change outside a DevSwarm session.</strong></p> <ul> <li><strong>Mesh substrate (previously unreleased v0.57 work).</strong> Per-project <code>repoKey</code>-keyed shared store: every linked worktree of one repo shares one store, so any worktree can message any other (all-to-all), not just its own parent/child. Windows-hardened path canonicalization.</li> <li><strong>Mesh CLI.</strong> <code>send --to <meshId>|--to-primary|--broadcast [--urgency low|normal| high|urgent]</code>, <code>roster [--ack]</code>, <code>mesh read</code>, <code>heartbeat <id> --summary</code>. Rows carry <code>{from, to, type, message, timestamp, urgency}</code>. Spoof-resistant cwd-derived <code>from</code>; fail-closed <code>--to</code>.</li> <li><strong>#36 STRUCTURAL.</strong> <code>devswarm-parent-gate</code>/<code>devswarm-parent-inbox</code> now scope to the caller's OWN project via <code>repoKey</code>, replacing a spoofable env-var filter that leaked workspaces across projects.</li> <li><strong>Mesh-only messaging enforcement.</strong> <code>command-guard</code> blocks agent-initiated <code>hivecontrol message-child</code>/<code>message-parent</code> in all contexts and redirects to the mesh CLI. Matching dequotes to the shell-effective argv, closing quote-based bypasses (<code>"message-parent"</code>, split-quote, quoted verb) — which also closes the same pre-existing gap in the <code>monitor</code>/<code>read-messages</code> guards. Lifecycle commands are never blocked; coordination commands stay exempt from the heavy-command gate.</li> <li><strong>Per-turn communication override</strong> (both roles) re-asserting mesh-only messaging and a mesh-poll resting posture. No external mechanism can wake a truly idle Claude Code session (anthropics/claude-code#44380), so idle-wake is not claimed.</li> <li>Thin lifecycle wrappers <code>spawn</code>/<code>merge</code>; <code>roster</code> folds <code>hivecontrol list</code>.</li> <li><code>reconcile</code> drains stranded per-worktree native queues into the store.</li> <li><code>archive-request</code> is now a store write (zero native calls).</li> <li>Supervisor escalates (never kills) on urgent/high unread — direct or broadcast.</li> <li><code>child-gate</code> gains a projection-only already-reported check; heartbeat sender identity is ownership-validated, so a forged heartbeat can no longer satisfy another workspace's Stop-gate.</li> <li>Emitted coordination commands now use an absolute CLI path (they were unrunnable from an agent's working directory).</li> <li>SQLite writes retry past <code>busy_timeout</code> exhaustion under concurrent writers.</li> <li><strong>Codex parity.</strong> The guard-block is shared and fires on Codex; the per-turn override hooks remain Claude-only (full parity tracked as a follow-up).</li> </ul> <h2 id="0560">0.56.0<a class="headerlink" href="#0560" title="Permanent link">¶</a></h2> <p><strong>DevSwarm archive handshake (send-only, never mechanical on either side); ack-ownership guard is now cwd-ground-truth (closes an env-spoof cursor-corruption path); the Primary sees its own inbound because the ingest daemon self-registers its own store row; child descriptor registration + heartbeat-key fix close two discovery/collision bugs; doctor/update now heal a drifted ingest script, not just a missing one; the Fable Reviewer seat is back.</strong></p> <ul> <li><strong>Archive handshake (both roles).</strong> New <code>devswarm.js archive-request <childId|branch></code> verb — SEND-ONLY: posts a <code>[[ANTIHALL_ARCHIVE_REQUEST]]</code>-marked message to a done child; it never verifies merged/tested/deployed itself (the Primary checks per its OWN repo policy before sending). Child side: <code>devswarm-child-turn.js</code> scans unread for the marker and injects a distinct segment telling the child to confirm with ITS user, then run <code>devswarm.js archive <id></code>. anti-hall NEVER archives mechanically on either side.</li> <li><strong>Ack-ownership guard — cwd-ground-truth (P0-hardened).</strong> <code>devswarm.js</code>'s <code>callerIdentity(env, cwd)</code> now treats cwd as the source of truth: when cwd resolves to a real git worktree, identity is derived from that worktree and a <code>DEVSWARM_BUILDER_ID</code> env var naming a <em>different</em> workspace is ignored, never trusted to override — closing the env-spoof path where a workspace could set <code>DEVSWARM_BUILDER_ID=<other-id></code> to impersonate another workspace and ack its cursor (also closes the related env-inheritance case, e.g. a subshell/subagent inheriting a stale builder id). <code>DEVSWARM_BUILDER_ID</code> is honored only when it can't contradict cwd — cwd already agrees, or cwd resolves to no worktree at all. The explicit <code>--ack-as-owner</code> operator override for a legitimate cross-workspace ack (e.g. a supervisor clearing a dead workspace's backlog) is preserved.</li> <li><strong>Primary sees its OWN inbound.</strong> <code>devswarm-parent-inbox.js</code> + <code>devswarm-parent-gate.js</code>: the Primary now surfaces AND Stop-gates on its own unread parent/peer messages, with the same imperative "STOP and read FIRST" wording as the child gate. This actually surfaces end-to-end because the ingest daemon now <strong>self-registers</strong> its own <code>primary-<hash></code> workspace row at startup (merge-preserving — a prior explicit <code>register-primary</code> call is never clobbered): previously no runtime path ever created that row, so messages landed in the store but the workspace itself never existed for the summary projection to surface. The daemon's own identity is <strong>worktree-ground-truth</strong> — it registers/ingests under the <code>primary-<hash></code> id derived from ITS OWN resolved worktree, never an inherited child <code>DEVSWARM_BUILDER_ID</code> — closing the same class of identity-spoof/inheritance bug the ack-ownership guard closes above.</li> <li><strong>Child descriptor registration (#31) + imperative unread.</strong> <code>devswarm-child-turn.js</code> mechanically writes/refreshes the child's own descriptor every turn (merge-preserving) so the parent can always discover it; unread-parent-message wording escalated advisory → imperative. Child-gate STRICT mode (<code>ANTIHALL_DEVSWARM_CHILD_GATE_STRICT</code>, default on) backs the durable- inbox check with one bounded, non-destructive native <code>message-count</code> probe (fail-open).</li> <li><strong>Heartbeat key fix.</strong> <code>devswarm-child-turn.js</code>'s <code>heartbeatKey</code> now keys by <code>DEVSWARM_BUILDER_ID</code> (was the source branch) — sibling children no longer collide and parent-join resolves correctly.</li> <li><strong>Ingest daemon durability + doctor/update self-heal parity.</strong> <code>install-devswarm-ingest.js</code>'s <code>resolveStableScript</code> now bakes the daemon's script path to the git marketplace-clone path (updated in place by the updater) instead of the version-pinned cache dir that gets renamed on update — fixes an orphaned crash-loop after an update. The daemon now logs (startup / lock-refusal / ERROR+stack) to <code>~/.anti-hall/devswarm-ingest.log</code> (systemd <code>StandardOutput</code>/<code>StandardError</code> + cron redirect; previously discarded to <code>/dev/null</code>). <code>doctor --repair</code> and <code>update.js</code>'s <code>healIngestDaemon()</code> gained a NEW <code>unstable-script</code> classification (<code>classifyIngestUnit</code>) — the baked <code>ExecStart</code> script still exists but no longer matches the current stable path (config drift, not just an absent/wrong-path unit) — and both now detect + migrate it to the stable path, DevSwarm-session-gated, fail-open.</li> <li><strong>Migration <code>--mark-read</code>.</strong> <code>migrate-state.js</code> / <code>devswarm-migrate.js</code> gained an opt-in flag/env to advance a freshly-imported legacy backlog's cursor to already-seen (default OFF, byte-identical output when absent) so a catch-up migration doesn't trip the parent neglect-gate.</li> <li><strong>Fable seat re-enabled.</strong> Fable 5 is now available, so the earlier 0.43.2 policy-disable (over-restrictive/refusal-prone per community feedback) is reversed: the deadly-loop / ship-it Reviewer seat once again tries Fable first when <code>args.fableAvailable === true</code>, falling back to Sonnet 5 then Opus (fallback chain unchanged). Both Claude (<code>skills/deadly-loop/references/ deadly-loop.workflow.js</code>, <code>skills/ship-it/references/ship-it.workflow.js</code>) and the Codex port (<code>codex/skills/anti-hall-deadly-loop/SKILL.md</code>, <code>codex/skills/anti-hall-model-policy/SKILL.md</code>).</li> <li><strong>Doc correction.</strong> Removed the stale claim (across <code>docs/KB-devswarm-hivecontrol.md</code>, <code>README.md</code>, <code>llms.txt</code>, <code>docs/KB.md</code>) that <code>devswarm-child-gate</code> silences the Stop-gate on a fresh child heartbeat — that behavior was reverted in v0.54.1; the gate always demands a report, bounded only by its per-episode cap.</li> <li><strong>Codex parity.</strong> Mirror updates in <code>plugins/anti-hall/codex/skills/anti-hall-devswarm/SKILL.md</code> and the model-policy/deadly-loop Codex skills above.</li> <li><strong>Cross-project scoping (#36).</strong> DevSwarm workspace descriptors now record their <code>repoId</code> (<code>DEVSWARM_REPO_ID</code>), and <code>devswarm-parent-gate</code> / <code>devswarm-parent-inbox</code> filter their per-turn enumeration to the current session's project — a Primary in one project no longer sees, gates on, or is nagged by another project's workspaces. Fail-open + back-compat: an untagged descriptor (or a session with no <code>DEVSWARM_REPO_ID</code>) still surfaces exactly as before, and children re-stamp their <code>repoId</code> within one turn.</li> <li>Docs sweep: <code>README.md</code>, <code>llms.txt</code>, <code>docs/KB-devswarm-hivecontrol.md</code>, <code>docs/KB.md</code>, and <code>skills/devswarm/SKILL.md</code> / <code>skills/update/SKILL.md</code> updated to match.</li> <li><strong>Cross-platform CI fixes.</strong> <code>devswarm-child-gate</code>'s STRICT native <code>message-count</code> probe now resolves the <code>hivecontrol</code> shim on Windows (<code>spawnSync</code> <code>shell</code> on win32 → PATHEXT); the statusline base-command dispatcher no longer discards a successful base command (exit 0 + output) when writing our stdin to a non-reading child races to a benign broken-pipe error — judged errno-agnostically by exit status and stdout presence, since the OS reports it differently per platform (<code>EPIPE</code> on POSIX; <code>EOF</code> on Windows, where libuv maps <code>ERROR_BROKEN_PIPE</code> → <code>UV_EOF</code>). The hook-test helper (<code>tests/helpers/spawn-hook.js</code>) now gives cmd.exe its full Windows system env (<code>SystemRoot</code>/<code>ComSpec</code>/<code>PATHEXT</code>/etc, inherited from the parent) on win32 so the STRICT probe's shell spawn can actually start, while stripping every <code>DEVSWARM_*</code>/<code>ANTIHALL_*</code> key first to keep test isolation intact; the statusline's top-level stdout error handler now tolerates the same benign broken-pipe condition (<code>EOF</code> on Windows, <code>EPIPE</code> on POSIX) instead of rethrowing.</li> <li>Full suite: 1487 tests, 1485 pass, 2 skipped (Windows-gated doctor no-ops), 0 fail.</li> </ul> <h2 id="0550">0.55.0<a class="headerlink" href="#0550" title="Permanent link">¶</a></h2> <p><strong>DevSwarm Primary read path ships (the Primary can finally read its own inbox); per-project ingest daemon fixes multi-repo coverage; a mechanical single-consumer read-guard closes the raw-read bypass; the supervisor actively escalates into the parent's store; and <code>doctor</code> now DIAGNOSES then REPAIRS every aspect it safely can.</strong></p> <ul> <li><strong>DevSwarm Primary read path (NEW).</strong> <code>store.listMessages(id, {sinceCursor})</code> — implemented on BOTH backends (sqlite <code>devswarm-store.js</code>, journal/NDJSON, the latter deduping by hash on read for parity with <code>messageCount</code>) — is the read-BACK side of the store: message <code>body</code> was written but never readable back until now. <code>devswarm.js inbox messages <id></code> reads it non-destructively (<code>--unread</code> limits to the tail past a durable ACK cursor under <code>cursors/<id>.json</code>); <code>inbox read-primary <id></code> additionally ACKs — advancing both the cursor file and the store's own cursor (<code>store.setCursor</code>) so <code>deriveSummary</code>'s projected <code>unread</code> count (what the live workspace table shows) drops immediately, not just on the next unrelated store write. <code>register-primary</code> registers the CURRENT worktree's Primary/parent descriptor under its per-worktree <code>primary-<hash></code> id (reusing <code>register</code>'s validation + store-upsert path) so <code>migrate</code> can fold a stranded legacy NDJSON inbox into the store under the same id.</li> <li><strong>Per-project ingest daemon (Option A — fixes multi-repo coverage).</strong> The daemon previously covered only the ONE worktree it was installed from, silently — a second repo had no reception at all. <code>install-devswarm-ingest.js</code> now derives an <strong>additive, per-worktree</strong> unit identity: <code>worktreeHash(wt)</code> (an 8-hex SHA-256 of the worktree's realpath) feeds a per-worktree LaunchAgent label <code>com.anti-hall.devswarm-ingest.<hash></code> (macOS), systemd unit / cron marker <code>anti-hall-devswarm-ingest-<hash></code> (Linux), a per-worktree <code>O_EXCL</code> lock (<code>locks/ingest-<hash>.lock</code>), and <code>primaryWorkspaceId(wt)</code> = <code>primary-<hash></code> (the store partition key for that worktree's own reception queue — replacing an early hardcoded <code>'primary'</code> that collided rows across repos). Installing from a second repo now creates a NEW unit rather than overwriting the first's; <code>devswarm-ingest.js</code> computes the identical hash from its own resolved worktree so lock path and workspace id agree byte-for-byte with what the installer baked in. New <code>listInstalledIngestUnits()</code> enumerates every installed unit (legacy hash-less AND per-worktree, across launchd/systemd/cron) for readback — <code>doctor</code>/<code>doctor-repair</code> use this rather than re-deriving. The daemon writes a per-worktree LIVENESS HEARTBEAT (<code>heartbeats/ingest-<hash>.json</code>, atomic tmp+rename) every bounded sweep cycle regardless of whether anything was ingested — a live-but-quiet daemon no longer false-reads as stale. <strong>Multi-repo coverage means installing separately from each repo/worktree</strong> — there is still no single daemon that covers more than the worktree it was launched from.</li> <li><strong>Mechanical single-consumer read-guard.</strong> New PreToolUse-Read hook <code>hooks/inbox-read-guard.js</code> blocks a direct Read-tool read of the raw DevSwarm inbox NDJSON or store (sqlite db + sidecars, journal NDJSON) — the harm is cursor desync + a store-layering violation, not queue-draining (the inbox is append-only). <code>command-guard.js</code> gained a parallel Bash-side branch (<code>detectProtectedFileRead</code>, reusing the shared <code>devswarm-inbox-paths.js</code> classifier, with the same quote-neutralization + <code>bash -c</code>/<code>eval</code>/<code>$()</code> recursion as the existing hivecontrol detector) that blocks <code>cat</code>/<code>head</code>/<code>tail</code>/<code>grep</code>/<code>sed</code>/<code>awk</code>/etc. reads of those same paths — closing the shell-side bypass the Read-tool guard can't see. <code>hivecontrol workspace read-messages</code> now blocks UNCONDITIONALLY under DevSwarm, same as <code>monitor</code> (previously it only blocked with durable-inbox evidence present); both reasons now redirect to the <code>devswarm.js inbox pull</code>/<code>read</code> wrapper and document the <code>DISABLE_ANTIHALL_DEVSWARM=1</code> kill-switch. Both guards share the <code>devswarm-read-guard</code> skip name (immune to a blanket <code>{all}</code> skip, like git-guard) and are fully fail-open on any internal error.</li> <li><strong>Supervisor active escalation.</strong> <code>recovery.js</code>'s <code>pokeOrEscalate</code> — on the TRANSITION into <code>escalated</code> (nudge attempts exhausted) — now also mechanically appends a notice into the PARENT/Primary's store (<code>notifyParentEscalation</code>, resolving the parent's <code>primary-<hash></code> id via <code>install-devswarm-ingest.js</code>), so an idle Primary sees "child X idle Nm — reassign or archive" without the child taking a turn. Idempotent two ways: caller-gated to the actual verdict transition, plus a stable per-escalation content hash (<code>escalate:<childId>:<staleSince></code>) so a race between two sweep observers still lands one row. Fully fail-open — a store-write failure never crashes the sweep or blocks the already-persisted escalation.</li> <li><strong>Freshness banner</strong> in <code>devswarm-parent-inbox.js</code>, rewired from <code>summary.json</code>'s <code>generatedAt</code> (which only advances on <code>inserted>0</code> and false-read a live-but-quiet daemon as stale) to the daemon's own per-worktree heartbeat file. Resolves the current worktree via a pure <code>fs</code> walk-up for <code>.git</code> (no git spawn on this hot path) hashed the same way the installer did; when the heartbeat is missing or older than 3 minutes, a <code>⚠ DEVSWARM STALE DATA</code> banner is injected above the live workspace table. Gated on at least one active workspace so an idle system never false-alarms; fully fail-open.</li> <li><strong>doctor repair mode.</strong> <code>doctor</code> now DIAGNOSES then REPAIRS every aspect it safely can. New module <code>hooks/lib/doctor-repair.js</code> (mirrors the <code>companion/lib/doctor-devswarm.js</code> require-and-call pattern doctor already uses). Exports <code>runRepairs()</code> plus the testable <code>readInstalledIngestWorkingDir()</code> / <code>classifyIngestUnit()</code>.</li> <li><strong>Flags on <code>hooks/doctor.js</code>:</strong> <code>--check</code> = PURE read-only (today's behavior, the CI/test path — mutates nothing); <code>--dry-run</code> = detection + print what WOULD be fixed, writes nothing (threads each installer's own <code>--dry-run</code>, migrate-state <code>dryRun:true</code>); <code>--fix</code>/<code>--repair</code> = explicit aliases for the default auto-apply path; <strong>plain <code>doctor</code> (no flags) now auto-applies fixes.</strong> <code>--quiet</code> preserved.</li> <li><strong>Two safety classes.</strong> AUTO-SAFE (always, honors dry-run): legacy/GSD/DevSwarm-store state migration; statusline install <strong>only when no statusLine is configured in any scope</strong> (a custom statusLine is never overridden); idempotent relaunch of an ALREADY-installed supervisor; Codex hook refresh when a <code>.codex/config.toml</code> exists but hooks are unwired (never creates a new <code>.codex</code>). GATED (only when <code>isDevswarmActive(env)</code> AND <code>resolveWorktree(cwd)</code> is a git worktree): ingest daemon install, the <strong>v0.54.1 wrong-path ingest heal</strong> (a unit whose <code>WorkingDirectory</code> no longer resolves inside a worktree is rebuilt from the correct one), stale ExecStart script, and supervisor FIRST-install. A closed gate REPORTS the gap + the exact manual command and mutates nothing. REPORT-ONLY: the MCP orphan reaper (kills on a timer — never auto-installed).</li> <li><strong>Ingest-daemon awareness added to doctor</strong> (previously the ingest daemon was surfaced only via the supervisor path); doctor now reads the installed ingest unit back and classifies it <code>ok</code> / <code>wrong-path</code> / <code>stale-script</code> / <code>absent</code>.</li> <li><strong>doctor runtime health (new module <code>companion/lib/doctor-runtime.js</code>, five REPORT-ONLY checks, never mutates).</strong> 1) DB/store health — sqlite <code>PRAGMA quick_check</code> (run in a child process so its <code>node:sqlite</code> experimental warning never reaches doctor's own stderr) plus a journal-backend torn-line scan and a store<->summary parity check (flags only <code>summary.total > store.total</code>). 2) Data staleness — <code>summary.json</code>'s <code>generatedAt</code> vs the per-worktree ingest heartbeat, gated on the daemon actually running AND <code>unread > 0</code> so an idle system never false-alarms. 3) Daemons RUNNING, not just installed — <code>launchctl list</code> PID probe for the continuous ingest daemon, <code>launchctl</code>/<code>systemctl</code> loaded-state probe for the periodic supervisor, and a heartbeat-or-<code>ps</code>-fallback probe for the cron fallback — enumerated per-worktree via <code>listInstalledIngestUnits()</code>. 4) Second-consumer detection — scans the process table for more than one <code>hivecontrol workspace monitor</code> process (which would split the destructive native queue), cross-checked against the <code>locks/ingest*.lock</code> holder PID; report-only, never kills. 5) Foreign-plugin/hook conflict scan (UNCONDITIONAL, not gated on DevSwarm) — cross-references other ENABLED plugins' <code>hooks.json</code> and skill directories against anti-hall's own, flagging a competing <code>PreToolUse</code>-on-Bash or <code>Stop</code> hook, or a skill-name collision; privacy-scoped to plugin name + event + matcher + hook basename only, never full command strings or file contents.</li> <li><strong>Verify-after-fix:</strong> each applied fix RE-RUNS the relevant detection before reporting <code>FIXED</code> (a spawned installer's exit code is not trusted — <code>launchctl load</code> can warn). A <code>FAILED</code> repair keeps the exit code non-zero. Every fix is try/catch fail-open.</li> <li><strong>Backward-compat:</strong> CI gates only on <code>node --test</code> (no doctor invocation), and the existing <code>tests/hooks/doctor.test.js</code> read-only assertions were re-pointed to <code>--check</code>; new <code>tests/hooks/doctor-repair.test.js</code> covers the auto-fix / gate / dry-run behavior. Windows daemon fixes are documented no-ops.</li> <li><strong>Dual-platform (doctor):</strong> <code>skills/doctor/SKILL.md</code> + <code>codex/skills/anti-hall-doctor/SKILL.md</code> document the flags, classes, and gate (the Codex mirror notes the gate is effectively always closed for gpt-5.x sessions).</li> <li><strong>DevSwarm store is now PHYSICALLY PER-PROJECT.</strong> Each git worktree gets its own <code>store/<hash>/devswarm.db</code> (sqlite) or <code>store/<hash>/journal/</code> (journal backend), plus its own <code>summaries/<hash>.json</code> — deliberately outside <code>store/</code> so the inbox read-guard's <code>store/**</code> deny doesn't also swallow summaries (<code>companion/lib/devswarm-store.js</code>). <code><hash></code> is the worktree hash unwrapped from a <code>primary-<hash></code> id, or <code>sha256(id).slice(0,8)</code> for any other workspace id. Upgrading from the old single global store is handled by a NON-DESTRUCTIVE <code>migrateGlobalStoreToPerProject</code> (<code>companion/devswarm-migrate.js</code>): it reads whichever legacy backend(s) are actually present directly under <code>store/</code>, splits rows by <code>workspace_id</code> into each project's own per-hash store, and leaves the legacy global store byte-for-byte intact as a backup (never deleted); idempotent via content-hash dedupe so a re-run copies nothing new.</li> <li><strong>Docs sweep.</strong> <code>skills/devswarm/SKILL.md</code> + <code>codex/skills/anti-hall-devswarm/SKILL.md</code> expanded to a comprehensive CLI reference (<code>inbox messages</code>/<code>read-primary</code>/<code>register-primary</code>, the per-worktree ingest identity, the single-consumer read-guard, and the hook-vs-daemon design split: mechanical parent/child triggers fire per-event via hooks, the ingest/supervisor loops run on their own interval via daemons). <code>docs/KB-devswarm-hivecontrol.md</code> gained the largest update — the read path, per-worktree ingest identity (with an explicit correction that an installed daemon must never be described as "verified functioning" for the whole machine or for other repos — only for the worktree it drains), the read-guard, and active escalation. <code>docs/KB.md</code> version row, <code>README.md</code>, and <code>llms.txt</code> updated to match.</li> </ul> <h2 id="0542">0.54.2<a class="headerlink" href="#0542" title="Permanent link">¶</a></h2> <p><strong>Child reception ships (native queue → durable inbox); the ingest daemon actually functions (<code>WorkingDirectory</code> fix); lock hardening closes a torn-read double-consumer race.</strong></p> <ul> <li><strong>Child reception (SHIPPED — replaces v0.54.1's surfacing-only).</strong> <code>node scripts/devswarm.js inbox pull <id></code> (<code>companion/lib/devswarm-pull.js</code>, <code>pullOnce</code>) is the drain v0.54.1 left as an explicit follow-up: a bounded, guard-safe pull that folds a child's NATIVE parent→child queue into its durable inbox + store. Each pull auto-ensures the descriptor (idempotent — reuses <code>register</code>'s write path), then runs a non-destructive <code>message-count</code> gate FIRST (count <code>0</code> returns without ever calling <code>read-messages</code>); on count <code>>0</code>, ONE bounded <code>read-messages</code> (finite 10 s timeout — never the hanging <code>monitor</code>); appends the batch to the durable inbox in one atomic NDJSON <code>appendFileSync</code>, idempotent by embedded content hash (reused verbatim from the ingest daemon so both paths dedupe identically); a per-id <code>O_EXCL</code> lock so a child never drains its own queue twice concurrently. <code>devswarm-child-turn</code> now statically nudges the child to run it every turn (no spawn on the hot path — the nudge is a string, not a call). Honest limits, not hidden: reception is <strong>pull, not push</strong> (a parent message is seen at most one child turn late, since a child can't host the blocking <code>monitor</code> daemon on its turn thread), and there is a <strong>destructive-read crash-window</strong> — <code>read-messages</code> marks the native messages read before <code>pullOnce</code> durably persists them, so a crash in that window loses them from the native side without landing in the durable inbox; the <code>message-count</code> gate minimizes but cannot close the window (hivecontrol exposes no non-destructive full read).</li> <li><strong>Ingest daemon now actually functions (fix).</strong> v0.54.1's auto-installer baked no <code>WorkingDirectory</code> into the plist/systemd unit/cron line, so launchd/systemd/cron ran the daemon from <code>$HOME</code> — not a git repo — and every <code>hivecontrol workspace monitor</code> call failed "Not in a git repository", draining nothing. <code>install-devswarm-ingest.js</code> now resolves the git worktree the installer was RUN from (<code>git rev-parse --show-toplevel</code> against the install-time cwd) and bakes it into the unit: macOS <code>WorkingDirectory</code> plist key, Linux systemd <code>WorkingDirectory=</code> (escaped the same way as <code>ExecStart</code>), and a <code>cd '<worktree>' &&</code> prefix on the cron fallback line (single-quoted, no injection hole reopened). Refuses to install (fail-open — log + skip, exit 0) if the install-time cwd doesn't resolve to a git worktree — no daemon is better than one that silently drains nothing.</li> <li><strong>Lock hardening.</strong> <code>withMessagesLock</code> (journal backend) no longer ever runs the append critical section UNLOCKED when contention is exhausted: a genuine fs error opening the lock (e.g. <code>EPERM</code>) now fails closed with a distinct <code>ELOCKFS</code>, and exhausting the retry budget throws a distinct, retryable <code>ELOCKUNAVAIL</code> that <code>appendMessage</code> retries with jittered backoff (idempotent by hash, so a retry can only add a row once). Also fixed a torn-read-steal race: a lock file is briefly 0 bytes between <code>openSync('wx')</code> and the write, and a concurrent reader that caught that empty window was wrongly treated as "holder absent" and stealable; readers now fall back to the lock file's mtime when the content is torn/unparseable and steal only on a genuinely OLD mtime, never a fresh torn read. The same torn-read-safe steal propagated to the <code>devswarm-ingest</code> / <code>devswarm-migrate</code> / <code>inbox pull</code> <code>O_EXCL</code> locks — closes a double-consumer race on the destructive native queue shared by any two of those. This is the exact race behind the CI failure caught on <code>windows-latest / node 24.x</code> for v0.54.1's "concurrent writers never duplicate a deduped hash row" test (<code>actual: 2, expected: 1</code>) — verified against the actual failed run before writing this note.</li> <li><strong>Reception silent-loss now surfaced (hardening).</strong> <code>pullOnce</code> count-gates a destructive <code>read-messages</code> on the non-destructive <code>message-count</code>, then relies on <code>normalizeMonitorPayload</code> — but an unhandled read-messages shape normalized to <code>[]</code>, so the messages were marked-read natively yet never persisted and the drain returned a quiet <code>imported:0, ok:true</code>. It now RECONCILES: when a drain recovers fewer messages than the native count (<code>imported+duplicate < message-count</code>) it surfaces a loud <code>lost</code> field, logs a best-effort telemetry line, and returns <code>ok:false</code>. The two are the same native unread metric (<code>message-count</code> counts the unread set; <code>read-messages</code> reads that same set), so a shortfall is real loss, never a benign count-vs-read race (a message arriving between the two calls can only make <code>read-messages</code> return more, never fewer).</li> <li><strong>Ingest daemon no longer crash-loops on a store-lock error (hardening).</strong> The store's messages lock fails CLOSED on <code>ELOCKFS</code> (a genuine fs/<code>EPERM</code> error) and <code>ELOCKUNAVAIL</code> (contention budget exhausted) — correct, but an uncaught throw propagated out of <code>runIngestLoop</code>, exiting the daemon so launchd/systemd re-exec'd it every restart interval, hammering the same wedged lock. The ingest loop now CATCHES those two known-retryable lock signals, logs, and continues to the next poll (the native queue buffers and replay is idempotent by hash); any other error still propagates so real bugs are not swallowed.</li> </ul> <h2 id="0541">0.54.1<a class="headerlink" href="#0541" title="Permanent link">¶</a></h2> <p><strong>DevSwarm substrate follow-up — ingest daemon auto-install, child-gate over-nag fix, partial child reception surfacing, live active-workspace table.</strong></p> <ul> <li><strong>Ingest daemon auto-installer (<code>companion/install-devswarm-ingest.js</code>).</strong> The <code>devswarm-ingest</code> daemon (the one native consumer wrapping <code>hivecontrol workspace monitor</code> into the store) now AUTO-installs/refreshes on <code>/anti-hall:update</code> inside an active DevSwarm session — same no-offer, no-ask posture as the supervisor installer, same <code>isDevswarmActive(process.env)</code> gate. Continuous daemon (not a periodic sweep): macOS LaunchAgent with <code>KeepAlive</code>, Linux <code>systemd --user</code> <code>.service</code> with <code>Restart=always</code> (cron fallback on systemctl-less hosts — every-minute tick, so a cron-only Linux host has up to ~60s revive gap after a crash). Idempotent (<code>launchctl unload && load</code> / <code>systemctl daemon-reload</code> + <code>restart</code>), so it first-installs when absent and refreshes an already-running daemon to the current build. <code>capability-scan.js</code> now detects BOTH shapes of a Linux systemd unit (<code>.timer</code> for the periodic supervisor, <code>.service</code> for the continuous ingest daemon) so either is reported correctly. Closes the gap where the daemon existed in code (0.54.0) but nothing ever started it.</li> <li><strong><code>devswarm-child-gate</code> over-nag fix.</strong> The child Stop-gate now stays SILENT when the child's own turn-authored heartbeat (<code>devswarm-child-turn</code>'s <code>heartbeats/<branch>.json</code>) is FRESH (<5 min old) — it no longer forces a duplicate heartbeat on every single Stop. The forced-ack now fires only for the genuinely unreported case (no heartbeat yet, or one stale past the freshness window).</li> <li><strong>Child inbox reception surfacing — PARTIAL, not full reception.</strong> <code>devswarm-child-turn</code> now does a non-destructive unread check against the child's own durable descriptor inbox and, when unread > 0, tells the child how many unread parent message(s) it has and the safe (non-draining) way to read them via the CLI's <code>inbox read</code> primitive. This is surfacing-only and forward-compatible: nothing shipped yet actually DRAINS the child's native parent→child queue into that durable inbox (native reads are destructive and guard-blocked; nothing currently populates the child's durable inbox from the native side). Full child-side reception is an explicit <strong>v0.54.2 follow-up</strong> — do not read this as "child message reception now works."</li> <li><strong>Live active-workspace table (<code>devswarm-parent-inbox</code>).</strong> The Primary's UserPromptSubmit hook now injects a compact status table EVERY turn — one row per active workspace, columns workspace / status (<code>escalated</code> > <code>stale</code> > <code>archive-ready</code><blockquote> <p><code>active</code>, attention-needing rows sorted first) / finishing rate (required completion gates met/total, plus an optional heartbeat progress-percent) / unread count / last-activity (relative age). Capped at 12 rows (<code>+N more</code>, logged, never silently truncated); empty output when no active workspace exists; read-only, fail-open, zero git calls on the hot path (reads <code>summary.json</code> + the liveness verdict + heartbeat files only).</p> </blockquote> </li> </ul> <h2 id="0540">0.54.0<a class="headerlink" href="#0540" title="Permanent link">¶</a></h2> <p><strong>DevSwarm coordination substrate — mechanical parent/child triggers, SQLite-backed store, ingest daemon, CLI, auto-safe migration, archive-ready recommendation.</strong></p> <ul> <li><strong>Mechanical parent/child trigger hooks.</strong> Four new opt-in hooks, all fail-open and empty-when-idle (own state files, no output unless the DevSwarm workspace is actually active):</li> <li><code>devswarm-parent-inbox</code> (UserPromptSubmit) — injects unread/idle status for the parent, plus a persistent reminder once a workspace becomes archive-ready.</li> <li><code>devswarm-parent-gate</code> (Stop) — capped forced-ack on unread/stale state, surfaces supervisor verdicts; does not touch git on Stop.</li> <li><code>devswarm-child-turn</code> (UserPromptSubmit) — child heartbeat.</li> <li><code>devswarm-child-gate</code> (Stop) — heartbeat forced-ack for the child.</li> <li><strong><code>devswarm-inbox-cursor</code></strong> — the read/ack primitive underlying the forced-ack clear path shared by the parent/child gate hooks.</li> <li><strong>SQLite dual-backend store</strong> (<code>node:sqlite</code> feature-detected, journal-file fallback when unavailable, WAL mode) with an atomic <code>summary.json</code> projection; hooks read only the projection, never the DB directly.</li> <li><strong><code>devswarm-ingest</code> daemon</strong> — wraps <code>hivecontrol workspace monitor</code> behind a single-native-consumer lock (never steals a live lock; heartbeats while running) with dedupe.</li> <li><strong><code>scripts/devswarm.js</code> CLI</strong> — <code>register</code> / <code>heartbeat</code> / <code>inbox</code> (<code>read</code>/<code>ack</code>/<code>count</code>) / <code>workspaces</code> / <code>nudge</code> / <code>archive</code> / <code>gate</code> / <code>archive-ignore</code> / <code>migrate</code>; JSON output throughout (CLI-over-MCP, per project convention).</li> <li><strong>Automatic-but-safe migration</strong> wired into the updater: idempotent, non-destructive, single-consumer-locked, verifies row counts before switching over.</li> <li><strong>Archive-ready recommendation.</strong> Tracks per-workspace completion gates (done/merged/tests_passed plus consumer-defined gates, e.g. deployed); once met, PERSISTENTLY reminds the user to archive the workspace in the DevSwarm app (GUI-only teardown), with a per-workspace <code>archive-ignore</code> opt-out. Never auto-archives.</li> <li><strong><code>system-briefing</code> companion skill</strong> — derived live from the current hook/KB map; hivecontrol KB refreshed to 2.3.4.</li> <li>Note: the 4 DevSwarm Phase-1 hooks are documented-non-mirror on the Codex port (env-gated inert there; tested).</li> <li><strong>Cross-model review hardening:</strong> ingest live-pid lock guard + heartbeat, migration false-verify-on-unreadable fix, <code>summary.json</code> unique-tmp write, child-gate fail-open, journal-dedupe lock, CLI register validation.</li> </ul> <h2 id="0530">0.53.0<a class="headerlink" href="#0530" title="Permanent link">¶</a></h2> <p><strong>DevSwarm destructive-read redirect (command-guard) + topology-aware edit-guard wording.</strong></p> <ul> <li><strong><code>command-guard</code> — new DevSwarm destructive-read redirect.</strong> Under a DevSwarm-active session (<code>isDevswarmActive</code>), the guard now redirects the two CONSUMING native <code>hivecontrol</code> inbox reads before its own skip/coordinator gate, in ALL contexts (not coordinator-only — a delegated read drains the queue identically):</li> <li><code>hivecontrol workspace monitor</code> blocks UNCONDITIONALLY (a no-timeout long-poll that hangs the shell AND consumes the native queue).</li> <li><code>hivecontrol workspace read-messages</code> blocks ONLY when durable-inbox evidence exists (<code>ANTIHALL_DEVSWARM_INBOX_CMD</code> non-empty, OR a <code>~/.anti-hall/devswarm/workspaces/*.json</code> descriptor with a truthy <code>inboxPath</code>); otherwise a harmless single-consumer read is ALLOWED (fail-OPEN-to-allow).</li> <li>Matching reuses the heavy-path quote-neutralization + <code>bash -c</code>/<code>eval</code>/<code>$()</code>/backtick recursion, so a grep/echo of the command as quoted DATA does not false-positive while smuggled forms still block. Closed-vocabulary reason never echoes input.</li> <li>Its own <code>devswarm-read-guard</code> skip name, now added to skip-guard's <code>DESTRUCTIVE</code> set — a blanket <code>all</code> skip does NOT silence it (it prevents irreversible native-queue drain / data loss); an explicit <code>{"devswarm-read-guard": <ttl>}</code> skip is required, like git-guard.</li> <li>Fully fail-open: any error falls through and never blocks a turn.</li> <li><strong><code>edit-guard</code> — topology-aware DevSwarm block wording.</strong> The DevSwarm block reason now distinguishes the Primary (<code>the primary/main orchestrator</code>) from a child workspace (<code>the sub-orchestrator</code>) via <code>isChildWorkspace</code>; both still block, only the noun changes.</li> <li><strong>deadly-loop / ship-it — Codex CRITIC seat pinned to <code>gpt-5.6-sol</code>.</strong> The adversarial Critic seat in both workflows now pins the flagship reasoning model via the brief prefix <code>--fresh --model gpt-5.6-sol</code> (deadly-loop <code>investigateAgent</code>, ship-it <code>criticAgent</code> — the single Codex-critic call site in each). This <strong>supersedes</strong> the 0.52.0 note that the workflows "neither pins a Codex model by design": the Critic is now deliberately pinned. The Codex IMPLEMENTER seat (ship-it <code>buildAgent</code>) stays unpinned (<code>gpt-5.6-terra</code>); the availability-fallback (Codex-unavailable → Opus) and cross-model self-review guard are untouched, so the sol pin can never leak onto an Opus seat. Both <code>MODEL-POLICY.md</code> copies updated in lockstep. On codex CLI v0.143.0 the <code>-m</code> pin may emit "Model metadata not found" (fallback metadata) per <code>docs/KB-gpt-5.6.md</code> — acceptable.</li> </ul> <h2 id="0520">0.52.0<a class="headerlink" href="#0520" title="Permanent link">¶</a></h2> <p><strong>New KB + Codex model-tier migration to GPT-5.6 (Sol/Terra/Luna).</strong></p> <ul> <li><strong>New <code>docs/KB-gpt-5.6.md</code></strong> — a cited knowledge-base doc on OpenAI's GPT-5.6 lineup (Sol/Terra/Luna): tier breakdown, official model IDs (<code>gpt-5.6-sol</code>, <code>gpt-5.6-terra</code>, <code>gpt-5.6-luna</code>; bare <code>gpt-5.6</code> resolves to Sol), pricing (Sol $5/$30, Terra $2.50/$15, Luna $1/$6), GA date 2026-07-09, with honest confidence bands (primary vs secondary vs press sourcing) and the local codex #31873 picker/cache caveat noted explicitly.</li> <li><strong>Codex port model-tier migration to GPT-5.6</strong> across the codex skills, README, the <code>MODEL-POLICY.md</code> line, and the limit-conserve downshift hook (updated in lockstep with its test): planning/debate <code>gpt-5.5</code>→<code>gpt-5.6-sol</code>; implementation <code>gpt-5.4</code>→<code>gpt-5.6-terra</code> — both are same-price capability upgrades over their predecessors. The cheap tier KEEPS <code>gpt-5.4-mini</code> as the default (Luna is offered only selectively — it's +33% over <code>gpt-5.4-mini</code>). The #31873 degraded-metadata caveat is carried forward, and <code>gpt-5.5</code> is kept as the recommended non-degraded fallback until the codex CLI cache issue is fixed.</li> <li><strong>Deadly-loop / ship-it workflows unchanged.</strong> Neither pins a Codex model by design — every Codex seat in those workflows uses <code>agentType codex:codex-rescue</code>, which resolves its own backend.</li> <li><strong>Correction:</strong> GPT-5.6 reached GA on 2026-07-09. The 0.49.0 entry above (lines 156–159, "Codex model routing — re-verified, unchanged") stated GPT-5.6 was "preview/select-partners-only, not GA" at the time — that note is now superseded by this entry; it is left unedited as a historical record.</li> </ul> <h2 id="0511">0.51.1<a class="headerlink" href="#0511" title="Permanent link">¶</a></h2> <p><strong>P0 fix — <code>codex-availability</code> PATH probe falsely reported a directory as an executable.</strong></p> <ul> <li><strong><code>hooks/codex-availability.js</code>.</strong> <code>probeCodexOnPath()</code> matched a candidate with <code>fs.accessSync(candidate, X_OK)</code> and no <code>isFile()</code> check. On POSIX a DIRECTORY named <code>codex</code> on PATH has the execute/search bit set, so <code>X_OK</code> succeeded and the hook wrote <code>available:true</code> / emitted the prefer-Codex <code>additionalContext</code> even though no runnable binary existed — contradicting the hook's own comment that it matches "a real executable (never a shell alias)". Fixed: a candidate now counts only if <code>fs.statSync(candidate).isFile()</code> is true (symlinks to a real binary still resolve via <code>statSync</code>'s follow-symlink behavior), and on POSIX the execute bit (<code>X_OK</code>) is additionally required; on Windows <code>isFile()</code> plus the <code>PATHEXT</code>-extension match is sufficient. Still fully fail-open (any stat/access throw just means "not a match here"; the hook always exits 0).</li> <li><strong>Reference probe parity.</strong> The byte-identical <code>codex-available.js</code> reference snippet in both <code>skills/MODEL-POLICY.md</code> and <code>skills/deadly-loop/references/MODEL-POLICY.md</code> carried the same bug and got the same <code>isFile()</code> fix, edited identically so the two copies remain byte-for-byte the same.</li> <li><strong>New regression test</strong> <code>tests/hooks/codex-availability.test.js</code>: a directory named <code>codex</code> on PATH (the bug) now asserts <code>available:false</code> and no emitted context; a real executable file named <code>codex</code> on PATH asserts <code>available:true</code>; an empty PATH asserts fail-open (<code>available:false</code>, exit 0, no throw).</li> </ul> <p><strong>Docs — <code>fable-availability.js</code> documented as intentionally Claude-only.</strong> No behavior change: <code>hooks/fable-availability.js</code> probes <code>~/.claude.json</code> for a Claude Fable model entitlement to inform the Claude Reviewer-seat fallback, which is irrelevant to gpt-5.x Codex/OMX sessions and (like the DevSwarm liveness supervisor) has no Codex mirror by design — Fable routing is itself policy-disabled (see <code>MODEL-POLICY.md</code>). Documented in <code>plugins/anti-hall/codex/README.md</code>'s intentional-parity list, with a matching inline comment in <code>codex/install-codex.js</code>.</p> <h2 id="0510">0.51.0<a class="headerlink" href="#0510" title="Permanent link">¶</a></h2> <p><strong>"Strongly-defaulted" Codex utilization</strong> — a new <code>codex-availability</code> SessionStart hook (Claude + Codex port) plus guidance edits pushing Codex toward being the default worker pool rather than an occasional afterthought. This strongly defaults Codex usage but does NOT mechanically guarantee it on the inline-Skill path — a saved Workflow template is still required for the enforced <code>codexUp</code> Critic wiring.</p> <ul> <li><strong>New hook <code>hooks/codex-availability.js</code>.</strong> A pure-Node, OS-agnostic PATH probe (mirrors the probe documented in <code>MODEL-POLICY.md</code>) that checks once per session whether a real <code>codex</code> executable is on PATH, writes <code>~/.anti-hall/codex-availability.json</code> (<code>{available, checkedAt, source:"path-probe"}</code>), and — when <code>available:true</code> — emits an honest <code>additionalContext</code>: this proves the binary is reachable, NOT that Codex is authenticated/ready, so a runtime spawn returning null still falls back to Opus/Sonnet. Fails open (any error -> exit 0, no output). Registered in both <code>hooks/hooks.json</code> (after <code>fable-availability.js</code>) and the Codex port's <code>codex/hooks/hooks.json</code>.</li> <li><strong><code>doctor</code> gains a non-failing "saved workflow template" advisory.</strong> Looks for <code>~/.claude/workflows/deadly-loop*.js</code> / <code>ship-it*.js</code> (and the <code><cwd>/.claude/workflows/</code> equivalents); if none are saved, prints an advisory (warning, not a failure — exit code unchanged) that the inline SKILL path leaves the Critic seat as unenforced LLM guidance that can silently degrade to Opus, and points at <code>/workflows</code> for the enforced wiring.</li> <li><strong>Guidance edits</strong> (<code>MODEL-POLICY.md</code> — both copies, <code>skills/orchestration/SKILL.md</code>, <code>skills/deadly-loop/SKILL.md</code>, <code>skills/ship-it/SKILL.md</code>): read the <code>codex-availability</code> fact before re-probing; when available, default the Critic seat and everyday correctness-review/implementation load to <code>codex:codex-rescue</code> with Opus as the fallback taken only on a null spawn (never a default); prefer the Workflow tool with a saved deadly-loop/ship-it template over ad-hoc inline Agent fan-out for big/parallel work.</li> </ul> <h2 id="0500">0.50.0<a class="headerlink" href="#0500" title="Permanent link">¶</a></h2> <p><strong>New always-on guard <code>edit-guard</code> — the coordinator can no longer edit files directly.</strong> Mirrors what <code>command-guard</code> already does for heavy Bash: a PreToolUse hook mechanically blocks a coordinator/orchestrator from calling <code>Edit</code>/<code>Write</code>/<code>MultiEdit</code>/<code>NotebookEdit</code> directly. Only coordinators are gated — subagents always pass, which prevents infinite agent nesting — and it fails open on ambiguity.</p> <ul> <li><strong>Shared coordinator detection.</strong> <code>hooks/coordinator-detect.js</code> is extracted from <code>command-guard.js</code> so both guards share the exact same coordinator/subagent discriminator (behavior-preserving refactor, no change to <code>command-guard</code>'s own logic).</li> <li><strong>Decision A — normally-skippable.</strong> <code>edit-guard</code> is intentionally NOT added to skip-guard's DESTRUCTIVE set; it's skippable like its sibling <code>command-guard</code>. Skip name is <code>edit-guard</code> (write <code>~/.anti-hall/skip.json</code> <code>{"edit-guard": <expiry-ms>}</code>).</li> <li><strong>Decision B — built-in allowlist.</strong> Coordinator-owned config/plan/memory paths are always allowed through: <code>CLAUDE.md</code>, <code>AGENTS.md</code>, <code>GEMINI.md</code>, <code>.claude/**</code>, <code>.omc/**</code>, <code>.anti-hall/**</code>, <code>PLAN.md</code>, <code>STATE.json</code>. Extensible via env <code>ANTIHALL_EDIT_GUARD_ALLOW</code> (colon/comma-separated globs).</li> <li><strong>DevSwarm-aware.</strong> Enforcement is always on in both standalone and DevSwarm sessions; only the block-message wording adapts via <code>isDevswarmActive()</code> — the guard is agnostic and works identically with or without DevSwarm.</li> <li><strong>New hooks.json entry.</strong> A separate PreToolUse matcher <code>Write|Edit|MultiEdit|NotebookEdit -> edit-guard.js</code> was added; the existing <code>api-guard</code>/<code>ship-it-guard</code> matcher is left untouched, so <code>NotebookEdit</code> isn't dragged onto them.</li> <li><strong>Doctor + tests.</strong> <code>doctor</code> gained a live behavioral self-test for <code>edit-guard</code> (the first edit-time behavioral probe); <code>tests/hooks/edit-guard.test.js</code> adds 17 tests.</li> <li><strong>Codex parity gap (documented, not claimed).</strong> <code>edit-guard</code> shares the same documented <code>apply_patch</code> payload-adapter gap as <code>api-guard</code>/<code>ship-it-guard</code> — it is NOT registered in the Codex hooks yet (a real Codex <code>apply_patch</code> PreToolUse payload must be captured first). No hard-block parity is claimed for Codex edit-time.</li> </ul> <h2 id="0490">0.49.0<a class="headerlink" href="#0490" title="Permanent link">¶</a></h2> <p><strong>DevSwarm recovery redesign — the supervisor NO LONGER auto-kills.</strong> The automatic path is now layered and non-destructive (detect → poke → escalate-to-parent); the hardened kill+<code>--resume</code> survives only as a deliberate on-demand CLI. Plus a child self-report hook, a fully autonomous updater, and a mid-turn-kill safety hardening.</p> <h3 id="-behavior-change-automatic-kill-removed">⚠️ Behavior change: automatic kill removed<a class="headerlink" href="#-behavior-change-automatic-kill-removed" title="Permanent link">¶</a></h3> <p>The 0.47–0.48 supervisor auto-killed + <code>--resume</code>d a workspace it judged wedged. <strong>That automatic kill is removed.</strong> Rationale: the real-world failure is overwhelmingly <em>idle-not-wedged</em> (a finished child with a dead listener), where a poke recovers it non-destructively; a blind automatic kill is too aggressive. If you relied on automatic recovery, it now <strong>escalates to the parent/you</strong> instead — and you kill deliberately via the new CLI. An already-installed 0.48 supervisor picks up this no-kill behavior automatically on its next sweep (the companion is an ephemeral per-interval spawn, not a persistent daemon); old on-disk <code>recovering</code> verdicts are simply superseded.</p> <h3 id="the-layered-liveness-model-automatic-path--never-kills">The layered liveness model (automatic path — never kills)<a class="headerlink" href="#the-layered-liveness-model-automatic-path--never-kills" title="Permanent link">¶</a></h3> <ul> <li><strong>Layer 1 — child self-report</strong> (new SessionStart hook <code>hooks/devswarm-child-role.js</code> + <code>hooks/lib/devswarm-role.js</code>): a DevSwarm <strong>child workspace</strong> sub-orchestrator is reminded to message its parent "idle — reassign or archive me" when idle. Role-gated by <code>DEVSWARM_SOURCE_BRANCH</code> — fires <strong>only</strong> on child workspaces, never the Primary/main orchestrator, never subagents.</li> <li><strong>Layer 2 — supervisor poke</strong> (non-destructive): on <code>stale</code>, if the descriptor carries an optional <code>nudgeCommand</code>, the sweep fires it (the consumer's real wake channel) — no pid targeting at all.</li> <li><strong>Layer 3 — escalate-to-parent</strong>: after the poke budget is exhausted, write <code>status:"escalated"</code> + a <code>recovery.log</code> line, and fire an optional <code>escalateCommand</code>. <strong>The automatic sweep stops here.</strong></li> </ul> <h3 id="on-demand-recovery-cli--the-only-kill-path">On-demand recovery CLI — the ONLY kill path<a class="headerlink" href="#on-demand-recovery-cli--the-only-kill-path" title="Permanent link">¶</a></h3> <ul> <li><code>node companion/devswarm-recover.js <workspace-id></code> — a deliberate, human/parent-invoked kill + <code>claude --resume</code> for one named workspace. Reuses the deadly-loop-hardened confirm-gate (exactly-one-candidate-or-abstain, identity-binding, cwd match, TOCTOU re-confirm before every signal, single-writer lock, recovery cap), with one relaxation: it targets <strong>interactive</strong> sessions too (DevSwarm launches children interactively; naming the id is itself the deliberate override of the human-takeover protection).</li> <li><strong>Mid-turn-kill hardening</strong> (from a live test that caught a real double-execution risk): the CLI <strong>group-kills</strong> the process tree (a lone <code>claude</code> SIGKILL otherwise orphans the in-flight tool subprocess), and the resume prompt is <strong>prepended with a state-check guardrail</strong> ("before re-running any side-effecting command, verify with a read-only check whether it already completed") — which measurably stopped the model from blindly re-running a mutating command on resume.</li> </ul> <h3 id="autonomous-updater">Autonomous updater<a class="headerlink" href="#autonomous-updater" title="Permanent link">¶</a></h3> <ul> <li><code>/anti-hall:update</code> (+ Codex mirror) now <strong>autonomously installs or refreshes</strong> the supervisor when the update runs inside a DevSwarm session (<code>DEVSWARM_REPO_ID</code>) — no prompt, no manual step (the installer is idempotent and reloads, so the next sweep runs the new code). Safe unprompted precisely because the daemon never kills. Outside a DevSwarm session it does nothing.</li> </ul> <h3 id="config--doctor">Config + doctor<a class="headerlink" href="#config--doctor" title="Permanent link">¶</a></h3> <ul> <li>Sweep env: dropped <code>ANTIHALL_DEVSWARM_MAX_RECOVERIES</code>/<code>GRACE_SEC</code>/<code>STUCK_SEC</code> (kill-era knobs); added <code>ANTIHALL_DEVSWARM_NUDGE_MAX_ATTEMPTS</code>/<code>NUDGE_WINDOW_SEC</code>/<code>NUDGE_COOLDOWN_SEC</code>. The on-demand CLI keeps its own cap/grace, decoupled from the sweep.</li> <li>Verdict enum is now <code>alive | stale | nudged | ambiguous | escalated</code> (dropped <code>recovering</code>); doctor maps <code>nudged</code> → WARN and drops the stuck-recovering logic. A genuine <strong>simplification</strong> — removing automatic kill also removes most of the precision-kill machinery from the sweep path.</li> </ul> <h3 id="codex-model-routing--re-verified-unchanged">Codex model routing — re-verified, unchanged<a class="headerlink" href="#codex-model-routing--re-verified-unchanged" title="Permanent link">¶</a></h3> <ul> <li>Confirmed against official sources that the current Codex flagship is still <strong><code>gpt-5.5</code></strong> (GPT-5.6 is preview/select-partners-only, not GA). The plugin's <code>gpt-5.5</code>→<code>gpt-5.4</code>→<code>gpt-5.4-mini</code> routing already matches; no change. (Re-check at the next release if 5.6 goes GA.)</li> </ul> <h2 id="0480">0.48.0<a class="headerlink" href="#0480" title="Permanent link">¶</a></h2> <p><strong>DevSwarm addons rounded out + plugin self-awareness: tunable supervisor thresholds, a new <code>/anti-hall:devswarm</code> skill, a capability-aware smart update, an environment-aware doctor, and account-change-aware limit-conservation. Every optional integration stays dormant unless it's actually in use.</strong></p> <h3 id="devswarm-liveness-supervisor--thresholds-now-env-tunable">DevSwarm liveness supervisor — thresholds now env-tunable<a class="headerlink" href="#devswarm-liveness-supervisor--thresholds-now-env-tunable" title="Permanent link">¶</a></h3> <ul> <li>The supervisor's runtime thresholds (previously hardcoded) are configurable, fail-open to the same defaults: <code>ANTIHALL_DEVSWARM_IDLE_SEC</code> (900), <code>ANTIHALL_DEVSWARM_COOLDOWN_SEC</code> (600), <code>ANTIHALL_DEVSWARM_MAX_RECOVERIES</code> (3), <code>ANTIHALL_DEVSWARM_GRACE_SEC</code> (5), and doctor's <code>ANTIHALL_DEVSWARM_STUCK_SEC</code> (1800). (<code>ANTIHALL_DEVSWARM_INTERVAL</code> was already tunable.) Also fixed a latent gap — <code>graceMs</code> wasn't being threaded into recovery at all before.</li> </ul> <h3 id="new-skill-anti-halldevswarm--codex-mirror">New skill: <code>/anti-hall:devswarm</code> (+ Codex mirror)<a class="headerlink" href="#new-skill-anti-halldevswarm--codex-mirror" title="Permanent link">¶</a></h3> <ul> <li>Explains the whole OPTIONAL DevSwarm integration — the hivecontrol KB, the (designed, unbuilt) workspace-tier orchestration, and the shipped liveness supervisor — plus a consumer <strong>activation checklist</strong>: the install command, the <code>~/.anti-hall/devswarm/workspaces/<id>.json</code> descriptor schema (required vs load-bearing fields), env gates, keep-fresh rules, config/tuning env, the safety model, and where to watch outputs. A DevSwarm-orchestrating agent invokes this to get the as-built activation contract — no separate hand-off doc needed.</li> </ul> <h3 id="smart-capability-aware-update">Smart, capability-aware update<a class="headerlink" href="#smart-capability-aware-update" title="Permanent link">¶</a></h3> <ul> <li><code>/anti-hall:update</code> (+ Codex mirror) now runs a <strong>dynamic capability scan</strong> after pulling: it discovers opt-in companions from <code>companion/install-*.js</code> (no hardcoded list), checks which are actually installed on this machine (launchd / systemd / cron artifacts), plus statusline state and pending state-migrations, and reports "available vs active + how to enable." Idempotent migrations run automatically; opt-in installs are <em>guided</em>, never auto-run (a process-killer must stay opt-in). New pure-Node <code>scripts/capability-scan.js</code>.</li> </ul> <h3 id="comprehensive-environment-aware-doctor">Comprehensive, environment-aware doctor<a class="headerlink" href="#comprehensive-environment-aware-doctor" title="Permanent link">¶</a></h3> <ul> <li><code>doctor.js</code> now detects and tests each present integration and cleanly skips the absent ones (no noise for a plain user): <strong>OMC</strong> (enabled + whether a loop is active), <strong>Codex/OMX</strong> (config + anti-hall hooks wired), and <strong>DevSwarm</strong> — including whether the supervisor companion is installed and an explicit per-workspace <strong>listener-presence</strong> line (an honest WARN when the listener state can't be observed, never a fabricated PASS).</li> </ul> <h3 id="limit-conservation-is-account-aware">Limit-conservation is account-aware<a class="headerlink" href="#limit-conservation-is-account-aware" title="Permanent link">¶</a></h3> <ul> <li>anti-hall's weekly-limit conservation nudge now <strong>deactivates when the Claude account changes</strong> (<code>userID</code> in <code>~/.claude.json</code>) until the usage cache is refreshed under the new account — fixing stale over-restriction after an account switch. Fail-open (any ambiguity → prior behavior), biased toward <em>not</em> over-restricting, with an <code>ANTIHALL_LIMIT_ACCOUNT_CHECK=off</code> kill-switch. Its firing depends on <code>userID</code> rotating on switch (a safe no-op if it doesn't).</li> </ul> <h2 id="0471">0.47.1<a class="headerlink" href="#0471" title="Permanent link">¶</a></h2> <p><strong>Windows CI fix for the 0.47.0 DevSwarm liveness supervisor.</strong> <code>encodeWorktreePath()</code> (<code>companion/lib/target-session.js</code>) stripped only <code>/</code> and <code>.</code> when mapping a worktree path to its <code>~/.claude/projects/<encoded></code> transcript dir — not <code>\</code> or the drive-letter <code>:</code>. On Windows an absolute <code>C:\…</code> worktree path therefore stayed a multi-component absolute segment, and <code>path.join</code> (which does not anchor on a later absolute-looking argument) produced a doubled <code>…\.claude\projects\C:\…</code> path → the doctor self-test and the liveness detector ENOENT'd on all Windows CI jobs (macOS + Linux were unaffected and green). Widened the strip to <code>/[/\\:.]/g</code> so the encoded segment is always one flat, filesystem-legal component on every platform — a no-op on POSIX (which has no <code>\</code>/<code>:</code>). <code>node --test</code> = 855 pass / 0 fail / 2 skip; green across ubuntu/macos/windows × node 18/20/22/24.</p> <h2 id="0470">0.47.0<a class="headerlink" href="#0470" title="Permanent link">¶</a></h2> <p><strong>New OPT-IN DevSwarm liveness supervisor — a pure-Node companion that detects a wedged/idle DevSwarm workspace agent from its OUTBOUND activity and precisely recovers it (targeted kill + <code>claude --resume</code>), fully dormant unless DevSwarm is in use. Hardened by a 3-seat deadly-loop that found 15 issues (incl. a wrong-victim P0) — all fixed with non-vacuous tests BEFORE any code shipped. <code>node --test</code> = 855 pass / 0 fail / 2 Windows-skips (+74 tests).</strong></p> <h3 id="what-it-is">What it is<a class="headerlink" href="#what-it-is" title="Permanent link">¶</a></h3> <ul> <li>A background companion (<code>companion/devswarm-supervisor.js</code>, installed via <code>companion/install-devswarm-supervisor.js</code> — the same opt-in launchd / systemd-user / cron model as the <code>mcp-reaper</code>). Per active DevSwarm workspace it computes liveness from <strong>OUTBOUND</strong> signals (the agent's own session-transcript activity + git/worktree activity — not inbound messages, which are blind to the hang) and recovers the agent only when BOTH signals are idle past a threshold AND the workspace has pending work it should be servicing.</li> <li>Works around an upstream Claude Code core hang (process stays alive but can't service the next turn, so crash-restarters never fire) — labeled a documented workaround for <code>claude-code#39755</code>, to re-evaluate when a real upstream fix lands.</li> </ul> <h3 id="optional--exactly-like-omcomx">Optional — exactly like OMC/OMX<a class="headerlink" href="#optional--exactly-like-omcomx" title="Permanent link">¶</a></h3> <ul> <li>Entirely feature-detected via <code>DEVSWARM_REPO_ID</code> (new <code>hooks/lib/devswarm-detect.js</code>, modeled on <code>omc-detect.js</code>). When DevSwarm is not in use, the supervisor and its doctor check are completely dormant — zero effect, nothing installed, no context. Opt-in install; kill-switch + config env. Documented as optional in README + llms.txt alongside OMC/OMX.</li> </ul> <h3 id="safety-it-kills-processes-so-it-is-hardened-accordingly">Safety (it kills processes, so it is hardened accordingly)<a class="headerlink" href="#safety-it-kills-processes-so-it-is-hardened-accordingly" title="Permanent link">¶</a></h3> <ul> <li><strong>Precise targeting, never a broad <code>pkill</code>:</strong> maps a workspace to exactly one <code>claude</code> pid via the process's own argv session-id + a real cwd match, and <strong>ABSTAINS</strong> on any ambiguity (0 or >1 candidates).</li> <li><strong>Identity-bound:</strong> the target's session-id must match the workspace's declared id AND the process must be <strong>headless (<code>-p</code>)</strong> — so a human who takes over a worktree interactively is never killed.</li> <li>Re-confirms identity on fresh data immediately <strong>before every kill signal</strong> (PID-recycle/TOCTOU defense, mirroring <code>mcp-reaper</code>); <strong>group-kills</strong> so it never orphans MCP children; <strong>detached <code>--resume</code></strong> so it can't time-out and kill legitimate long-running work; <strong>stale-lock steal</strong> (mirroring <code>swarm-guard</code>) so a supervisor crash can't permanently disable recovery; a per-workspace recovery cap that escalates instead of looping; single-flight sweep; descriptor-id sanitization; bounded <code>ps</code>/<code>lsof</code> probes. <strong>Fail-open throughout</strong> — any error logs and continues, never kills a healthy agent. <strong>Windows = detection-only</strong> (a running process's cwd isn't obtainable in pure Node there), matching the <code>mcp-reaper</code> platform stance.</li> <li><strong>doctor</strong> gains a per-active-workspace listener + liveness PASS/WARN/FAIL check (dormant when no DevSwarm) and now syntax-checks the new <code>lib/</code> files.</li> </ul> <h3 id="the-seam--anti-hall-stays-generic-the-consumer-owns-the-devswarm-glue">The seam — anti-hall stays generic; the consumer owns the DevSwarm glue<a class="headerlink" href="#the-seam--anti-hall-stays-generic-the-consumer-owns-the-devswarm-glue" title="Permanent link">¶</a></h3> <ul> <li>anti-hall ships only the generic, agnostic supervisor. The DevSwarm-specific transport (inbox daemon, done-report contract, <code>hivecontrol</code> wiring) stays consumer-side. Interface: the consumer publishes a per-workspace descriptor <code>~/.anti-hall/devswarm/workspaces/<id>.json = { id, worktreePath, inboxPath, cursorPath, sessionId }</code>; anti-hall derives pid/session itself and writes a liveness verdict + recovery log. Full design + implementation plan under <code>docs/superpowers/</code>.</li> </ul> <h2 id="0460">0.46.0<a class="headerlink" href="#0460" title="Permanent link">¶</a></h2> <p><strong>Whole-plugin ultracode audit remediated: 2 critical guard/data-loss bugs + 21 correctness/parity fixes, plus 3 opt-in DevSwarm/hivecontrol integration hooks. Every fix ships with a non-vacuous test; <code>node --test</code> = 783 pass / 0 fail / 2 Windows-skips.</strong></p> <h3 id="critical-p0">Critical (P0)<a class="headerlink" href="#critical-p0" title="Permanent link">¶</a></h3> <ul> <li><strong>git-guard <code>bash -c</code>/<code>sh -c</code> bypass closed.</strong> <code>scanCommand()</code> unwrapped only <code>eval</code>, so <code>bash -c "git push --force"</code> and <code>bash -c '… Co-Authored-By: Claude …'</code> both passed (exit 0) — a total bypass of the one guard the repo treats as non-skippable. Added <code>extractShellCPayload()</code> (SHELL_VERBS <code>bash/sh/zsh/dash/ksh/ash</code> + <code>-c</code>/<code>--command</code>) with depth-bounded recursion, mirroring <code>command-guard.js</code> / <code>graphify-guard.js</code>. Also fixed <code>splitSegments</code> splitting on the <code>&</code> in <code>2>&1</code> / <code>>&2</code> / <code>&></code>, which had orphaned a trailing <code>--force</code> into a non-<code>git</code> segment. Both classes now block (exit 2).</li> <li><strong>statusline uninstall no longer clobbers a committed settings file.</strong> <code>uninstall --project</code> targeted <code>.claude/settings.json</code> while the installer writes <code>.claude/settings.local.json</code>, and Strategy A overwrote it with the global base — no backup, no check it was even anti-hall's. Now resolves <code>settings.local.json</code> first, only rewrites when the current value is the anti-hall dispatcher, and backs up before mutating.</li> </ul> <h3 id="correctness--parity-p1">Correctness & parity (P1)<a class="headerlink" href="#correctness--parity-p1" title="Permanent link">¶</a></h3> <ul> <li><strong>Guards:</strong> command-guard adds a narrow, path-anchored exception for the coordinator-owned helpers <code>statusline/phase.js</code> + <code>hooks/agent-watchdog.js</code> (arbitrary <code>node *.js</code> still blocked in coordinator context); tasklist-guard's <code>BASH_WORK_RE</code> no longer counts read-only Bash (quoted content, <code>2></code> stderr redirects, mid-command write-verbs) as file-changing work.</li> <li><strong>Workflows / skills:</strong> <code>ship-it.workflow.js</code> now begins with <code>export const meta</code> (the Workflow runtime previously rejected it, so <code>/ship-it</code> could not run), its audit gate blocks on new P0 <strong>or</strong> P1 (was P0-only), and Codex-implemented code is reviewed by an Opus Critic (implementer ≠ reviewer). <code>deadly-loop.workflow.js</code> re-checks a respawned seat for drift before trusting it as a live GO. <code>deadly-loop/SKILL.md</code> swarm-mode Reviewer is Sonnet 5 (Fable routing is policy-disabled).</li> <li><strong>Codex port:</strong> every Codex skill resolves the plugin root via a <code>$ANTI_HALL_ROOT</code> preamble — verified against official Codex docs that <code>${PLUGIN_ROOT}</code> is expanded only for hook commands, never in skill bodies — instead of clone-relative <code>node plugins/anti-hall/…</code> paths that break once the port is installed as a plugin. The Codex ship-it <code>STATE.json</code> gained the 0.45.0 <code>gate:"locked"|"not-run"</code> field + PLAN.md <code>## Goal coverage</code> (anti-false-completion parity); <code>install-codex.js</code> dedup is backslash-safe so a Windows re-install no longer duplicates all 15 hook groups.</li> <li><strong>Statusline:</strong> <code>install --consolidate</code> on the already-installed path no longer throws a TDZ <code>ReferenceError</code>.</li> <li><strong>Docs accuracy:</strong> model-routing-guard default documented correctly (strict by default since v0.35.0; <code>=advisory</code> opts out — was documented backwards); the dead OMC companion link repointed to the real upstream; the <code>docs/KB.md</code> version row refreshed; and a private project codename scrubbed from a committed doc (agnostic-mandate fix).</li> </ul> <h3 id="devswarm--hivecontrol-integration-opt-in-off-by-default">DevSwarm / hivecontrol integration (opt-in, off by default)<a class="headerlink" href="#devswarm--hivecontrol-integration-opt-in-off-by-default" title="Permanent link">¶</a></h3> <ul> <li><strong>merge-gate</strong> also recognizes <code>hivecontrol workspace merge-into-source</code> / <code>merge-from-source</code> (stays default-OFF via <code>ANTIHALL_MERGE_GATE</code>; <code>skip.json</code> override intact).</li> <li><strong>deadly-loop</strong> writes an explicitly-<strong>advisory</strong> <code>~/.anti-hall/approvals/<repo>@<HEAD-sha>.json</code> on convergence (<code>"proof": false</code> — an audit record, never an authorization token; the reader must enforce its own real gating). Mirrored into the Codex deadly-loop skill.</li> <li><strong>swarm-guard</strong> appends a blocked-spawn trip to a separate <code>~/.anti-hall/swarm-trips.log</code> so a cap trip leaves a forensic trace (the rate window and <code>SPAWN_CAP</code> are untouched — observation only).</li> </ul> <h3 id="also-in-this-release-carried-forward-cleanup--docs">Also in this release (carried-forward cleanup + docs)<a class="headerlink" href="#also-in-this-release-carried-forward-cleanup--docs" title="Permanent link">¶</a></h3> <ul> <li><strong>Codex <code>anti-hall-feature-launch</code> fully removed</strong> — completes the 0.45.0 GSD/feature-launch retirement on the Codex side. The skill is deleted and every reference (<code>codex/README.md</code>, <code>docs/KB-omx.md</code>, <code>codex/skills/anti-hall-omx/SKILL.md</code>, and the plugin-manifest description) now points at <code>anti-hall-ship-it</code>.</li> <li><strong>graphify guidance modernized.</strong> <code>doctor.js</code>, <code>graphify-reminder.js</code>, <code>graphify-session.js</code>, and the SessionStart primer now say <code>graphify update .</code> (code graph) / <code>/graphify . --update --obsidian</code> (docs + Obsidian) instead of the retired <code>/graphify --obsidian</code>, and the legacy <code>.planning/graphs/</code> fallback is dropped (GSD is gone) — with a regression test asserting <code>.planning/graphs</code> alone no longer triggers the graphify-first injection.</li> <li><strong>New DevSwarm / hivecontrol docs.</strong> <code>docs/KB-devswarm-hivecontrol.md</code> — a reference KB for the <code>hivecontrol</code> CLI (v2.3.3), the <code>.devswarm/config.json</code> schema, and workspace role-detection — plus the approved DevSwarm-aware workspace-tier orchestration design + implementation plan under <code>docs/superpowers/</code>. The feature itself is not built yet; these are the hardened design/plan for a separate later effort.</li> </ul> <h2 id="0450">0.45.0<a class="headerlink" href="#0450" title="Permanent link">¶</a></h2> <p><strong>GSD discontinued and removed as a live dependency; 4 new research KBs; a 17-item cross-KB feature audit resolved; a plugin-wide KB-contradiction sweep; 2 real bugs fixed in <code>deadly-loop.workflow.js</code>.</strong></p> <h3 id="gsd-planning-removed--owner-decision-gsd-itself-is-discontinued">GSD (<code>.planning/</code>) removed — owner decision, GSD itself is discontinued<a class="headerlink" href="#gsd-planning-removed--owner-decision-gsd-itself-is-discontinued" title="Permanent link">¶</a></h3> <ul> <li><code>ship-it/SKILL.md</code>, <code>ship-it-guard.js</code>, and <code>docs/KB.md</code> no longer recognize <code>.planning/PLAN.md</code> — <code>PLAN.md</code> at the repo root is the only location. <code>graphify-guard.js</code>, <code>graphify-reminder.js</code>, and <code>doctor.js</code> no longer treat <code>.planning/graphs/</code> as an alternate graph location.</li> <li><code>scripts/migrate-state.js</code> gained <code>migrateGsdPlanning()</code>: folds an existing <code>.planning/</code> tree into <code>.anti-hall/history/legacy/planning/</code>, then <strong>deletes each source file once its own copy is verified byte-identical</strong> — the <code>.planning/</code> directory (and any subdirectory) is never removed, only the individual files inside it, one at a time. A file whose copy can't be verified is left in place, reported as <code>verify-failed</code>, never silently deleted. Wired into <code>/anti-hall:update</code>'s existing migration step (same command, no new invocation needed).</li> <li><code>AGENTS.md</code>'s dual-platform-parity section now documents two accepted structural limitations (Codex has no Dynamic-Workflows equivalent; Claude→Codex integration is one-directional) so future work stops re-litigating them, and explicitly extends the parity mandate to KBs.</li> </ul> <h3 id="4-new-research-kbs-all-source-count-verified-none-padded">4 new research KBs (all source-count-verified, none padded)<a class="headerlink" href="#4-new-research-kbs-all-source-count-verified-none-padded" title="Permanent link">¶</a></h3> <ul> <li><code>docs/KB-model-modes.md</code> (47 sources) — Claude/Codex effort levels, Plan Mode, the Workflow tool + "ultracode" (confirmed officially documented), <code>/code-review ultra</code>. Found and fixed a real terminology-drift bug (prompt text said "max thinking"/"max reasoning" when the actual code passes <code>effort:"high"</code>/<code>"xhigh"</code> — Codex has no <code>max</code> tier at all) across 10 files.</li> <li><code>docs/KB-overengineering.md</code> (15 sources) — general SE + AI-agent-specific causes of overengineering, empirical bloat measurement.</li> <li><code>docs/KB-false-completion.md</code> (21 sources) — reward hacking / specification-gaming research; identified that ship-it v2's own <code>STATE.json</code> protocol (shipped last release) was prose-only and mechanically unenforced — the exact failure mode the research catalogs, now fixed (below).</li> <li><code>docs/KB-goal-setting.md</code> (15 sources) — goal-clarity research; identified a gap between Step 1's intent and Step 2's decomposed phases with no reconciliation check — now fixed (below).</li> </ul> <h3 id="kb-contradiction-sweep--13-real-contradictions-found-and-fixed-across-15-kb-docs">KB-contradiction sweep — 13 real contradictions found and fixed across 15 KB docs<a class="headerlink" href="#kb-contradiction-sweep--13-real-contradictions-found-and-fixed-across-15-kb-docs" title="Permanent link">¶</a></h3> <p>Including a within-document self-contradiction in <code>KB-sonnet-5.md</code> (TL;DR vs its own cited source), a benchmark figure mis-attributed to the wrong model, a pre-Sonnet-5 routing doc never flagged as superseded, and a stale error-code claim contradicting its own correction three sections later. All 13 independently re-verified against the actual working tree (grep/ls counts), not just re-read.</p> <h3 id="feature-improvement-audit--all-17-findings-resolved">Feature-improvement audit — all 17 findings resolved<a class="headerlink" href="#feature-improvement-audit--all-17-findings-resolved" title="Permanent link">¶</a></h3> <ul> <li><strong>Doc/text fixes:</strong> <code>deadly-loop-multi</code> retry-arithmetic (wrong tier referenced), <code>task-guard.js</code> under-description, <code>gpt-5.3-codex-spark</code> presented as interchangeable with <code>gpt-5.4-mini</code> (fixed across 9 Codex files + added an Effort column to the Codex model-policy table), a 4-location Fable stale-language sweep, <code>update</code>/<code>doctor</code> SKILL.md delegation instructions missing an explicit model (live-verified to trigger <code>model-routing-guard.js</code>'s own block).</li> <li><strong>Real code fixes:</strong> <code>ship-it.workflow.js</code>'s <code>buildAgent()</code> now pins explicit effort (<code>medium</code>/<code>high</code>, opt-up via <code>phase.effort</code>) instead of silently inheriting a model default; <code>ship-it-guard.js</code> gained a plan-conformance advisory (flags — never blocks — an edit outside every phase's declared <code>files:</code>); <code>graphify-guard.js</code> is now subagent-aware (a delegated subagent's search no longer burns the coordinator's one-time nudge) and re-arms its nudge after ~240KB of transcript growth instead of firing once per session forever; <code>task-tracker.js</code> gained the same token-growth re-arm trigger alongside its wall-clock one; <code>doctor.js</code> now sums the footprint of every registered SessionStart hook (was undercounting) and gained a behavioral test for <code>version-alert.js</code>; Codex-native deadly-loop documents its same-model-TRIO as a disclosed degraded config with an opt-in cross-model escalation path for L-tier hard-risk work.</li> <li><strong>Empirical verification, not assumed:</strong> ran a real experiment confirming SubagentStart's <code>additionalContext</code> genuinely reaches a spawned subagent's model context (cross-checked against 37 historical subagent transcripts) — <code>docs/KB-claude-codex.md</code> updated accordingly.</li> <li><strong><code>ship-it/SKILL.md</code> protocol additions</strong> (from the two false-completion/goal-setting KB findings above): <code>STATE.json</code> gained a <code>gate:"locked"|"not-run"</code> field, set only after Step 5's deadly-loop actually locks — a resumed phase can no longer be trusted <code>done</code> from inference alone. <code>PLAN.md</code>'s template gained a "Goal coverage" field mapping each intent clause to the phase that proves it, plus a new Step 3 Reviewer-checklist bullet verifying that mapping.</li> </ul> <h3 id="2-real-bugs-fixed-in-deadly-loopworkflowjs-not-doc-only--logic-changes">2 real bugs fixed in <code>deadly-loop.workflow.js</code> (not doc-only — logic changes)<a class="headerlink" href="#2-real-bugs-fixed-in-deadly-loopworkflowjs-not-doc-only--logic-changes" title="Permanent link">¶</a></h3> <ul> <li>The Opus Reviewer-fallback silently inherited the Sonnet-5 seat's <code>effort:"xhigh"</code> via object spread instead of using Opus's own tier — fixed with a fresh options object (<code>effort:"high"</code>).</li> <li>Phase 2b (Argue) dispatched using each seat's <em>static</em> role-defined opts, never the opts that actually answered in Phase 2a — meant a seat whose primary model failed over during Investigate would silently retry the dead model in Argue with zero fallback and zero signal in <code>verdictSummary</code>. Fixed by threading the resolved per-seat opts from 2a into 2b.</li> </ul> <h3 id="gsd-removal-part-2--statusline--codex-side-parity-gaps">GSD removal, part 2 — statusline + Codex-side parity gaps<a class="headerlink" href="#gsd-removal-part-2--statusline--codex-side-parity-gaps" title="Permanent link">¶</a></h3> <p>A deeper sweep found GSD support was more extensive than the first pass caught: - <code>statusline-monorepo.js</code> was an entire dedicated rendering mode for GSD state (<code>readGsdState</code>/<code>formatGsdState</code>/<code>.planning/config.json</code>) — stripped down to its non-GSD content (model/context/current-task/dir); <code>.gsd/</code>/<code>.planning/</code> dropped as monorepo-detection triggers in <code>statusline.js</code> (<code>.gitmodules</code> still triggers monorepo mode for real git-submodule projects); <code>statusline-rich.js</code>'s GSD phase chip removed. - Two Codex-side skills contradicted their own stated intent: <code>anti-hall-feature-launch</code> explicitly said "do not invoke GSD commands, GSD was removed" while still instructing every artifact write into <code>.planning/</code> — fixed to <code>.anti-hall/feature-launch/</code>. <code>anti-hall-ship-it</code> still preferred <code>.planning/PLAN.md</code> — fixed to repo-root only, matching the Claude-side fix. - <code>install-statusline</code>'s SKILL.md and llms.txt's Codex README were also swept. - Historical/research docs (<code>docs/gsd-distilled.md</code>, <code>docs/KB-claude-codex.md</code> §12, <code>docs/superpowers-planning.md</code>) were deliberately left untouched — they document GSD's design as research provenance for anti-hall's own swarm/debate model, not live coexistence, matching this repo's "historical artifacts are frozen" convention.</p> <h3 id="process-note">Process note<a class="headerlink" href="#process-note" title="Permanent link">¶</a></h3> <p>Codex's independent verification of this release's implementation batch flagged one genuine defect (not environmental noise): a probe-record fixture claimed a grep returned zero matches when it actually returns one — corrected in place, logged rather than silently patched.</p> <p>Suite 729 (727 pass / 0 fail / 2 skip, up from 703 at the start of this release).</p> <h2 id="0440">0.44.0<a class="headerlink" href="#0440" title="Permanent link">¶</a></h2> <p><strong>Ship-it v2: resumable state, a global P0+P1 convergence gate, Codex-primary build seats, and legacy-state migration.</strong></p> <ul> <li><strong>Global deadly-loop convergence gate now blocks on confirmed P0 <em>or</em> P1</strong> (was P0-only). <code>deadly-loop.workflow.js</code>'s <code>VERDICT_SCHEMA</code> already tagged every finding P0/P1/P2/P3; the gate simply never checked P1 until now. This is a plugin-wide change — every deadly-loop consumer (ship-it, root-cause debugging, deadly-loop-multi) now requires zero NEW P0s AND P1s to converge, not just P0s. Mirrored in the Codex-native <code>anti-hall-deadly-loop</code> skill.</li> <li><strong>ship-it v2, L-tier only:</strong></li> <li>Resumable <code>.anti-hall/ship-it/<slug>/STATE.json</code> (coordinator-owned protocol — Dynamic Workflow scripts have no filesystem access, so this lives in <code>SKILL.md</code>, not the <code>.workflow.js</code> template): <code>plan_hash</code> (drift detection against <code>PLAN.md</code>), per-phase status, and an <code>escalations</code> counter. Reuses the existing <code>~/.anti-hall/agents/*.json</code> heartbeat convention (<code>agent-watchdog.js</code>) rather than inventing a new one.</li> <li>P2-severity findings from a converged deadly-loop no longer vanish — they're appended to <code>.anti-hall/ship-it/<slug>/decisions.md</code>.</li> <li>Build-to-plan escalations (a fix that needs re-planning, not just another fix-wave) are capped at 2; the 3rd stops and surfaces to the owner instead of looping.</li> <li>Step 6 now auto-writes a session-history entry (via the existing per-session system) plus <code>.anti-hall/ship-it/<slug>/SUMMARY.md</code>, then triggers <code>/graphify --obsidian --update</code>.</li> <li><strong>Build seats now try Codex-primary / Sonnet-5-failover</strong> (<code>buildAgent()</code> in <code>ship-it.workflow.js</code>, mirroring the existing <code>criticAgent()</code> fallback shape) instead of always going straight to Sonnet 5 — matching MODEL-POLICY.md's already-documented implementation-seat routing. Surfaced a real cross-model gap: when a phase's build falls back to Sonnet 5, that phase's Reviewer seat (also Sonnet-5-by-default) would otherwise be reviewing its own model's output. Fixed — the Reviewer now skips Sonnet 5 and goes straight to Opus for any phase Sonnet 5 itself built.</li> <li><strong><code>scripts/migrate-state.js</code></strong> (new) — non-destructive (copy-only, never deletes/moves), idempotent script that folds legacy root <code>.anti-hall-progress.md</code> / <code>.anti-hall-history.md</code> files into <code>.anti-hall/history/legacy/</code> for repos that haven't been through that migration yet. Wired into the <code>update</code> skill as a post-pull step.</li> <li><strong>Codex-native <code>anti-hall-ship-it</code> parity</strong> — same L-tier resumable-state convention, P0+P1 LOCK threshold, P2 decisions log, migrate-state.js reference, and wrap-up/summarize step, written as protocol text (this port has no Dynamic Workflow runtime, so nothing here relies on one). Codex-native ship-it is already the Codex-primary implementer by construction, so the cross-model self-review guard above doesn't apply on this port.</li> <li>Doc sweep: README (root + plugin) and llms.txt updated for the above; also corrected drifted test-pass counts (623/625 → 701/703) that had gone stale since 0.43.0.</li> </ul> <p>Suite 703 (701 pass / 0 fail / 2 skip, up from 688 — 15 new tests across the deadly-loop and ship-it workflow suites plus the new migrate-state suite).</p> <h2 id="0441">0.44.1<a class="headerlink" href="#0441" title="Permanent link">¶</a></h2> <p><strong>Fixed a Windows-only CI failure in 0.44.0: the new determinism test's comment-stripping regex silently no-op'd on CRLF checkouts.</strong></p> <ul> <li><code>tests/hooks/ship-it-workflow.test.js</code>'s determinism test stripped <code>//</code> comments by splitting on <code>\n</code> then matching <code>\/\/.*$</code> per line. On a Windows (CRLF) checkout each split line keeps a trailing <code>\r</code>; since regex <code>.</code> never matches a line terminator, <code>.*</code> can't reach the line's true end, so <code>$</code> (no multiline flag) never matches and the strip silently does nothing. The file's own doc-comment ("no Date.now() / Math.random() / argless new Date()") then survives verbatim into the "code" being scanned, and the test fails on every Windows Node version (18.x/20.x/22.x/24.x) — confirmed via <code>gh run view</code> after v0.44.0's push, reproduced locally with a CRLF-line simulation, and root-caused by comparing against <code>deadly-loop-workflow.test.js</code>'s equivalent test, which uses a <code>gm</code>-flag regex over the whole string instead (CRLF-safe by construction) and passed on the same run.</li> <li>Fix: normalize <code>\r\n</code> → <code>\n</code> before splitting, matching the CRLF-safe approach already used elsewhere. v0.44.0's tag is left as-is (not retagged) since it was never propagated to the marketplace or given a GitHub Release.</li> </ul> <p>Suite 703 (701 pass / 0 fail / 2 skip locally); Windows CI re-verified green after this fix.</p> <h2 id="0432">0.43.2<a class="headerlink" href="#0432" title="Permanent link">¶</a></h2> <p><strong>Fable routing policy-disabled: negative community feedback (over-restrictive/refusal-prone).</strong></p> <ul> <li><code>reviewerAgent()</code>/<code>buildFormation()</code> in ship-it.workflow.js and deadly-loop.workflow.js no longer attempt Fable for the Reviewer seat, even when <code>args.fableAvailable === true</code>. Reason: a soft refusal from an over-restrictive model would pass StructuredOutput schema validation as a "successful" verdict and get silently treated as real analysis -- worse than the already-handled unavailable/null case, and not worth building refusal-detection for given the community's reported experience with Fable's current behavior. Sonnet 5 is now the fixed primary Reviewer regardless of the flag; Opus stays the final fallback.</li> <li>The <code>fable-availability.js</code> SessionStart hook and its cache stay in place for visibility (informational only, no longer acted on) -- easy to re-enable if Fable's track record improves.</li> <li>MODEL-POLICY.md (both copies) documents the policy decision and its reasoning.</li> <li>Fixed a doc-currency gap the 0.43.0 doc sweep missed: <code>orchestration/SKILL.md</code> still described the OLD single <code>.anti-hall-progress.md</code> file as the enforced mechanism; now references the per-session path.</li> </ul> <p>2 tests updated to assert the new behavior (Fable never attempted); suite 688 (686 pass / 0 fail / 2 skip).</p> <h2 id="0431">0.43.1<a class="headerlink" href="#0431" title="Permanent link">¶</a></h2> <p><strong>History-entry writes now delegate to a cheap model; fixed a stale path left by 0.43.0.</strong></p> <ul> <li><code>verify-first-full.js</code>'s always-injected orchestration rule B previously told the coordinator to compose and append history entries itself (an expensive-model tokens spent on a mechanical write). Now it says to delegate the write to a cheap model (Haiku): hand it the cause/fix/ verification facts, let it compose and append the entry.</li> <li>The same instruction still referenced the OLD flat-file path (<code>.anti-hall-history.md</code>) from before 0.43.0's per-session restructuring -- missed by that release's doc sweep since it's a hardcoded string, not a doc file. Now points at <code>.anti-hall/history/<date>/<session-id>.md</code>.</li> </ul> <p>Suite 688 (686 pass / 0 fail / 2 skip), no test/code changes beyond the instruction string.</p> <h2 id="0430">0.43.0<a class="headerlink" href="#0430" title="Permanent link">¶</a></h2> <p><strong>Fable-5 availability flag (auto-detected, gated) + collision-free per-session progress/history.</strong></p> <h3 id="new-automatic-fable-5-detection--no-manual-reconsider-this-seat-needed">New: automatic Fable-5 detection — no manual "reconsider this seat" needed<a class="headerlink" href="#new-automatic-fable-5-detection--no-manual-reconsider-this-seat-needed" title="Permanent link">¶</a></h3> <p><code>fable-availability.js</code> (new SessionStart hook) reads <code>~/.claude.json</code>'s <code>modelAccessCache</code>/ <code>additionalModelOptionsCache</code> -- the exact same data Claude Code's own <code>/model</code> selector renders from -- ONCE per session (never re-probed every turn), fail-open, silent unless Fable 5 is actually available. When it is, the flag threads into ship-it/deadly-loop Workflow invocations via <code>args.fableAvailable</code>, and the Reviewer seat's fallback chain automatically extends to Fable 5 → Sonnet 5 → Opus (previously Sonnet 5 → Opus). No live API probe as the primary path (many Claude Code sessions have no <code>ANTHROPIC_API_KEY</code> to probe with); this reuses Claude Code's own already-maintained entitlement cache.</p> <h3 id="new-per-session-progresshistory-files--collision-free-across-concurrent-sessions">New: per-session progress/history files — collision-free across concurrent sessions<a class="headerlink" href="#new-per-session-progresshistory-files--collision-free-across-concurrent-sessions" title="Permanent link">¶</a></h3> <p><code>.anti-hall-progress.md</code> and the fix/history ledger used to be single shared files at the repo root -- two Claude Code sessions running concurrently on the same project could clobber each other's writes. Now each session gets its own file: <code>.anti-hall/progress/<date>/<session-id>.md</code> and <code>.anti-hall/history/<date>/<session-id>.md</code>. <code>tasklist-guard.js</code>'s freshness check now targets the current session's own path (session_id + UTC date, sanitized). Both get a running <code>INDEX.md</code> maintained via single-line atomic appends only (<code>fs.appendFileSync</code> with the <code>'a'</code> flag, idempotent, never read-modify-rewrite -- safe even if two sessions finish at the same moment). The old root-level ledgers are preserved untouched (<code>.anti-hall-history.md</code> also copied to <code>.anti-hall/history/legacy/pre-2026-07-01.md</code>); local <code>CLAUDE.md</code> now points at both new <code>INDEX.md</code> files as the entry point for finding prior session work.</p> <h3 id="new-progress-file-pruning--datedtimed-auto-archived-into-history">New: progress-file pruning — dated+timed, auto-archived into history<a class="headerlink" href="#new-progress-file-pruning--datedtimed-auto-archived-into-history" title="Permanent link">¶</a></h3> <p>Per-session progress files solve the single-huge-file problem but would otherwise accumulate unboundedly (one small file per session, forever). <code>progress-prune.js</code> (new SessionStart hook, per-cwd 24h-throttled) archives stale ones automatically: a newly-created progress file now carries a <code><!-- session: ... | started: <ISO-8601 UTC> --></code> header (dated <em>and</em> timed, human- readable); today's UTC date-folder is never touched; a past-date file is only pruned once its mtime is more than 6 hours stale (so a session still running across a midnight boundary is never touched mid-flight); pruning always appends the file's full content to that session's own history ledger under an "Archived progress" heading <em>before</em> deleting it — never deletes if the archive append fails, so no data is ever lost, only relocated.</p> <h3 id="hardening">Hardening<a class="headerlink" href="#hardening" title="Permanent link">¶</a></h3> <ul> <li>Fixed a real gap the adversarial verify step caught: the history-side index maintenance was dead code (<code>kind !== 'progress'</code> guard) despite being part of the spec -- now both progress and history are mechanically indexed, existence-triggered for history (no per-turn freshness concept applies there) so nothing depends on the agent remembering a manual bookkeeping step.</li> <li><code>ship-it/SKILL.md</code>'s Reviewer-seat prose updated to describe the automatic fallback chain instead of the stale "if Fable returns, reconsider this seat" note.</li> </ul> <p>17 new tests (fable-availability.test.js, ship-it-workflow.test.js, progress-prune.test.js, + additions to deadly-loop-workflow.test.js and tasklist-guard.test.js); suite 688 (686 pass / 0 fail / 2 skip).</p> <h2 id="0421">0.42.1<a class="headerlink" href="#0421" title="Permanent link">¶</a></h2> <p><strong>Bug-fix patch: root-causes the recurring ubuntu/node22 statusline flake and hardens <code>harvest-debt.js</code> / <code>eval/rescore.js</code>.</strong></p> <h3 id="fixed-statusline-base-command-timeout-flake">Fixed: statusline base-command timeout flake<a class="headerlink" href="#fixed-statusline-base-command-timeout-flake" title="Permanent link">¶</a></h3> <p><code>runBaseCommand</code>'s <code>spawnSync</code> timeout raised from 3000ms to 10000ms. Root-caused via code read (not a re-run guess) to CI runner contention on the <code>ubuntu-latest</code>/node22 job specifically; the flake had recurred 3 times across prior releases.</p> <h3 id="hardened-harvest-debtjs">Hardened: <code>harvest-debt.js</code><a class="headerlink" href="#hardened-harvest-debtjs" title="Permanent link">¶</a></h3> <ul> <li>Fixed a multi-marker-per-line drop (only the first <code>// anti-hall:</code> marker on a line was picked up).</li> <li>Fixed a comment-closer (<code>--></code>, <code>*/</code>) leaking into the parsed <code>when</code> field.</li> <li>Oversized files are now skipped, not silently truncated.</li> <li>Directory walk converted from recursive to iterative (avoids stack-depth risk on deep trees).</li> <li>Added tests covering <code><!-- --></code> and <code>--</code> marker styles.</li> </ul> <h3 id="hardened-evalrescorejs">Hardened: <code>eval/rescore.js</code><a class="headerlink" href="#hardened-evalrescorejs" title="Permanent link">¶</a></h3> <ul> <li><code>--selftest</code> now warns (instead of silently disagreeing) on a non-boolean <code>fabricated</code> value, a missing <code>condition</code>, or a duplicate <code>task_id</code> in the aggregate.</li> <li>Removed a duplicate <code>'use strict'</code> directive.</li> <li>Added missing selftest branch tests for the above.</li> </ul> <p>Full suite: 664 tests, 662 pass, 2 skipped, 0 fail.</p> <h2 id="0420">0.42.0<a class="headerlink" href="#0420" title="Permanent link">¶</a></h2> <p><strong>Sonnet 5 routing update + <code>KB-sonnet-5.md</code>; retires the Fable-disabled temp patch.</strong></p> <h3 id="new-docskb-sonnet-5md--model-routing-kb-claude--codex">New: <code>docs/KB-sonnet-5.md</code> — model routing KB (Claude + Codex)<a class="headerlink" href="#new-docskb-sonnet-5md--model-routing-kb-claude--codex" title="Permanent link">¶</a></h3> <p>Benchmark tables for Opus 4.8 / Sonnet 5 / Haiku 4.5 <strong>and</strong> the parallel Codex table (gpt-5.5 / gpt-5.4 / gpt-5.4-mini), effort-tier behavior, pricing, a task→model decision matrix, switch thresholds, the anti-hall seat routing, and cross-platform equivalence (gpt-5.5↔Opus, gpt-5.4↔Sonnet 5, gpt-5.4-mini↔Haiku). 17 sources (2 official Anthropic + 3 official OpenAI); benchmark figures are directional (system-card PDFs unparsed) and source-tagged. Indexed in <code>docs/KB.md</code> + <code>llms.txt</code>.</p> <h3 id="routing-sonnet-5-lands-fable-temp-retired">Routing: Sonnet 5 lands; Fable-temp retired<a class="headerlink" href="#routing-sonnet-5-lands-fable-temp-retired" title="Permanent link">¶</a></h3> <p>The <code>sonnet</code> tier token now resolves to <code>claude-sonnet-5</code>. MODEL-POLICY (both byte-identical copies) + the ship-it/deadly-loop workflow seats updated: <strong>implementation → Codex primary, Sonnet 5 failover</strong>; <strong>deadly-loop Reviewer + M/secondary planning → Sonnet 5 @xhigh</strong> (Auditor stays Opus @high, Critic stays Codex); Opus keeps top-level planning / root-cause / deep-debug. New rules: (1) the code implementer and its correctness reviewer are <strong>always different models</strong> (cross-model, no self-review); (2) Codex is the primary implementer (own limit → conserves the Claude bucket), failover to Sonnet 5 on unavailable/rate-limited — backoff, never retry-loop; (3) <strong>never run Sonnet 5 at <code>max</code> inside a loop</strong> (TTFT ~163s). The long-standing <code>TEMP(fable-disabled)</code> patch on the Reviewer seat is <strong>removed</strong> (Sonnet 5 fills it; a one-line "if Fable returns, reconsider" pointer remains).</p> <h3 id="limit-conservation-main-model-downshift-directive">Limit-conservation: main-model downshift directive<a class="headerlink" href="#limit-conservation-main-model-downshift-directive" title="Permanent link">¶</a></h3> <p>When conservation is active, <code>limit-conserve-inject.js</code> now advises downshifting the <strong>main coordinator</strong> off the flagship (Claude Opus → Sonnet 5; Codex gpt-5.5 → gpt-5.4) to preserve the flagship weekly bucket — <strong>but only to a 1M-context target</strong> (never gpt-5.4-mini / 400k), so no context is lost. The flagship stays for delegated hard seats + escalation. Conditional advisory (the model isn't exposed in the UserPromptSubmit payload); the agent surfaces it (it can't self-<code>/model</code>). Codex mirror in the <code>anti-hall-context-conserve</code> skill.</p> <h3 id="dual-platform-parity">Dual-platform parity<a class="headerlink" href="#dual-platform-parity" title="Permanent link">¶</a></h3> <p>Committed the standing <strong>parity mandate</strong> to <code>AGENTS.md</code>: every plan/work covers both platform variations (Claude + Codex) AND both orchestration layers (OMC ↔ OMX); model-routing artifacts get a Claude table AND a Codex table. Codex <code>anti-hall-model-policy</code> skill updated with the Claude-side mapping.</p> <h2 id="0411">0.41.1<a class="headerlink" href="#0411" title="Permanent link">¶</a></h2> <p><strong>CI fix for Codex installer tests.</strong></p> <ul> <li>Fixed the new Codex installer regression test to accept both POSIX and Windows path separators when checking generated hook commands. No runtime behavior change.</li> </ul> <h2 id="0410">0.41.0<a class="headerlink" href="#0410" title="Permanent link">¶</a></h2> <p><strong>Codex/OMX port without disturbing the Claude plugin surface.</strong></p> <h3 id="new-codex-native-plugin-layer">New: Codex-native plugin layer<a class="headerlink" href="#new-codex-native-plugin-layer" title="Permanent link">¶</a></h3> <ul> <li>Added <code>plugins/anti-hall/.codex-plugin/plugin.json</code> and <code>plugins/anti-hall/codex/install-codex.js</code>. The installer writes Codex <code>.codex/hooks.json</code> plus <code>[features].hooks = true</code> for the supported Codex hook subset.</li> <li>Added plugin-scoped Codex hooks at <code>plugins/anti-hall/codex/hooks/hooks.json</code> using <code>${PLUGIN_ROOT}</code> commands, matching installed Codex plugin examples.</li> <li>Added Codex repo marketplace compatibility via <code>.agents/plugins/marketplace.json</code>, pointing at <code>./plugins/anti-hall</code> with the official local marketplace shape.</li> <li>Added project/global install and dry-run modes for Codex activation. The installer preserves unrelated hook groups and replaces stale anti-hall hook groups.</li> <li>Added regression coverage in <code>tests/codex/install-codex.test.js</code> for dry-run behavior, hook/config writing, and merge behavior.</li> </ul> <h3 id="new-codex-skills-and-omx-workflow-mapping">New: Codex skills and OMX workflow mapping<a class="headerlink" href="#new-codex-skills-and-omx-workflow-mapping" title="Permanent link">¶</a></h3> <ul> <li>Added Codex-native skills for activate, doctor, update, root-cause, orchestration, deadly-loop, ship-it, model policy, context-conserve, feature-launch, OMX integration, OMC integration, statusline install, flutter-debug, simplify, and debt.</li> <li><code>anti-hall-context-conserve</code> ports the limit/context conservation behavior to Codex model routing and output hygiene.</li> <li><code>anti-hall-feature-launch</code> replaces the removed GSD path with a Codex/OMX planning protocol, graphify-first setup, <code>gpt-5.5</code> debate gates, phased execution, and launch verification.</li> </ul> <h3 id="new-codex-kb-equivalents">New: Codex KB equivalents<a class="headerlink" href="#new-codex-kb-equivalents" title="Permanent link">¶</a></h3> <ul> <li>Added <code>docs/CODEX-KB-MIGRATION-MAP.md</code> to classify existing KBs into Claude-specific, Codex/agnostic, and historical buckets.</li> <li>Added source-audited Codex KBs: <code>docs/KB-codex-platform-hooks-plugins.md</code>, <code>docs/KB-codex-workflow-orchestration.md</code>, and <code>docs/KB-omx.md</code>. Each has at least 10 sources and at least 2 official OpenAI sources.</li> <li>Updated <code>docs/KB.md</code> and <code>llms.txt</code> to include the Codex KBs.</li> </ul> <h3 id="codex-parity-boundary">Codex parity boundary<a class="headerlink" href="#codex-parity-boundary" title="Permanent link">¶</a></h3> <p>Current official Codex docs expose more hook events than the first migration note used for the initial port, including edit and subagent lifecycle hook names. This release intentionally registers only the Codex hook subset whose anti-hall payload contracts are currently adapted/tested: SessionStart, UserPromptSubmit, Bash PreToolUse, and Stop. Edit-time guards (<code>api-guard</code>, <code>ship-it-guard</code>) and lifecycle/compaction hooks are tracked as documented-but-not-yet-adapted Codex parity work until Codex payload adapters and tests prove them. Claude Workflow JS remains non-portable; use Codex skills/native subagents/OMX/scripts instead.</p> <p>Codex/OMX statusline parity is explicitly bounded: Claude Code supports a command-backed <code>statusLine</code>, so anti-hall can append the <code>AH: Vx.y.z</code> chip there. Codex <code>[tui].status_line</code> is documented as built-in footer item IDs only, so the Codex port documents the limitation rather than injecting an unsupported custom item.</p> <h3 id="claude-compatibility">Claude compatibility<a class="headerlink" href="#claude-compatibility" title="Permanent link">¶</a></h3> <p>The Claude manifest remains present and versioned, and the Claude hook/skill files remain in their existing locations. The Codex port lives in separate <code>codex/</code> and <code>.codex-plugin/</code> paths.</p> <h2 id="0400">0.40.0<a class="headerlink" href="#0400" title="Permanent link">¶</a></h2> <p><strong>Ponytail-derived heavier features: two new skills (<code>simplify</code>, <code>debt</code>) + a zero-API eval rescorer.</strong></p> <h3 id="new-skill-anti-hallsimplify--measured-behavior-preserving-simplification">New skill: <code>/anti-hall:simplify</code> — measured, behavior-preserving simplification<a class="headerlink" href="#new-skill-anti-hallsimplify--measured-behavior-preserving-simplification" title="Permanent link">¶</a></h3> <p>A harvest-then-prove pass over recently-changed (or named) code. Each finding gets exactly one tag — <code>delete:</code> (dead code), <code>stdlib:</code> (reinvented stdlib), <code>native:</code> (reimplemented builtin), <code>yagni:</code> (premature generality), <code>shrink:</code> (verbose equivalent), <code>slop:</code> (AI-slop filler) — the safe set is applied, the SAME tests are re-run, and the result is scored as a single <code>net: -N lines</code>. Crucially the score is the <strong>measured</strong> post-apply diff delta (<code>git diff --shortstat</code>), never a projected "you'll save ~X" estimate — that's the exact unverifiable saved-X claim verify-first rule 10 forbids. Behavior-preserving by contract: anything that removes a capability is declined as a scope change, not applied.</p> <h3 id="new-skill-anti-halldebt--a-register-for-deliberate-budgeted-debt">New skill: <code>/anti-hall:debt</code> — a register for <em>deliberate</em>, budgeted debt<a class="headerlink" href="#new-skill-anti-halldebt--a-register-for-deliberate-budgeted-debt" title="Permanent link">¶</a></h3> <p>Introduces the <code>// anti-hall: <ceiling>,<when></code> marker — a budgeted, harvestable alternative to vague TODOs. <code><ceiling></code> is the limit you consciously accepted (e.g. <code>30 lines</code>, <code>O(n^2)</code>); <code><when></code> is the concrete payback trigger (e.g. <code>when >3 callers</code>). New <code>plugins/anti-hall/scripts/harvest-debt.js</code> (pure Node, comment-syntax-agnostic across <code>//</code> <code>#</code> <code>--</code> <code>/* */</code> <code><!-- --></code>) greps the tree, parses each marker, and flags <strong>rot-risk</strong> (<code>no-trigger</code>) when a marker has no <code><when></code> <em>or</em> sits in code git-untouched past a staleness threshold (default 90 days; fail-open when git is absent). Explicitly <strong>not</strong> a license to skip real work — lazy TODO/stub patterns remain blockers; the marker is the narrow, defensible exception.</p> <h3 id="new-evalrescorejs--recompute-eval-stats-with-zero-api-calls">New: <code>eval/rescore.js</code> — recompute eval stats with zero API calls<a class="headerlink" href="#new-evalrescorejs--recompute-eval-stats-with-zero-api-calls" title="Permanent link">¶</a></h3> <p>Recomputes the summary block (protocol/baseline fabrication rates, delta, per-task differences) straight from saved <code>records[].fabricated</code> — no answer calls <em>and</em> no judge calls (distinct from <code>grade.js</code>, which re-calls the judge). Aggregates across multiple result files. <code>--selftest</code> is an integrity gate: it re-derives each file's summary from its own records and fails (exit 1) on any schema violation or count/rate mismatch, catching hand-edited or corrupted result files. 21 new tests (<code>tests/eval/rescore.test.js</code>, <code>tests/hooks/harvest-debt.test.js</code>); suite 646 (644 pass / 2 skip).</p> <h2 id="0390">0.39.0<a class="headerlink" href="#0390" title="Permanent link">¶</a></h2> <p><strong>Subagents now receive the verify-first Iron Law (new SubagentStart hook) + guard/hardening refinements.</strong></p> <h3 id="new-subagentstart-re-injection--discipline-finally-reaches-subagents">New: SubagentStart re-injection — discipline finally reaches subagents<a class="headerlink" href="#new-subagentstart-re-injection--discipline-finally-reaches-subagents" title="Permanent link">¶</a></h3> <p><code>verify-first-full.js</code> was SessionStart-only, so every Task-spawned subagent ran verify-first-UNAWARE. New <code>verify-first-subagent.js</code> (SubagentStart hook) injects the Iron Law + rationalization table + positive rules + scope-fidelity into each spawned subagent — but DELIBERATELY omits the orchestration "delegate everything" block (subagents are workers; re-injecting it would recreate deep nesting). The shared core is extracted to <code>verify-first-core.js</code> (one source of truth for both hooks, no drift). 9 tests. (SubagentStart confirmed as a real Claude Code event via docs/KB-claude-codex.md §1.1.)</p> <h3 id="model-routing-guard-researchexplore-nudge-no-longer-false-positives-on-write-tasks">model-routing-guard: research→Explore nudge no longer false-positives on write tasks<a class="headerlink" href="#model-routing-guard-researchexplore-nudge-no-longer-false-positives-on-write-tasks" title="Permanent link">¶</a></h3> <p>The 0.37.0 anti-nesting advisory nudged any research-shaped <code>general-purpose</code> spawn toward Explore — but Explore can't write, so release/commit/build agents were wrongly nudged. Now suppressed when the spawn has write/execute signals.</p> <h3 id="hardening-ponytail-derived">Hardening (ponytail-derived)<a class="headerlink" href="#hardening-ponytail-derived" title="Permanent link">¶</a></h3> <ul> <li><code>install-statusline.js</code>: <code>isShellSafe()</code> allowlist guards paths before embedding them in a settings.json command (unsafe → manual-setup fallback).</li> <li>verify-first rule 10: never display a per-run "you saved X tokens/lines" number — the unbuilt baseline was never run; cite a benchmark median with provenance, or say it's unmeasured.</li> </ul> <h2 id="0382">0.38.2<a class="headerlink" href="#0382" title="Permanent link">¶</a></h2> <p><strong>Fix: <code>agentsRunning()</code> was inert — the parallel-orchestration guard exemptions never fired.</strong></p> <p><code>agentsRunning()</code> (consumed by task-guard, tasklist-guard, task-tracker) reads <code>~/.anti-hall/agents/*.json</code> heartbeats, but NOTHING wrote them, so it always returned false — the 0.36.1 multiple-in_progress exemption and the idle-neglect "no agents running" signal were no-ops, and the Stop guards nagged even while background agents were actively working. Fix: - <code>phase-tracker.js</code> now writes a rolling heartbeat <code>~/.anti-hall/agents/recent-spawn.json</code> (<code>{ts}</code>) on every Agent/Task spawn (fail-open; the existing <code>agent-spawns.log</code> write is unchanged). <code>agentsRunning()</code> returns true for ~20 min after the most recent spawn = active orchestration, so the agent-aware exemptions ACTUALLY fire now. (A single agent running >20 min with no new spawn is a known limitation — a per-subagent refresh is a future enhancement.) - <code>task-guard.js</code>: the generic "open tasks remain at Stop" block now suppresses when agents are live (it was firing on in_progress/owned tasks regardless of agent state). Genuinely-neglected work (open tasks + no live agent) still blocks. - +6 tests.</p> <h2 id="0381">0.38.1<a class="headerlink" href="#0381" title="Permanent link">¶</a></h2> <p>Test + docs maintenance: de-coupled the statusline minor/major-ahead test fixtures from the hardcoded version (now derived from plugin.json so they never go stale on a bump); refreshed stale test-count references in README/llms.txt/KB.md.</p> <h2 id="0380">0.38.0<a class="headerlink" href="#0380" title="Permanent link">¶</a></h2> <p><strong>Limit-conservation mode + consolidated statusline merge + OMC as recommended optional dependency.</strong></p> <h3 id="new-limit-conserve-inject-hook-userpromptsubmit--limit-conservejs-helper">New: <code>limit-conserve-inject</code> hook (UserPromptSubmit) + <code>limit-conserve.js</code> helper<a class="headerlink" href="#new-limit-conserve-inject-hook-userpromptsubmit--limit-conservejs-helper" title="Permanent link">¶</a></h3> <p><code>limit-conserve-inject.js</code> — a UserPromptSubmit hook that injects a token-conservation nudge when the session context usage is at or above a threshold. Env knobs:</p> <ul> <li><code>ANTIHALL_LIMIT_CONSERVE</code> — <code>auto</code> (default), <code>on</code>, or <code>off</code>. In <code>auto</code> mode the hook reads the OMC usage cache (<code>~/.anti-hall/omc-usage-cache.json</code>) to detect the current context percentage; <code>on</code> forces the nudge unconditionally; <code>off</code> disables it.</li> <li><code>ANTIHALL_LIMIT_THRESHOLD</code> — integer percentage (default <code>85</code>). The nudge fires only when detected context usage ≥ this value.</li> </ul> <p>Auto mode requires an OMC installation that populates the usage cache; without it the hook operates in manual mode (<code>on</code>/<code>off</code> only) and auto silently behaves as off. Skip-guard hatch: <code>limit-conserve</code>. UserPromptSubmit hooks 2 → 3.</p> <p><code>limit-conserve.js</code> is the shared helper consumed by the hook (reads the OMC usage cache, applies threshold logic). It is not itself a hook.</p> <h3 id="statusline-consolidated-merge-mode---consolidate">Statusline: consolidated merge mode (<code>--consolidate</code>)<a class="headerlink" href="#statusline-consolidated-merge-mode---consolidate" title="Permanent link">¶</a></h3> <p><code>install-statusline --consolidate</code> merges the anti-hall statusline with an existing statusline (e.g., the OMC HUD) instead of replacing it. The existing base <code>statusLine</code> value is read from <code>ANTIHALL_STATUSLINE_BASE</code> (env) or detected from the current settings; the anti-hall bar is appended as an additional component. The resolved base is persisted to <code>~/.anti-hall/consolidated-base.json</code> so subsequent sessions can restore it without re-reading the env var.</p> <p>New env knob: <code>ANTIHALL_STATUSLINE_BASE</code> — explicitly sets the base statusline expression when using consolidated mode.</p> <h3 id="omc-recommended-optional-dependency">OMC: recommended optional dependency<a class="headerlink" href="#omc-recommended-optional-dependency" title="Permanent link">¶</a></h3> <p>oh-my-claudecode (OMC) is now explicitly documented as a <strong>recommended optional</strong> dependency. Anti-hall is and remains fully standalone without it. Two features unlock automatic behavior when OMC is installed:</p> <ol> <li><code>limit-conserve</code> auto mode — reads the OMC usage cache to detect the live context percentage.</li> <li>Consolidated statusline mode — the version chip and base-merge work with the OMC HUD out of the box.</li> </ol> <p>Without OMC, both features fall back to manual/off behavior; no errors, no breaking change.</p> <h3 id="temporary-fable-removed-from-all-spawn-sites">Temporary: Fable removed from all spawn sites<a class="headerlink" href="#temporary-fable-removed-from-all-spawn-sites" title="Permanent link">¶</a></h3> <ul> <li>Temporary: Fable removed from all spawn sites (Anthropic disabled it) — Reviewer/flagship-Claude seats in <code>deadly-loop</code> and <code>ship-it</code> run on Opus until re-enabled. All changed spawn sites are marked <code>TEMP(fable-disabled 2026-06-29)</code> for easy grep-revert when Fable is restored.</li> </ul> <h2 id="0370">0.37.0<a class="headerlink" href="#0370" title="Permanent link">¶</a></h2> <p><strong>Version awareness + priority-aware guards + anti-nesting backstop + cmux/OMC KBs.</strong></p> <h3 id="new-sessionstart-version-alert-hook">New: SessionStart <code>version-alert</code> hook<a class="headerlink" href="#new-sessionstart-version-alert-hook" title="Permanent link">¶</a></h3> <p><code>version-alert.js</code> (+ detached <code>version-alert-refresh.js</code>): a NON-BLOCKING SessionStart check that alerts when a newer anti-hall version is available. Reads the running version vs a cached latest (<code>~/.anti-hall/version-check.json</code>); if behind, emits a one-line "vX available — /anti-hall:update". When the cache is absent/stale it spawns a DETACHED, unref'd <code>git ls-remote --tags</code> refresh and stays silent that session — SessionStart never blocks or does synchronous network. Off-switch <code>ANTIHALL_VERSION_ALERT=off</code>; skip-guard hatch. SessionStart hooks 2 -> 3. 8 tests.</p> <h3 id="statusline-version-chip-with-update-indicator">Statusline: version chip with update indicator<a class="headerlink" href="#statusline-version-chip-with-update-indicator" title="Permanent link">¶</a></h3> <p>The statusline shows <code>AH: Vx.y.z</code> between the cost chip and the email segment. When the version-check cache shows a newer release: <code>★ AH: …</code> in YELLOW for a new MINOR, RED for a new MAJOR; plain dim otherwise (fail-open if no cache).</p> <h3 id="priority-aware-stop-guards-less-nag-noise">Priority-aware Stop guards (less nag noise)<a class="headerlink" href="#priority-aware-stop-guards-less-nag-noise" title="Permanent link">¶</a></h3> <p><code>task-guard</code> (idle-neglect) and <code>tasklist-guard</code> (multi-in_progress) now read each task's <code>metadata.priority</code> and only chase ACTIONABLE P0/P1 work — a backlog of P2/deferred tasks no longer triggers a nudge, and P2 in_progress doesn't count toward the stale-multi check. Missing/garbage priority is treated as actionable (P1) so a real high-priority task is never under-nagged. Encodes "priority = check the top first/more often, never neglect the rest, don't nag about backlog."</p> <h3 id="anti-nesting-backstop-in-model-routing-guard">Anti-nesting backstop in <code>model-routing-guard</code><a class="headerlink" href="#anti-nesting-backstop-in-model-routing-guard" title="Permanent link">¶</a></h3> <p>A new advisory fires when research/read-only-shaped work is spawned as <code>general-purpose</code> (carries the Agent tool, can recurse) — nudging to use the <code>Explore</code> agent type (has WebSearch/WebFetch but NO Agent tool, so it structurally cannot nest). Advisory only; the structural complement to rule M's anti-deep-nesting discipline.</p> <h3 id="new-kbs">New KBs<a class="headerlink" href="#new-kbs" title="Permanent link">¶</a></h3> <p><code>docs/KB-cmux.md</code> (cmux terminal multiplexer) + <code>docs/KB-omc.md</code> (oh-my-claudecode + the cmux+OMC+Claude stack). Both agnostic; registered in llms.txt + docs/KB.md.</p> <h2 id="0361">0.36.1<a class="headerlink" href="#0361" title="Permanent link">¶</a></h2> <p><strong>Fix: <code>tasklist-guard</code> multiple-in_progress false-positive was crippling parallel orchestration.</strong></p> <p>The Stop-hook <code>hasStaleInProgress</code> sub-cause fired on ANY 2+ in_progress tasks and told the agent to "keep one task in_progress at a time" — i.e. to <strong>serialize</strong>, the exact opposite of the parallel fan-out anti-hall itself promotes. With background agents legitimately working N tasks at once, it nagged on every Stop and pushed agents to collapse their parallel work. Fix:</p> <ul> <li><strong>Exempt the multi-in_progress block when a live background-agent heartbeat exists</strong> — reuse <code>agentsRunning()</code> (<code>~/.anti-hall/agents/*.json</code> fresh within 20 min). Multiple in_progress is CORRECT while agents are live; only genuinely STALLED in_progress (no live agent) now flags.</li> <li><strong>Rewrite the message</strong> from "keep one in_progress at a time" (serialize) to "dispatch a background agent for EACH so they run in PARALLEL, or set idle ones back to pending; priority = check it first and more often, never pause the rest."</li> <li>Fail-open: an <code>agentsRunning()</code> error reads as not-running (can only permit a nudge, never wrongly silence a real stall). +2 tests.</li> </ul> <h2 id="0360">0.36.0<a class="headerlink" href="#0360" title="Permanent link">¶</a></h2> <p><strong>Codex everyday-routing + Workflow model-distribution discipline + new <code>codex-nudge</code> advisory hook.</strong></p> <h3 id="new-hook-codex-nudge-stop-advisory">New hook: <code>codex-nudge</code> (Stop, advisory)<a class="headerlink" href="#new-hook-codex-nudge-stop-advisory" title="Permanent link">¶</a></h3> <p>A loop-safe, fail-open Stop hook that nudges ONCE per session to get an independent OpenAI-Codex second opinion when the session shipped a substantial code change (>= <code>ANTIHALL_CODEX_NUDGE_MIN</code>, default 3 code-file edits) with no Codex review (no <code>codex:codex-rescue</code> spawn / codex skill). Mechanizes the everyday-routing policy for the MAIN agent (the deadly-loop/ship-it Critic seat already covers those skills). Codex is the cross-model correctness reviewer (off-by-one, races, subtle bugs); Opus keeps architecture/design review. Deduped on the edited-file signature, hard cap 2 nudges/session. Off-switch <code>ANTIHALL_CODEX_NUDGE=off</code>; skip-guard hatch <code>codex-nudge</code>. 10 tests + doctor smoke test. Stop hooks: 5 -> 6.</p> <h3 id="everyday-routing--workflow-model-distribution-discipline">Everyday-routing + Workflow model-distribution discipline<a class="headerlink" href="#everyday-routing--workflow-model-distribution-discipline" title="Permanent link">¶</a></h3> <p><code>verify-first-full.js</code> gains orchestration rules M (shallow+wide; a subagent is a worker that does not re-delegate; lift 3+ nested/parallel spawns into a deterministic Workflow; Explore for read-only) and N (distribute models per seat — implementation->sonnet, correctness/verify review->Codex, planning/architecture->opus; NEVER an all-Opus fan-out; the model-routing guard does NOT police models inside a workflow review fan-out, so it is an authoring responsibility), a <code>model-routing</code> ALWAYS-APPLY bullet, and two per-turn nudges. Reconciled against a 13-source Codex-vs-Opus coding KB (<code>docs/KB-codex-vs-opus-coding.md</code>) + the Workflow KB (<code>docs/KB-claude-workflow-orchestration.md</code>).</p> <h3 id="fix-ship-it-build-seats-set-an-explicit-model">Fix: ship-it build seats set an explicit model<a class="headerlink" href="#fix-ship-it-build-seats-set-an-explicit-model" title="Permanent link">¶</a></h3> <p><code>ship-it.workflow.js</code> implementation seats omitted <code>model</code> — under strict model-routing (default since 0.35.0) a mechanical omitted-model spawn is BLOCKED, so the build could self-block, and an omitted model inherits the flagship orchestrator. Build seats now set <code>model: phase.model || 'sonnet'</code> (implementation -> Sonnet per the KB; override per phase).</p> <h3 id="model-policy-everyday-routing-section">MODEL-POLICY: everyday routing section<a class="headerlink" href="#model-policy-everyday-routing-section" title="Permanent link">¶</a></h3> <p><code>skills/MODEL-POLICY.md</code> (+ the deadly-loop copy) gains an "Everyday agent routing" section: Codex = second-opinion/correctness review (always) + bounded code-apply/terminal/migration; Opus = planning/architecture/design + design-level review; Sonnet = implementation; Haiku = trivial/nav. The TRIO debate roster is unchanged.</p> <h2 id="0351">0.35.1<a class="headerlink" href="#0351" title="Permanent link">¶</a></h2> <p><strong>ship-it Workflow Fable→Opus availability fallback (bug fix) + Workflow orchestration KB.</strong></p> <ul> <li><strong>Fixed: the ship-it deadly-loop gate broke when Fable was unavailable.</strong> The Reviewer seat in <code>skills/ship-it/references/ship-it.workflow.js</code> hardcoded <code>model: 'fable'</code> with no guard, so when the <code>fable</code> tier token is disabled at the account level the Reviewer spawn died and the whole per-phase gate could not complete. <code>MODEL-POLICY.md</code> <em>documented</em> a fable✗→Opus fallback matrix, but it was never wired into the workflow script (only the Codex seat had a runtime probe). Added <code>reviewerAgent(p)</code>: it attempts the latest flagship and, on a terminal <code>null</code> return (the Workflow contract's unavailable-model signal), falls back to an Opus Reviewer — or short-circuits straight to Opus when the coordinator passes <code>args.fableAvailable === false</code> (no wasted spawn). Floor stays Opus; never a cheaper model. Validated: 3 branches (Fable healthy / Fable null / coordinator-off), 8/8 assertions green against the real shipped function. <code>ship-it/SKILL.md</code> snippet updated to show the wrapper.</li> <li><strong>New: <code>docs/KB-claude-workflow-orchestration.md</code></strong> — a 14-source (8 official Anthropic) knowledge base on programmatic multi-agent orchestration (the <code>Workflow</code> tool): what it is vs ad-hoc subagent spawns, when to use it vs a single/shallow agent, the core patterns (orchestrator-worker, pipeline/parallel, map-reduce, loop-until-done, adversarial verify), the ~15× token / 90.2% / depth-nesting cost numbers, and a "how to make an agent utilize it more" section. Reconciled against in-repo live verification (Workflow runs under Opus 4.8 — it is <strong>not</strong> Fable-bound). Registered in <code>llms.txt</code> and the <code>docs/KB.md</code> doc table.</li> </ul> <h2 id="0350">0.35.0<a class="headerlink" href="#0350" title="Permanent link">¶</a></h2> <p><strong>Strict-by-default model routing + <code>anti-hall:activate</code> first-run setup skill.</strong></p> <h3 id="model-routing-guard-strict-is-now-the-default">model-routing-guard: strict is now the default<a class="headerlink" href="#model-routing-guard-strict-is-now-the-default" title="Permanent link">¶</a></h3> <p><code>model-routing-guard.js</code> row-2 behavior flipped: <strong>strict mode is the default</strong> as of v0.35.0. Previously strict was opt-in (<code>ANTIHALL_MODEL_ROUTING=strict</code>); now advisory is the opt-out (<code>ANTIHALL_MODEL_ROUTING=advisory</code>).</p> <ul> <li><strong>Before (≤ 0.34.1):</strong> omitted-model mechanical spawns → advisory by default; set <code>ANTIHALL_MODEL_ROUTING=strict</code> to block.</li> <li><strong>After (≥ 0.35.0):</strong> omitted-model mechanical spawns → <strong>blocked unconditionally</strong> by default; set <code>ANTIHALL_MODEL_ROUTING=advisory</code> to revert to advisory-only.</li> </ul> <p>Rationale: an omitted model silently inherits the orchestrator's model. On a flagship orchestrator this produces an all-flagship swarm with no warning and no signal — the most common and most expensive misroute. Strict is the right default; advisory opt-out covers projects where the orchestrator is verifiably cheap-modeled.</p> <p>The row-1 behavior (explicit flagship + mechanical task → block, debate-role exemption downgrades to advisory) is unchanged. The row-2 strict block message now says "default" and names <code>ANTIHALL_MODEL_ROUTING=advisory</code> as the remedy. All existing tests updated.</p> <h3 id="new-skill-anti-hallactivate">New skill: <code>anti-hall:activate</code><a class="headerlink" href="#new-skill-anti-hallactivate" title="Permanent link">¶</a></h3> <p><code>skills/activate/SKILL.md</code> — one-shot, idempotent first-run setup. User-invoked only; <strong>never</strong> auto-runs as a SessionStart side-effect (the always-on hooks need no activation). What it does:</p> <ul> <li>Checks <code>~/.claude/settings.json</code> for an existing <code>statusLine</code>. If none → installs the anti-hall statusline at user scope (delegates to <code>install-statusline.js --user</code>). If one exists (conflict) → reports it and lets the user choose: wrap as line 1 (global), install at project scope, or skip.</li> <li>Reports model-routing state: strict (default, unset) or advisory (opt-out set). No action taken — informational only.</li> <li>Writes <code>~/.anti-hall/activated.json</code> sentinel so re-runs report "already activated" and exit 0 immediately (idempotent).</li> <li>Prints a "restart Claude Code" note only when the statusline was actually changed.</li> </ul> <p>Reuses <code>install-statusline.js</code> — no reimplemented logic.</p> <h3 id="what-stays-opt-in-unchanged">What stays opt-in (unchanged)<a class="headerlink" href="#what-stays-opt-in-unchanged" title="Permanent link">¶</a></h3> <p><code>mcp-reaper</code>, <code>ANTIHALL_API_GUARD_THIRDPARTY</code>, <code>ANTIHALL_SHIPIT_GATE</code>, <code>ANTIHALL_MERGE_GATE</code>, <code>ANTIHALL_SEMANTIC_JUDGE</code> — all remain off by default; activate does not touch them. On-demand skills (deadly-loop, ship-it, etc.) unchanged.</p> <p>Skills shipped: 9 → 10.</p> <h2 id="0341">0.34.1<a class="headerlink" href="#0341" title="Permanent link">¶</a></h2> <p><strong>Honest-fix wave on the shipped <code>flutter-debug</code> agent/skill</strong> (retroactive hardening of v0.34.0).</p> <ul> <li><strong>FP1b resolution (NEGATIVE verdict, now stated honestly).</strong> The step-0 probe captured <code>flutter_driver_command</code> tap/screenshot FAILING against a plain debug app — they REQUIRE an in-app <code>enableFlutterDriverExtension()</code> before <code>runApp</code> (same invasiveness class as marionette, which is strictly richer — the only semantic input/screenshot route). The shipped degradation row, SKILL.md honest-scope, and the agent's honesty discipline now say so; the previous "tap + screenshot with NO app package" claim is removed. The <code>widget_inspector</code> tree-inspection-without-extension nuance is retained.</li> <li><strong>Private-identifier scrub + lint extension.</strong> Owner-private identifiers (the owner's Flutter app name, a configured AVD name) were scrubbed from the public probe record, the plan doc, and this CHANGELOG. The repo-agnostic lint test is extended to a denylist constant (no <code>/Users/</code> home paths, no private app/device names) and now covers the probe record, the plan doc, and the CHANGELOG 0.34.x sections — future names add one denylist entry.</li> <li><strong>doctor-conditionality + Windows-path test hardening.</strong> A subprocess test asserts the doctor's <code>flutter-debug</code> section is silent without a <code>pubspec.yaml</code> and present with one; a platform-monkeypatched test asserts BOTH <code>claude</code> CLI candidates failing on Windows ⇒ <code>cli-unavailable</code> (a manual-verify WARN), never a false FAIL.</li> <li><strong>Auto-apply pre-write safety checks.</strong> SKILL.md app-integration now skips when <code>MarionetteBinding</code> is already present, locates the entrypoint via <code>lib/main.dart</code>, and WARN-AND-STOPs (never guess-edits) on non-standard layouts or multiple <code>runApp()</code> call sites. The applied-diff visibility stays.</li> <li> <p><strong>marionette PATH-resolution warning.</strong> After a successful <code>dart pub global activate marionette_mcp</code>, preflight now verifies the wrapper actually resolves (<code>$PUB_CACHE/bin</code> ~ <code>~/.pub-cache/bin</code> or <code>$PATH</code>); if not, it WARNs with the exact manual <code>export PATH</code> fix line (fail-open).</p> </li> <li> <p><strong>Android support VERIFIED via live FP7 (marionette taps/screenshots on emulator).</strong> The live FP7 probe (2026-06-11) confirmed all 15 <code>ext.flutter.marionette.*</code> extensions registered on android_arm64; semantic tap drove a counter 0→1; screenshots returned valid PNG. Scope: one AVD / android_arm64 arch — physical device and other architectures not yet probed. All shipped Android-pending text updated to reflect the verified status with honest scope boundaries.</p> </li> <li><strong>Full E2E debug loop validated live (plant→reproduce→read→fix→hot-reload→verify).</strong> A <code>StateError</code> was planted in the scratch app, reproduced via marionette tap, <code>get_runtime_errors</code> captured "Bad state: planted bug", the fix was applied, <code>hot_reload</code> succeeded, 3 taps past the old throw point produced zero new errors, and the screenshot was verified. Closes the loop on every agent-loop step being exercised against a real AVD.</li> <li><strong><code>get_runtime_errors</code> timestamp guidance added.</strong> The tool accumulates errors since DTD connection — it does NOT reset on <code>hot_reload</code>. The agent's mandatory re-verify step now requires comparing each error's timestamp against the <code>hot_reload</code> time; only post-reload errors count as new failures. Added to the agent loop (step 7) and SKILL.md MCP-usage notes.</li> </ul> <p><strong>Transparency (house rules):</strong> v0.34.0 was tagged against a RED CI run; the corrective fix landed in commit <code>31a7451</code>. This 0.34.1 wave addresses the substance flagged in the retroactive review.</p> <h2 id="0340">0.34.0<a class="headerlink" href="#0340" title="Permanent link">¶</a></h2> <p><strong>New agent + skill: <code>/anti-hall:flutter-debug</code></strong> — drive a Flutter app in debug mode, close the fix loop via agent-controlled hot reload + visually verified UI changes.</p> <h3 id="workstream-a--agent--skill-architecture">Workstream A — Agent + Skill architecture<a class="headerlink" href="#workstream-a--agent--skill-architecture" title="Permanent link">¶</a></h3> <p><strong><code>agents/flutter-debug.md</code></strong> — anti-hall's first shipped agent. Self-contained executor persona carries the full debug-loop protocol (reproduce → read error → root-cause → fix → hot reload → <strong>visually re-verify</strong>), so direct spawns work without the skill. Frontmatter: <code>{name: flutter-debug, description: <>, model: sonnet}</code> (sonnet is the code-authoring floor — never haiku). <strong>No <code>tools</code> allowlist</strong> — MCP tool names vary by alias; a wrong allowlist bricks the agent (FP5 validates MCP reachability).</p> <p><strong><code>skills/flutter-debug/SKILL.md</code></strong> — user-facing workflow: setup orchestration, zero-setup MCP registration (scope-aware atomic <code>claude mcp add</code> per FP9/FP10), app-side marionette integration, capability tier degradation table, and escalation-report trigger. Loop protocol lives ONCE in the agent; the skill delegates to it (non-blocking coordinator).</p> <h3 id="workstream-b--mcp-strategy-zero-setup-via-scope-aware-claude-mcp-add">Workstream B — MCP strategy: zero-setup via scope-aware <code>claude mcp add</code><a class="headerlink" href="#workstream-b--mcp-strategy-zero-setup-via-scope-aware-claude-mcp-add" title="Permanent link">¶</a></h3> <p><strong>NO bundled <code>.mcp.json</code> for external servers</strong> (owner directive; duplicated processes when user already has dart/marionette registered). Each of dart + marionette is registered idempotently via FP9 flow: <code>claude mcp get <name></code> → absent everywhere ⇒ <code>claude mcp add --scope <chosen></code> (scope question asked once in main context; non-interactive = default local) → present in ANY scope ⇒ SKIP (user entries win by precedence).</p> <p><strong>Composition:</strong> official Dart MCP (reload/reload/get_runtime_errors/widget_inspector/vm_service/analyze_files) <strong>REQUIRED</strong> [2]; marionette_mcp ≥ 0.4.0 (semantic tap/enter_text/scroll_to + screenshots + logs) <strong>PRIMARY</strong> [5]; joshuayoes/ios-simulator-mcp (coordinate taps + screenshots, fallback, Flutter reliability UNVERIFIED [10]). Auto-apply: marionette <code>pubspec.yaml</code> dependency + upstream-verbatim kDebugMode init (FP6 verified; app-side, invisible to release builds).</p> <h3 id="workstream-c--preflight--doctor-integration">Workstream C — Preflight + doctor integration<a class="headerlink" href="#workstream-c--preflight--doctor-integration" title="Permanent link">¶</a></h3> <p><strong><code>scripts/preflight.js</code></strong> — pure Node ≥ 18, cross-platform, fail-open. <strong>ONE implementation, TWO entry points:</strong> exports its checks; <code>doctor.js</code> require()s and CALLS them in-process (not subprocess — matches the G test contract). Runs ONLY on Flutter projects (pubspec.yaml or flutter-debug in use). <strong>Checks:</strong> (1) <code>dart --version</code> ≥ 3.12 FULL / 3.9–3.11 WARN / <3.9 FAIL [1][3]; (2) FP9 idempotent MCP registration; (3) marionette host on PATH ≥ 0.4.0, auto-fix via <code>dart pub global activate marionette_mcp</code>; (4) iOS booted simulator + idb reachable [9][10]; (5) <strong>Android SDK explicit-path resolution</strong> (BINDING LESSON — bare-PATH probe once falsely reported "no tooling" on a machine with SDK 36 + AVD; now explicit env/path checks + <code>flutter doctor</code> cross-check [FP7]); (6) project sanity (pubspec.yaml present). Degradation table per missing capability. Honest on Android visual status (taps/screenshots PENDING FP7, now ACTIVE).</p> <p><strong>doctor.js § 6b:</strong> conditional flutter-debug section (silent in non-Flutter cwd). Calls preflight exports in-process (skipRegistration:true = read-only mode). Scope context inherited from the user's current work. Fail-open on any probe error.</p> <h3 id="workstream-d--the-debug-loop-agent-body-mirrors-kb-5--root-cause-discipline">Workstream D — The debug loop (agent body; mirrors KB §5 + root-cause discipline)<a class="headerlink" href="#workstream-d--the-debug-loop-agent-body-mirrors-kb-5--root-cause-discipline" title="Permanent link">¶</a></h3> <p>Agent executes: (0) preflight + announce tier; (1) run + DTD connect; (2) reproduce (marionette/coordinate/fallback semantic/coordinate/coordinate taps); <strong>screenshot BEFORE</strong> [5][10]; (3) read error routed by kind (exceptions → get_runtime_errors; layout → widget_inspector; prints → marionette get_logs / get_app_logs <strong>double-gated: launch_app-enabled path only</strong> [2]); (4) <strong>NO CAUSE NO FIX</strong> — root-cause with evidence (error + widget tree + code); (5) edit + verify with analyze_files [2]; (6) reload (hot_reload state-preserved, hot_restart for const/init/reset [2]); (7) <strong>MANDATORY re-verify:</strong> re-run get_runtime_errors until clean (matches official demo [1]); on visual tier, BEFORE/AFTER screenshot compare (visual MCP required for "verified", else "error-clear but visually unverified"); (8) <strong>escalation trigger:</strong> after 2 full iterations without proven root cause OR fix needs redesign → report <code>escalate: opus</code> with collected evidence (agent never respawns itself); (9) loop until clean. Report = tier + per-fix evidence (errors, screenshots).</p> <h3 id="workstream-e--step-0-probes-testsfixturesstep0-probe-record-v0340md">Workstream E — Step-0 probes (tests/fixtures/step0-probe-record-v0.34.0.md)<a class="headerlink" href="#workstream-e--step-0-probes-testsfixturesstep0-probe-record-v0340md" title="Permanent link">¶</a></h3> <ul> <li><strong>FP1 DONE:</strong> <code>flutter_driver_command</code> exposes tap + screenshot + semantic finders (ByValueKey/ByText/BySemanticsLabel…) IN-SCHEMA [2]; schema forbids guessing (widget_inspector first).</li> <li><strong>FP1b RESIDUAL:</strong> runtime vs plain debug app (no enableFlutterDriverExtension()). Gates no-package degradation. No shipped promise pre-FP1b.</li> <li><strong>FP2 UNVERIFIED:</strong> Flutter widgets in iOS a11y trees [10][9]. Grades SUPPLEMENT tier only.</li> <li><strong>FP3 EXCLUDED:</strong> mcp_flutter untested-merge warning [6]. Revisit gate = stable tag (tracked in KB staleness ledger).</li> <li><strong>FP4 CONFIRMED:</strong> lifecycle tools ABSENT from default list (disabled by default [2]). Manual <code>flutter run --print-dtd</code> stays the path. get_app_logs correction = double-gate logic.</li> <li><strong>FP5 GENERIC PASS:</strong> subagent spawns surface mcp__ tools via ToolSearch and invoke live. Architecture viable; skill-only fallback retained as contingency (executable spec: SKILL.md → CI test omitted, AC1/AC3 branches, CHANGELOG note).</li> <li><strong>FP6 DONE MATCH:</strong> regular dep, upstream-verbatim kDebugMode if/else (quoted in B). Release-safe per upstream (LogCollector optional, non-fatal).</li> <li><strong>FP7 UN-DEFERRED ACTIVE:</strong> bare-PATH false-negative corrected (SDK 36 at ~/Library/Android/sdk; a configured AVD; adb + emulator functional). Probe = boot AVD → marionette → tap/screenshot. Upstream silent on Android ⇒ no promise pre-FP7. Binding baked into preflight check 5.</li> <li><strong>FP9 CAPTURED:</strong> <code>claude mcp add</code> scope/precedence/atomicity semantics (full consequences in B registration flow; B2 shims dropped).</li> <li><strong>FP10 CAPTURED:</strong> ecosystem precedent (OMC bundles only its own server; ecc bundles 6 externals REJECTED for duplicate-server hazard). Directs decision to FP9 registration flow.</li> </ul> <h3 id="workstream-f--model-routing--escalation">Workstream F — Model routing + escalation<a class="headerlink" href="#workstream-f--model-routing--escalation" title="Permanent link">¶</a></h3> <p>Agent default <strong><code>model: sonnet</code></strong> (code-authoring floor). Escalation split: TRIGGER in agent body (D step 8, <code>escalate: opus</code>); RESPAWN via caller. SKILL.md instructs coordinator to respawn at <code>model: opus</code> on signal. Tier tokens only (v0.32.0 policy).</p> <h3 id="workstream-g--tests">Workstream G — Tests<a class="headerlink" href="#workstream-g--tests" title="Permanent link">¶</a></h3> <p><strong>41 new tests</strong> (≥19 required): agent frontmatter parse (1); SKILL.md lint (1); preflight dart full/warn/fail (3), degradation rows (7), malformed-output fail-open (2), Windows <code>.cmd</code> best-effort (1); FP9 registration flow (≥4): get-absent → add, get-present → SKIP, get-unparsable → no add, add hard-error surfaced; Android SDK-path units (2); doctor shared-checks (1); repo-agnostic lint (1). All 535 tests pass (2 skip unrelated).</p> <h3 id="workstream-h--documentation">Workstream H — Documentation<a class="headerlink" href="#workstream-h--documentation" title="Permanent link">¶</a></h3> <p><strong><code>docs/KB-flutter-claude-debug.md</code></strong> — 13 sources synthesized (KB + probe evidence cited via <code>[n]</code> = KB numbering + FP-ids). Drives every capability claim in the agent/skill/preflight (honesty contract: every promise traces to KB or probe).</p> <p><strong><code>docs/2026-06-10-v0.34.0-flutter-debug-plan.md</code></strong> — development log (rounds 1–3 converged GO×3, v5 baseline + 2026-06-11 amendments for FP9/FP10 MCP strategy + FP7 Android binding). For future context.</p> <p><strong><code>tests/fixtures/step0-probe-record-v0.34.0.md</code></strong> — raw probe captures (FP1–FP10 with dates + tooling inventory). Supports KB staleness tracking.</p> <p><strong>README / llms.txt:</strong> new agent listed; 8 skills → 9 skills (including flutter-debug); test count 461 → 535; total checks 49; doctor integration noted.</p> <p>Skills shipped: 8 → 9.</p> <h2 id="0330">0.33.0<a class="headerlink" href="#0330" title="Permanent link">¶</a></h2> <p><strong>New skill: <code>/anti-hall:update</code></strong> — in-session self-update with cache sync and changelog delta.</p> <p><code>scripts/update.js</code> (pure Node ≥ 18, cross-platform including Windows) implements the full update lifecycle: resolves the installed version from <code>installed_plugins.json</code> (v2 schema, harness-owned — read-only), <code>git pull --ff-only</code> the marketplace clone (fail-closed on dirty tree or non-fast-forward divergence — hard STOP with a clear message, no merge/rebase/force), mirrors the new version into the version-pinned cache so <code>/reload-plugins</code> can resolve it (semver-anchored, traversal-proof path join — no writes outside the clone dir + cache), extracts the CHANGELOG delta between installed and latest, and emits a JSON status + human summary. <code>--check</code> mode: <code>git fetch</code> + local-vs-remote version compare, no pull, no writes. Unknown failures are fail-closed (exit 1 with message) per a conservative posture; offline/no-git is reported and exits 0.</p> <p>Hardened by a 2-round deadly-swarm: a path-traversal P1 (unsanitized cache path join) and a live E2E registry-shape bug (v2 <code>installed_plugins.json</code> parsing) were caught and fixed before ship. 47 dedicated tests cover both modes, all STOP / fail-open / offline / already-up-to-date branches, traversal-proof cache sync, changelog extraction, and JSON status output.</p> <p>On a successful update the skill instructs the user to run <code>/reload-plugins</code> to load the new version in-session (hooks and statusline pick up changes from disk automatically; <code>/reload-plugins</code> refreshes the skill list and version label). Rarely, a harness build may require a restart instead — the skill says so when relevant and does not over-promise.</p> <p>Skills shipped: 7 → 8.</p> <h2 id="0321">0.32.1<a class="headerlink" href="#0321" title="Permanent link">¶</a></h2> <p>Docs: refresh test counts (459 pass / 461 total) across README (root+plugin), llms.txt, KB.md; index KB-fable-5.md and the v0.32.0 design plan in llms.txt + KB.md. No behavioral changes.</p> <h2 id="0320">0.32.0<a class="headerlink" href="#0320" title="Permanent link">¶</a></h2> <p><strong>Fable 5 awareness, model-routing guard, OMC-deference, TRIO debate roster, 3-phase deadly-swarm workflow, statusline segment-matching + latent-bug fix, <code>ANTIHALL_JUDGE_MODEL</code>.</strong></p> <h3 id="model-routing-guard-new-hook">model-routing-guard (new hook)<a class="headerlink" href="#model-routing-guard-new-hook" title="Permanent link">¶</a></h3> <p><code>model-routing-guard.js</code> — PreToolUse Agent/Task, always-on anti-waste net. Classifies spawn descriptions by keyword signals (mechanical vs complex) and nudges toward the cheapest model that fits the task shape:</p> <ul> <li><strong>Default (advisory):</strong> emits a <code>hookSpecificOutput.additionalContext</code> advisory when an explicit flagship model (<code>opus</code>/<code>fable</code>) is paired with a purely mechanical task (fetch, grep, build, deploy, run tests, etc.), or when <code>model</code> is omitted on a mechanical spawn (omitted model inherits the orchestrator's — on a flagship orchestrator that silently produces an all-flagship swarm).</li> <li><strong>Strict mode</strong> (<code>ANTIHALL_MODEL_ROUTING=strict</code>): upgrades omitted-model mechanical spawns to an unconditional block (exit 2). Opt-in and <strong>project-scoped</strong> — enable via the PROJECT's <code>.claude/settings.json</code> env block, NOT a global shell profile. A globally-exported strict blocks omitted-model mechanical spawns in EVERY project, including genuinely-cheap-orchestrator ones. Remedy: set an explicit cheap model on the spawn, or unset strict.</li> <li><strong>Debate-role exemption:</strong> row-1 blocks (explicit flagship + mechanical) are downgraded to advisory when a role-word (<code>reviewer</code>/<code>auditor</code>/<code>critic</code>/<code>debate</code>/<code>deadly-loop</code>) appears in the spawn <code>description</code> — TRIO debate seats legitimately need flagship models. The exemption does NOT apply to strict row-2 (role words must not defeat the user's explicit strict opt-in).</li> <li>Fail-open on any error; never blocks unknown model tokens (forward-compat).</li> <li>Hooks shipped: 21 → 23 (+<code>model-routing-guard.js</code>, +<code>omc-detect.js</code>).</li> </ul> <h3 id="omc-detect-new-shared-helper">omc-detect (new shared helper)<a class="headerlink" href="#omc-detect-new-shared-helper" title="Permanent link">¶</a></h3> <p><code>omc-detect.js</code> — shared pure-Node helper (not a hook) exported as <code>isOmcLoopActive({ cwd, sessionId })</code>. Returns <code>true</code> when an oh-my-claudecode autonomous loop (ralph, ultrawork, autopilot, ultraqa, team, ultrapilot, pipeline, omc-teams) is currently active AND fresh (any of <code>last_checked_at</code>/<code>updated_at</code>/<code>started_at</code> within 2 h, per OMC's own staleness rule) AND session-affinity-matched. Consumed by <code>task-guard</code> and <code>tasklist-guard</code> to suppress Stop-blocks to an advisory when an OMC loop is running — preventing the deadlock where the guard stops the loop it was meant to coexist with. Fail-open direction = NOT deferring (false): missing/malformed state = not active. Kill-switches: <code>DISABLE_OMC=1</code> or <code>OMC_SKIP_HOOKS</code> including <code>persistent-mode</code>. Version fragility documented in-file (state filenames stable since OMC 4.14.6).</p> <h3 id="trio-debate-roster-workstream-c">TRIO debate roster (Workstream C)<a class="headerlink" href="#trio-debate-roster-workstream-c" title="Permanent link">¶</a></h3> <p><code>MODEL-POLICY.md</code> (both copies) rewritten to a <strong>three-agent TRIO</strong>:</p> <table> <thead> <tr> <th>Role</th> <th>Model</th> <th>Thinking</th> <th>Persona</th> </tr> </thead> <tbody> <tr> <td><strong>Reviewer</strong></td> <td>latest flagship Claude (<code>model:"fable"</code>)</td> <td>adaptive, effort <code>xhigh</code> (→ <code>high</code>)</td> <td>correctness / architecture auditor</td> </tr> <tr> <td><strong>Auditor</strong></td> <td>latest Claude Opus (<code>model:"opus"</code>)</td> <td><code>xhigh</code></td> <td>divergent: regression & coupling hunter</td> </tr> <tr> <td><strong>Critic</strong></td> <td>latest OpenAI Codex</td> <td>max reasoning (<code>xhigh</code> → <code>high</code>)</td> <td>adversarial failure-mode hunter</td> </tr> </tbody> </table> <p>Floor for every seat = Opus. Availability fallback matrix covers all four availability combinations. Round governance: DEGRADED round (seat dead after retry) may iterate but cannot grant final GO. Dissent adjudication: single-seat re-run for evidence only (no code change); any fix wave → full TRIO next round.</p> <p><strong>Latest-model policy (owner directive):</strong> all spawn paths use harness tier tokens only (<code>fable</code>/<code>opus</code>/<code>sonnet</code>/<code>haiku</code>; Codex = latest the installed CLI reports) — resolved latest-at-call-time. NO versioned IDs in executable snippets. API call sites are the sole exception (no evergreen tier alias; <code>claude-haiku-4-5</code> is the alias-form for its tier).</p> <h3 id="deadly-loop-skill--workflow-workstream-e--e1">deadly-loop skill + workflow (Workstream E / E.1)<a class="headerlink" href="#deadly-loop-skill--workflow-workstream-e--e1" title="Permanent link">¶</a></h3> <p><code>skills/deadly-loop/SKILL.md</code> updated: - Phase B header + B0-B2 skeletons now describe the full TRIO (Reviewer + Auditor + Critic). - New <strong>Swarm mode</strong> section documents <code>references/deadly-loop.workflow.js</code> as the swarm-first path (plain Agent-tool path stays fully supported for no-consent sessions). - <strong>Three-phase architecture per round:</strong> (1) CONTEXT AGENT (<code>model:"sonnet"</code>, shared pack, graphify freshness check round-1 only); (2) THE DUEL (2a independent investigation + 2b structured argument); (3) CONVERGE & CONFIRM (VERDICT_SCHEMA dedup, RESPAWN-ON-DRIFT with objective criteria only). - Workflow consent friction, guard-coverage boundary, no cross-round caching — all stated honestly.</p> <p><code>references/deadly-loop.workflow.js</code> — new Workflow template (args contract): accepts <code>{round, multiplier, targetSHA, branch, scope, handoffPath, prevPackPath, findings, fixesApplied, contextMode, argue, respawnQuota, seats, codexAvailable}</code>. MODEL INVARIANT in file header: tier tokens only, never versioned IDs. DETERMINISM: no <code>Date.now()</code> / <code>Math.random()</code> / argless <code>new Date()</code>.</p> <h3 id="statusline-workstream-a2">statusline (Workstream A2)<a class="headerlink" href="#statusline-workstream-a2" title="Permanent link">¶</a></h3> <p><code>statusline-rich.js</code> <code>getModelName()</code>: token-segment matching (split model id on <code>-</code>, compare segments) for fable/opus/sonnet/haiku, fable-first. Fixes <code>confable</code>-class false-positive collisions. Latent bug fixed: the <code>:361-366</code> "pick max lastUsedAt" loop read a phantom field absent in real <code>~/.claude.json</code> (ts always 0, masked by last-key fall-through — P6 probe); simplified to explicit last-key with a comment stating the verified data shape (cumulative counters only, no timestamps).</p> <h3 id="antihall_judge_model">ANTIHALL_JUDGE_MODEL<a class="headerlink" href="#antihall_judge_model" title="Permanent link">¶</a></h3> <p><code>speculation-judge.js:194</code> now honors <code>ANTIHALL_JUDGE_MODEL</code> env override (house convention for all anti-hall LLM call sites). Default remains <code>claude-haiku-4-5</code> (alias-form, auto-tracking). <code>eval/run.js</code> / <code>eval/grade.js</code> were already env-overridable.</p> <h2 id="0311">0.31.1<a class="headerlink" href="#0311" title="Permanent link">¶</a></h2> <p><strong>Fix: statusline swarm-activity count is now per-session — no more cross-session/cross-project bleed.</strong></p> <p>The line-2 "orchestrating · N agents active" bar counted EVERY recent subagent spawn from EVERY Claude Code session on the machine, so a swarm running in one project showed its count on every other open project's statusline. Root cause: <code>phase-tracker.js</code> wrote bare <code>Date.now()</code> timestamps to the GLOBAL <code>~/.anti-hall/agent-spawns.log</code> with no session identity, and <code>phase-bar.js</code> <code>activityLine()</code> counted lines by TIME ONLY.</p> <p>Fix (per-session isolation): - <code>phase-tracker.js</code> now tags each spawn line as <code>"<ms> <tag>"</code>, where <code><tag></code> is the PreToolUse <code>session_id</code> (sanitized to <code>[A-Za-z0-9_-]</code>), falling back to <code>cwd-<sha1(cwd)[:12]></code>, else <code>unknown</code>. Retention prune (5 min) and fail-open / never-block behavior are unchanged; other sessions' fresh lines are preserved. - <code>phase-bar.js</code> <code>activityLine()</code> now derives THIS session's tag from the statusline's own session JSON (stdin: <code>session_id</code>, else <code>cwd</code>-hash) and counts ONLY entries whose tag matches AND fall within the 2-min activity window. LEGACY untagged lines belong to no session and are never counted (they age out). When no session identity is available, NO activity line is rendered (safer than a wrong count). Fail-open intact.</p> <h2 id="0310">0.31.0<a class="headerlink" href="#0310" title="Permanent link">¶</a></h2> <p><strong>Feature (OPT-IN, default OFF): <code>merge-gate</code> — a mechanical backstop for the v0.30.0 "false done" discipline.</strong></p> <p>Mechanizes the one <em>checkable</em> part of the false-done failure: the agent wrote a self-hedge ("first-pass" / "pending review" / "do not merge" / "needs your eyes") in its OWN recent output, then auto-merged anyway. <code>merge-gate</code> is a PreToolUse (Bash) hook that, <strong>only when <code>ANTIHALL_MERGE_GATE</code> ∈ {1,true,yes,on}</strong>, detects an auto-merge intent (<code>gh pr merge</code> incl. <code>--auto</code>, <code>gh pr review --approve</code>, <code>git merge --no-ff/--ff</code> into <code>main</code>/<code>master</code>/<code>develop</code>) and does a bounded (128 KB) tail-scan of the recent <strong>assistant</strong> transcript text for an UNRESOLVED hedge. A hedge is RESOLVED (and the merge allowed) when a resolution token follows it ("owner approved", "owner signed off", "fidelity verified", "verified against", "resolved:", "sign-off received"). Unresolved hedge + auto-merge → block (exit 2) with a verify-or-get-sign-off reason.</p> <p>HONEST limits (in the header): keyword-heuristic, bypassable (alternate merge syntax / heredoc / GitHub UI / API), <strong>default-OFF</strong>, fail-open on every error (no transcript, parse error, fs error, bad stdin). Cannot hard-loop — PreToolUse is single-shot and holds no state. A backstop on the v0.30.0 discipline, NOT a guarantee. Honors the <code>merge-gate</code> skip-hatch. Hooks shipped 19 → 20 (+<code>hooks.json</code> = 20 → 21 files); +18 tests.</p> <h2 id="0300">0.30.0<a class="headerlink" href="#0300" title="Permanent link">¶</a></h2> <p><strong>Fix (P0): "false done" — DONE now requires verification against the AGREED acceptance criteria, not tests-pass or a subagent "per-spec" report.</strong></p> <p>Fidelity that can't be mechanically verified (UI vs agreed mockup) is reported PENDING OWNER VERIFICATION, never folded into done as a hidden follow-up; coordinator verifies delegated acceptance (rule L); autonomy doesn't lower the bar. Grounded in a real field failure (UI shipped "done" off green behavior-tests + subagent self-reports, never compared to the agreed design HTML). Touches ship-it Step 4/6 + Autonomous-mode, always-on protocol rule 6, +1 per-turn nudge (17 → 18). A self-issued hedge (e.g., "first-pass / not pixel-perfect / pending review / needs your eyes") about a deliverable hard-blocks both its "done" status and any auto-merge; the coordinator's own written doubt is a verification signal.</p> <h2 id="0290">0.29.0<a class="headerlink" href="#0290" title="Permanent link">¶</a></h2> <p><strong>Feature (P0): <code>task-guard</code> now catches IDLE NEGLECT — the orchestrator sitting on dispatchable work instead of spinning parallel agents.</strong></p> <p>The #1 field pain: an autonomous run ends a turn with non-blocked, unassigned tasks pending and NO subagents running — it just stops instead of fanning out. The old Stop hook only knew "open tasks exist → generic nudge", which the model learned to ignore.</p> <p><code>task-guard</code> now <strong>classifies</strong> open tasks into <strong>ACTIONABLE NOW</strong> = status <code>pending</code> AND unowned (no <code>owner</code>, or owner is the main thread) AND no <strong>OPEN</strong> <code>blockedBy</code> (every blocker already in a done state). It also <strong>detects in-flight agents</strong> by scanning <code>~/.anti-hall/agents/*.json</code> for a FRESH heartbeat (numeric <code>ts</code>, or file mtime fallback, within ~20 min — same format <code>agent-watchdog.js</code> writes; absent dir = no agents).</p> <ul> <li><strong>Sharp block (the new condition):</strong> if ≥1 actionable-now task AND no agents running → <strong>IDLE NEGLECT</strong> → block naming those tasks: <em>"IDLE NEGLECT: N non-blocked, unassigned task(s) and NO agents running — dispatch them in PARALLEL NOW (one background agent each, cap ~min(16, cores-2)): <names>. Do not end the turn idle; only stop if a task truly needs the user (then say which + why)."</em></li> <li><strong>Gentle path (no nagging real work):</strong> if agents ARE in flight, or the only open tasks are blocked/owned/in_progress, it falls back to the existing generic drain nudge.</li> </ul> <p><strong>Loop-safety (cannot hard-loop):</strong> the idle-neglect block dedupes on a hash of (actionable-set + <code>"no-agents"</code>) so it re-fires only when that set actually changes; an absolute <code>MAX_BLOCKS</code> cap — raised modestly <strong>3 → 5</strong>, counting BOTH modes — guarantees a genuinely-stuck set goes quiet after a few nudges. All existing safety kept (top-level try/catch → exit 0, skip-hatch, per-session state under <code>~/.anti-hall/</code>, fail-open on any error, bounded transcript tail-read).</p> <p>Also: <strong>orchestration rule C</strong> (<code>verify-first-full.js</code> SessionStart protocol) reworded to demand <strong>PROACTIVE</strong> parallel dispatch — fire a background agent for a pending, unblocked, unassigned task the moment it exists, <em>without being asked</em>; ending a turn with such tasks and no agents running is explicitly named IDLE NEGLECT.</p> <p><strong>Complementary per-turn layer — <code>task-tracker</code> (UserPromptSubmit) now nudges BEFORE the turn, not only at Stop.</strong> The Stop-hook idle-neglect block fires after the model already decided to stop; <code>task-tracker</code> now reviews the actionable-now set on <em>every</em> prompt and, when ≥1 pending+unowned+unblocked task exists AND no agents are in flight, injects a SPECIFIC review line into <code>additionalContext</code>: <em>"TASK REVIEW (every turn): N non-blocked, unassigned pending task(s) — dispatch a background agent for EACH now, in parallel (cap ~min(16, cores-2)), unless already in-flight: <up to 4 names>. Do not leave them idle; only hold one if it truly needs the user."</em> It reuses task-guard's exact definitions — <code>normOwner</code> / <code>normBlockedBy</code> / <code>classifyOpen</code> (owner main/orchestrator/coordinator = ours; blocker open unless its task is done/completed/cancelled) and the <code>~/.anti-hall/agents/*.json</code> fresh-heartbeat check. Task subjects are control-char-stripped then <code>JSON.stringify</code>'d (inert quoted strings — no prompt injection). When 0 actionable, the existing generic discipline + open-tasks freshness note are unchanged. Reconstruction now also captures <code>owner</code>/<code>blockedBy</code> per task (status-only <code>TaskUpdate</code> does not clear them). Bounded (same 256 KB tail read) and fully fail-open — any error → existing generic text or nothing, never wedges the turn.</p> <p>Tests: +6 task-guard cases (actionable-now + no-agents → idle-neglect naming tasks; agents-running → generic not idle-neglect; all-blocked → generic; owned-pending → not actionable; idle-neglect dedupe; churn-cap loop-safety). +5 task-tracker cases (actionable-now → review line names tasks + says parallel; 0 actionable → generic only; owned/blocked → not listed; fresh agent → no review line; malformed transcript → fail-open). +E2E system test for the task-neglect enforcement (<code>tests/hooks/task-neglect-e2e.test.js</code>): one realistic 3-pending-task transcript driving BOTH real hooks together against shared, real fs heartbeat state — no-agents → tracker review line + guard idle-neglect block naming only the actionable task; fresh heartbeat planted → both back off; actionable in_progress → neither flags; loop-safety end-to-end (dedupe once, churn capped at MAX_BLOCKS). Suite <strong>340 passing / 342 total</strong>.</p> <h2 id="0282">0.28.2<a class="headerlink" href="#0282" title="Permanent link">¶</a></h2> <p><strong>Fix (P0): <code>ship-it</code> plan-approval gate now honors granted autonomy — no more blocking for a "go" already given.</strong></p> <p>Reported in the field: an autonomous run sat idle for hours at <code>ship-it</code>'s <code>ExitPlanMode</code> plan-approval gate, waiting on approval the owner had already granted ("full autonomy / build it / AFK"). Root cause: when the lean <code>ship-it</code> replaced feature-launch, it dropped feature-launch's autonomous-mode handling, so the plan-approval and brainstorm gates became <strong>unconditional human stops</strong> with no autonomy carve-out.</p> <p>Fix — the two human gates (Step 1 brainstorm, Step 3 <code>ExitPlanMode</code>) are now explicitly <strong>SOFT</strong>: under granted autonomy they are satisfied by <em>forming/recording the design</em> and <em>converging the plan via the deadly-loop to zero NEW P0s</em>, then the run <strong>proceeds straight into the build</strong> — it does not re-stop for an already-given go. A new <strong>Autonomous mode</strong> section makes this the rule: stop only at a <strong>hard safety boundary</strong> (destructive / irreversible / financial / secret / prod-deploy / force-push — these still NEVER autonomy-bypass) or a <strong>genuine ambiguity</strong>, never the plan-approval gate itself. Interactive behavior is unchanged (present plan, wait for approval).</p> <h2 id="0281">0.28.1<a class="headerlink" href="#0281" title="Permanent link">¶</a></h2> <p><strong>Loose-end closeout — cross-model (Codex) review fixes for <code>ship-it</code> + a full doc-currency pass. No new features.</strong></p> <p>A cross-model Codex review of the shipped <code>ship-it</code> skill + its workflow template caught five real correctness/honesty gaps, all fixed:</p> <ul> <li><strong>Plan-mode write claim corrected.</strong> Plan mode is read-only <em>for the repo</em> (it writes only to <code>~/.claude/plans/</code>), so <code>ship-it</code> no longer claims it can write the repo <code>PLAN.md</code> inside plan mode — the plan is drafted + presented via <code>ExitPlanMode</code>, and <code>PLAN.md</code> is written as the first action <em>after</em> approval.</li> <li><strong>Two-phase tier-sizing closed a gap.</strong> Even an S-classified change now does a ~30-second blast-radius sanity glance before the tier is locked, so a deceptively-large "simple" ask can't skip the re-tier (S stays lean otherwise).</li> <li><strong>Workflow template hardened.</strong> Takes <code>files[]</code> (not a diff string) and <code>validateGroup()</code> fails closed unless a fan-out group is conflict-free (disjoint files, unique labels, no intra-group dependency) before any <code>parallel()</code>.</li> <li><strong>Template honesty.</strong> The template is now labelled a SINGLE-PASS audit <em>scaffold</em> (shows the fan-out shape); the full iterate-to-zero-NEW-P0 fix-wave + D1.5 gate is run by invoking the <code>deadly-loop</code> skill — no false convergence claim.</li> <li><strong>Commit ownership clarified.</strong> The template never commits; agents return results and the coordinator commits serially on the main thread.</li> </ul> <p>Plus a thorough doc-currency pass: test counts reconciled to <strong>329 pass / 331 total / 2 platform-skip</strong>, and <code>ship-it-guard</code> added to every user-facing hook inventory (it was only in the KB).</p> <h2 id="0280">0.28.0<a class="headerlink" href="#0280" title="Permanent link">¶</a></h2> <p><strong><code>ship-it</code> v2 — 2-phase tier-sizing, an OPT-IN enforcement gate, and a copyable <code>/ship-it</code> workflow template.</strong></p> <p>This release hardens the <code>ship-it</code> workflow with a thin mechanical backstop and a reusable execution script, without changing its default behavior (the new gate is OFF by default).</p> <ul> <li><strong>2-phase tier-sizing (skill).</strong> Step 0 now decides the S/M/L tier <strong>twice</strong>: a PROVISIONAL tier from the initial prompt (sets how much exploration to do), then a CONFIRMED/REVISED tier at the blast-radius map after the Step-2 graphify-first research reveals the true blast radius. A deceptively-large "simple" ask (one-liner that touches auth, or fans out to many callers) gets upgraded; an over-estimate gets downgraded — re-decided <strong>before</strong> plan mode locks rigor.</li> <li><strong><code>ship-it-guard</code> hook (new; OPT-IN, default OFF).</strong> PreToolUse on <code>Write|Edit|MultiEdit</code>. A pure no-op (exit 0) unless <code>ANTIHALL_SHIPIT_GATE</code> ∈ {1,true,yes,on}. When ON: blocks (exit 2) a CODE edit on a <strong>hard-risk path</strong> (migration / auth / <code>.github/workflows</code> / security/crypto) when <strong>no <code>PLAN.md</code></strong> exists (repo root or <code>.planning/PLAN.md</code>) — nudging the agent to plan first. <strong>Honest limits (in the hook header):</strong> enforces artifact-EXISTENCE only, NOT plan quality (a stub <code># Plan</code> satisfies it); bypassable via a <code>Bash</code> heredoc write (PreToolUse sees Edit/Write/MultiEdit only); <strong>conservative</strong> — never gates ordinary single edits, docs, or tests; <strong>fail-open</strong> on any error; honors the shared <code>isSkipped('ship-it-guard')</code> escape-hatch. Registered in <code>hooks.json</code> (timeout 10).</li> <li><strong>Copyable <code>/ship-it</code> workflow template (new).</strong> Investigated feasibility: per the official <a href="https://code.claude.com/docs/en/workflows">Dynamic Workflows docs</a>, a plugin <strong>cannot ship</strong> a workflow command — there is no <code>workflows</code> field in <code>plugin.json</code>, and a workflow only becomes a <code>/command</code> by saving a <em>live run's</em> script via <code>/workflows</code> → <code>s</code> into <code>.claude/workflows/</code> (project) or <code>~/.claude/workflows/</code> (user). So <code>ship-it</code> ships a <strong>copyable template</strong> at <code>skills/ship-it/references/ship-it.workflow.js</code> (deterministic — no <code>Date.now</code>/<code>Math.random</code>/<code>new Date</code>; inputs via <code>args</code>) that automates the L-tier Step-4 build fan-out (<code>parallel([...])</code> over disjoint phases) + the Step-5 per-phase deadly-loop (Reviewer + Codex Critic via <code>agentType: "codex:codex-rescue"</code>), plus a one-line pointer in the skill telling the user to save it as <code>/ship-it</code>.</li> <li><strong>Tests:</strong> <code>tests/hooks/ship-it-guard.test.js</code> (13 cases — default-off no-op, ON+L-risk+no-PLAN ⇒ exit 2 + reason, PLAN.md present ⇒ allow, ordinary/doc/test file ⇒ allow, env-value parsing, fail-open on malformed/empty stdin, skip-hatch). Full suite: 331 tests, 0 fail.</li> </ul> <h2 id="0270">0.27.0<a class="headerlink" href="#0270" title="Permanent link">¶</a></h2> <p><strong><code>ship-it</code> workflow replaces <code>feature-launch</code>. deadly-loop gains a D1.5 verification gate.</strong></p> <p>This release retires the <code>feature-launch</code> skill and ships <code>ship-it</code> in its place — one lean, anti-hall-native workflow for shipping any change correctly, from a one-line fix to a multi-phase feature.</p> <ul> <li><strong><code>ship-it</code> (new, replaces <code>feature-launch</code>).</strong> A tier-scaled (S / M / L) fusion of the best of superpowers-planning + the GSD phase loop + the deadly-loop, with the bloat removed. Brainstorm + plan happen <strong>in plan mode</strong> (<code>ExitPlanMode</code> is the build-unlock approval gate; no code before approval); the plan is hardened with the deadly-loop <strong>before</strong> any code; on large (L) work the disjoint build phases and the per-phase deadly-loop fan out as a <strong>Workflow swarm</strong>; each phase is verified with fresh evidence and a vacuous-test guard, then hardened until zero NEW P0s. Wired to real Claude Code primitives (plan mode + the Workflow tool) plus anti-hall's own deadly-loop and always-on guards — <strong>standalone</strong>, no GSD/superpowers dependency. Hard safety boundaries (force-push, prod deploy, destructive/financial actions) never autonomy-bypass and are enforced by the always-on guards, which swarm agents inherit.</li> <li><strong><code>feature-launch</code> removed.</strong> Its bespoke <code>references/</code> (including the per-project <code>PRE-TOOL-USE-HOOK.md</code> template) are gone; <code>ship-it</code> relies on the always-on guards instead of a bespoke per-feature sentinel, and reuses the deadly-loop's <code>MODEL-POLICY.md</code>, A3 branch/SHA verification preamble, validation table, and D1.5 gate. The shared <code>MODEL-POLICY.md</code> is now duplicated in 2 places (canonical + <code>deadly-loop/references/</code>) instead of 3.</li> <li><strong>deadly-loop D1.5 verification gate.</strong> A GO verdict is no longer valid without a D1.5 check — fresh evidence (re-run the authoritative check this round, not a stale prior result) plus a vacuous-test guard (a passing test that asserts nothing, or never exercises the changed path, does not count). Inherited by <strong>all</strong> deadly-loop-driven workflows, including <code>ship-it</code>'s per-phase gates and <code>deadly-loop-multi</code>.</li> </ul> <h2 id="0260">0.26.0<a class="headerlink" href="#0260" title="Permanent link">¶</a></h2> <p><strong>Structured-return discipline (rule G) made concrete + measured.</strong></p> <p>Rule G (SYNTHESIZE, NEVER RELAY) already required subagents to return tight summaries under an OUTPUT BUDGET. This release makes the SUBSTANTIAL-return case concrete and grounds it in measurement, with no mechanical or behavioral change to any guard.</p> <ul> <li><strong>Schema now specified.</strong> For a SUBSTANTIAL return (a review/audit/research dump, many claims), rule G now names the exact compact shape to require: <code>{claim, evidence:"file:line", verdict, blockers/uncertainty, next}</code>.</li> <li><strong>Measured, not asserted.</strong> A deadly-loop-hardened study (S4) measured these structured subagent returns at <strong>~5× smaller than verbose prose with zero decision-relevant loss</strong>, judged on a claim/evidence/uncertainty/blockers/next rubric (N=8, directional). Rule G cites the figure inline.</li> <li><strong>Reconciles the earlier ~1.4× number.</strong> The prior pilot's ~1.43× density figure measured <em>small, already-summarized</em> content (where a schema is a minor lever, and JSON overhead can make tiny outputs LARGER). The ~5× figure is for <em>verbose</em> returns. Both hold; they measure different inputs. The <strong>prose-for-tiny caveat is kept</strong> — a single prose line still wins for a SMALL result; do not impose JSON there.</li> <li><strong>Enforce, don't just request.</strong> Rule G notes that passing a <strong>schema to the Agent/Task tool</strong> validates the structured return rather than merely asking for it. The biggest levers remain the output budget + no-raw-relay rule; the schema is the multiplier on large returns.</li> <li>The per-turn SYNTHESIZE nudge (#57) was updated in place to carry the schema + the ~5× figure. Nudge count unchanged (<strong>17</strong>).</li> </ul> <p>No guard mechanics or hook behavior changed — this is prompt discipline (rule G text + one nudge) only. <code>node --test</code> green: 316 passing (+2 platform-skipped), 318 total.</p> <h2 id="0252">0.25.2<a class="headerlink" href="#0252" title="Permanent link">¶</a></h2> <p><strong>CI green on all platforms — root-cause fixes for the macOS stdout race and Windows base-command env.</strong></p> <p>0.25.1's test-gating fixes were incomplete; CI stayed red on macOS node 18/20 and all Windows legs. Root-caused properly (from real CI logs) and fixed at the source:</p> <ul> <li><strong>macOS (hook source fix):</strong> <code>verify-first-full.js</code> and <code>graphify-session.js</code> emit a ~10 KB JSON payload, then <code>process.exit(0)</code>. On a pipe, <code>process.stdout.write</code> is async once the payload exceeds the OS pipe buffer, so <code>exit(0)</code> could race the flush and the reader saw empty/partial stdout (intermittent on macOS node 18/20). Switched both to a blocking <code>fs.writeSync(1, …)</code> so every byte is handed to the pipe before exit. The test-side <code>expectJson</code> retry is kept as defense-in-depth and corrected to re-spawn on any parse failure (empty OR partial), not just empty.</li> <li><strong>Windows (test harness fix):</strong> the statusline base-command test starved <code>cmd.exe</code> with a stripped env (a hand-picked <code>SystemRoot/ComSpec/PATHEXT</code> allowlist was insufficient), so the base command failed and the dispatcher fell back. The statusline test harness now inherits the full parent env and overrides only the HOME-pointing vars (statusline scripts read no Claude Code markers, so only HOME needs isolating).</li> </ul> <p>No user-facing behavior change beyond more reliable hook stdout. <code>node --test</code> green on Ubuntu, macOS, and Windows × Node 18/20/22/24.</p> <h2 id="0251">0.25.1<a class="headerlink" href="#0251" title="Permanent link">¶</a></h2> <p><strong>CI fix — cross-platform test gating. No plugin behavior change from 0.25.0.</strong></p> <p>The 0.25.0 test suite passed locally but failed on CI macOS + Windows: all failures were test-side environment assumptions, not plugin bugs (source behavior is correct on every platform). Fixed test-side only:</p> <ul> <li><strong>macOS:</strong> a graphify reflected-path test asserted a substring (<code>graphify-out</code>) that the hook's 80-char <code>sanitizePath</code> cap truncates under CI's long <code>/var/folders/.../T/</code> tmpdir (passed locally only because the dev tmpdir is short) — now asserts the always-in-cap base-dir name. Added an opt-in single retry for a macOS <code>spawnSync</code> empty-stdout pipe flake on the large-JSON SessionStart hook (node 18/20).</li> <li><strong>Windows:</strong> install-reaper <code>--dry-run</code> tests now assert the <code>win32</code> no-op message (the installer correctly short-circuits before reading flags); tests that <code>mkdir</code> a directory whose name contains characters illegal in Windows filenames (<code>"</code>, control, bidi) are <code>win32</code>-skipped; statusline base-command tests invoke <code>node "<abspath>"</code> (survives both <code>sh -c</code> and <code>cmd /c</code>) instead of a nested-quote <code>node -e</code>.</li> </ul> <p>Net: <code>node --test</code> green on Ubuntu, macOS, and Windows × Node 18/20/22/24.</p> <h2 id="0250">0.25.0<a class="headerlink" href="#0250" title="Permanent link">¶</a></h2> <p><strong>Always-on SCOPE & FIDELITY discipline + an opt-in mcp-reaper companion (macOS + Linux).</strong></p> <p><strong>A — SCOPE & FIDELITY discipline (prompt layer).</strong> A new always-on discipline injected by the SessionStart protocol (<code>verify-first-full.js</code>) and reinforced by 2 new per-turn nudges (<code>verify-first.js</code>, NUDGES 12 → 14; later 14 → 15 with the verify-delegated-work nudge in section C). It enforces: solve the ACTUAL problem with the <strong>simplest sufficient solution</strong> (over-engineering is confabulating work the user never asked for); <strong>intent over letter</strong> (serve what the user means; do the small reading and say what you skipped rather than guess-big); <strong>confirm before expanding scope</strong> (new platform / file / dependency / phase / abstraction); <strong>match rigor to blast radius</strong> (heavy process is for risky or large work, not a reflex on small asks); and <strong>finish what was asked / drop nothing silently</strong>. It is now named in the "ALWAYS APPLY" disciplines list alongside root-cause / orchestration / anti-sycophancy, and mirrored in <code>AGENTS.md</code> for Codex.</p> <p><strong>B — mcp-reaper companion (OPT-IN, macOS + Linux).</strong> A new pure-Node background <strong>companion</strong> (NOT a hook — an interval job via a macOS LaunchAgent / Linux <code>systemd --user</code> timer, cron fallback) that kills <strong>orphaned</strong> MCP-server processes — ones leaked when their spawner (a Claude / codex / npm / node session) exits without cleaning them up (on macOS these reparent to launchd and pile up over a workday). Files: <code>companion/mcp-reaper.js</code>, <code>companion/install-reaper.js</code>, <code>companion/README.md</code>. Install with <code>node plugins/anti-hall/companion/install-reaper.js</code> (<code>--uninstall</code> to remove). Env knobs: <code>MCP_REAP_DRYRUN=1</code>, <code>MCP_REAP_GRACE</code>, <code>ANTIHALL_REAPER_MATCH</code>, <code>ANTIHALL_REAPER_EXCLUDE</code> (regex of cmd substrings to NEVER reap). Also recognizes <strong>Python MCPs</strong> (<code>uvx</code>/<code>uv</code> + underscore <code>mcp_server_*</code> forms), not just Node.</p> <ul> <li><strong>Safety invariant:</strong> a process is reaped only if its command matches a generic MCP signature <strong>and</strong> its parent is a reaper/init (pid1 / launchd / <code>systemd --user</code> / WSL <code>Relay()</code>). Because Unix always reparents a dead process's children, a <em>live</em> MCP's parent is always a live spawner — never a reaper — so "parent is a reaper" means the spawner died, i.e. the MCP is a true orphan. Killing an in-use server is impossible by construction.</li> <li><strong>Limitation — service-managed MCPs:</strong> an MCP run as a macOS LaunchAgent / <code>systemd --user</code> unit / other OS service shares init (ppid 1) as a parent <strong>while alive</strong> — indistinguishable from a leaked orphan — so it could be reaped. Exclude it via <code>ANTIHALL_REAPER_EXCLUDE='your-service-name|another'</code> (case-insensitive regex of cmd substrings that are never reaped).</li> <li><strong>Windows is a documented no-op (rescope rationale):</strong> Windows has no parent-death reparenting <strong>and</strong> recycles PIDs, so external orphan detection is unsafe there — a matched signature on a recycled PID with an <code>init</code>-ish parent could kill an unrelated live process. The correct fix on Windows is <strong>Job Objects set by the spawner</strong>, which a companion cannot do. So the installer prints why and installs no scheduler.</li> </ul> <p><strong>C — VERIFY DELEGATED WORK discipline (prompt layer).</strong> Orchestration now requires the coordinator to <strong>independently verify delegated work</strong> — a subagent's "done / fixed / tests pass / N passing" is an UNVERIFIED CLAIM, never a fact. Before marking any delegated task complete, the coordinator RE-RUNS the authoritative check (or dispatches a separate verifier) and reconciles multiple workers against GROUND TRUTH, not against each other. Added as rule L in <code>verify-first-full.js</code>, folded into the "ALWAYS APPLY" orchestration summary, mirrored in <code>AGENTS.md</code>, and reinforced by 1 new per-turn nudge (<code>verify-first.js</code>, NUDGES 14 → <strong>15</strong>).</p> <p><strong>D — Background-default orchestration (prompt layer).</strong> Orchestration now defaults delegated heavy/parallel/long work to the <strong>background</strong> (the coordinator passes <code>run_in_background</code> itself) so the user needn't background it manually, while still verifying each on completion — never fire-and-forget. Extended into rule F in <code>verify-first-full.js</code>, mirrored in <code>AGENTS.md</code>, and reinforced by 1 new per-turn nudge (<code>verify-first.js</code>, NUDGES 15 → <strong>16</strong>).</p> <p><strong>E — Fix-history ledger discipline (prompt layer).</strong> The <code>tasklist-guard</code> Stop reminder now also prompts appending each <strong>completed task</strong> to <code>.anti-hall-history.md</code> — an append-only fix ledger (one entry per task: <strong>Cause / Fix / Verified</strong>) so the fix history persists for the knowledge layer. Same reminder, enriched (no new hard Stop-block condition), fully fail-open. Mirrored as full-protocol text in rule B of <code>verify-first-full.js</code> and documented in <code>docs/TASKLIST-GUARD.md</code> / <code>docs/KB.md</code>. The hook never creates the file — it is gitignored, local session state.</p> <p><strong>F — Message-context bloat prevention (#45, prompt layer).</strong> Orchestration rule G is rewritten from "synthesize, don't paste raw output" into <strong>SYNTHESIZE, NEVER RELAY</strong>: the coordinator reports findings in its own words and NEVER pastes a subagent's raw return into the user thread (that verbatim relay is the <strong>#1 cause of message-context bloat</strong>), and subagents must return TIGHT summaries under an explicit <strong>OUTPUT BUDGET</strong> (findings only; a compact <code>{claim, evidence:"file:line", verdict}</code> schema only when >~5 claims or >~200 tokens, else one prose line). Reinforced by 1 new per-turn nudge (<code>verify-first.js</code>, NUDGES 16 → <strong>17</strong>). <strong>Pilot finding:</strong> a compact JSON return schema is only <strong>~1.43× denser than prose on average</strong> (and <em>worse</em> for tiny outputs), so schema enforcement is a MINOR lever — the real levers are the output budget + the no-raw-relay rule, which is why this ships as prompt discipline, not a schema-enforcement system. A <code>PostToolUse</code>-on-<code>Task</code> hook to flag oversized subagent returns was evaluated and is <strong>NOT feasible</strong>: per the Claude Code hook contract (KB-claude-codex §1.4), <code>PostToolUse</code> stdout /<code>additionalContext</code> never reaches the model and <code>PostToolUse</code> cannot block — so it cannot inject a reminder back. No hook was faked; the discipline is the mechanism.</p> <p>Suite 318 total, <strong>316 passing</strong> (+2 platform-skipped).</p> <h2 id="0241">0.24.1<a class="headerlink" href="#0241" title="Permanent link">¶</a></h2> <p><strong>Doc-currency pass — descriptions + test counts brought to current reality.</strong></p> <p>No behavioral change. The <code>plugin.json</code> and <code>marketplace.json</code> descriptions predated api-guard and the v0.24.0 gh-PR guard — both now name the two signature mechanical guards (api-guard: fabricated-API blocking, opt-in 3rd-party; git-guard: AI self-credit in commits <strong>and</strong> gh pr/issue/release bodies). Test-count references corrected to <strong>173</strong> across README (root + plugin), <code>llms.txt</code>, and <code>docs/KB.md</code>; KB snapshot provenance refreshed to post-0.24.0.</p> <h2 id="0240">0.24.0<a class="headerlink" href="#0240" title="Permanent link">¶</a></h2> <p><strong>git-guard now also blocks AI self-credit in <code>gh</code> PR / issue / release bodies.</strong></p> <p>Previously git-guard blocked AI co-author trailers in <code>git commit</code> only; PRs created with <code>gh pr create --body "… 🤖 Generated with Claude Code …"</code> slipped through. Now the same self-credit markers (the <code>🤖 Generated with [Claude Code](…)</code> footer, <code>Co-Authored-By: Claude</code>, <code>noreply@anthropic.com</code>, a bare <code>claude.com/claude-code</code> link) are blocked in the inline <code>--body</code>/<code>--title</code>/<code>--notes</code> of <code>gh pr|issue|release create|edit|comment</code>. Reuses the existing commit markers + a bare-link marker; quote/segment-aware via git-guard's tokenizer. Inline values only — <code>--body-file</code>/<code>-F</code> and heredoc/command-substitution bodies put the text off the command line and are a documented fail-open limitation. +8 gh tests; suite 173/173.</p> <h2 id="0230">0.23.0<a class="headerlink" href="#0230" title="Permanent link">¶</a></h2> <p><strong>api-guard v2 — opt-in 3rd-party API verification + security hardening.</strong></p> <p>api-guard can now verify installed <strong>3rd-party</strong> packages (pandas, lodash, …) — the highest-hallucination class (38–80% per the literature) — not just stdlib/builtins. But a double deadly-loop (2 Opus + 2 Codex) <strong>proved that verifying a 3rd-party package = importing it = running its top-level code at edit time</strong> (a confirmed RCE, triggered even when the write is blocked; Codex bypassed the first fix via bare <code>node_modules</code> packages and <code>.pth</code> files). You cannot check an installed package's API without executing it. So 3rd-party checking is <strong>opt-in, off by default</strong>:</p> <ul> <li><strong>Default (unchanged):</strong> stdlib/builtin modules + JS globals only — importing those to introspect is side-effect-free, so the probe never runs untrusted code. Safe.</li> <li><strong>Opt-in:</strong> set <code>ANTIHALL_API_GUARD_THIRDPARTY=1</code> to also verify installed packages, accepting that referenced installed packages are imported at edit time.</li> </ul> <p><strong>Security hardening (both modes):</strong> - Local/relative modules are NEVER probed: path-spec rejection (<code>./x</code>, <code>/abs</code>, <code>..</code>, <code>C:\</code>), Python probe runs with <code>cwd=<tmp></code> + scrubs cwd/<code>''</code> from <code>sys.path</code>, so a bare <code>import localmod</code> can't resolve to a repo file. - <code>SAFE_ENV</code> now also strips <code>PYTHONUSERBASE</code>/<code>PYTHONSAFEPATH</code>, and the Python probe runs with <code>-s</code> (no user site) — closes the <code>.pth</code> startup-exec vector. - Batched <strong>one import per module</strong> (not per attribute); <strong>30s wall-clock deadline</strong> under the hook timeout (raised 30→45s); JS requires extracted from comment-stripped code; function/lambda-parameter and <code>with/except as</code> names excluded from checking (no false-block on <code>def f(pd): pd.x</code>).</p> <p>165 tests (3rd-party gated + RCE-no-execute regression), bench 21/21 catch / 0 FP. Known v2.1 gaps (fail-open, documented): dotted-submodule receivers (<code>os.path.x</code>) and JS member chains (<code>fs.promises.x</code>) are not yet checked.</p> <h2 id="0222">0.22.2<a class="headerlink" href="#0222" title="Permanent link">¶</a></h2> <p><strong>Fix: api-guard CI flake on Windows/node (probe timeout too tight).</strong></p> <p>After 0.22.1 the windows-latest leg still flaked intermittently (same commit green on <code>main</code>, red on the tag run): a COLD <code>python</code>/<code>node</code> spawn on a loaded Windows runner occasionally exceeded the <code>1500ms</code> probe timeout → the check timed out → fail-open → the <code>asyncio.run_all</code> test expecting a block got an allow. Raised <code>SPAWN_TIMEOUT_MS</code> to <code>5000ms</code> (a CEILING, not added latency — a normal spawn returns in <300ms; this just stops us giving up too early on a cold start), lowered <code>MAX_CHECKS</code> 12→6 to keep the worst case bounded, and bumped the hook's outer timeout 30→45s for headroom. Suite 159/159.</p> <h2 id="0221">0.22.1<a class="headerlink" href="#0221" title="Permanent link">¶</a></h2> <p><strong>Fix: api-guard now works on Windows (CI was red on the windows-latest matrix legs).</strong></p> <p>The probe env was a PATH-only allowlist, which is secure but too minimal for Windows — Python could not spawn without <code>APPDATA</code>/<code>LOCALAPPDATA</code>/<code>TEMP</code>/etc., so <code>pyBin()</code> returned null and the hook fail-opened (Python fabrications silently allowed) on Windows. Changed <code>SAFE_ENV</code> to a <strong>denylist</strong>: the full parent environment minus the interpreter-injection vectors (<code>NODE_OPTIONS</code>, <code>NODE_PATH</code>, <code>PYTHONSTARTUP</code>, <code>PYTHONPATH</code>, <code>PYTHONHOME</code>, <code>PYTHONINSPECT</code>, <code>PYTHONEXECUTABLE</code>). Keeps the security property (a poisoned env can't influence the existence check) while letting Python spawn on every OS. Suite 159/159; the exact failing case (<code>asyncio.run_all</code>) now blocks. Doc-only follow-up <code>547e184</code> (corrected stale test counts to 159) also rolled up here.</p> <h2 id="0220">0.22.0<a class="headerlink" href="#0220" title="Permanent link">¶</a></h2> <p><strong><code>api-guard</code> — a mechanical guard against API hallucination, built on eval evidence.</strong></p> <p>A controlled A/B eval (see <a href="https://github.com/talas9/anti-hall/tree/main/tools/eval"><code>tools/eval/</code></a>) established that the verify-first <em>prompt</em> does not reliably reduce API fabrication: across four rounds (incl. a powered 122-trap, tools-on run with a naive baseline) the protocol netted <strong>no statistically-significant reduction</strong> (McNemar p=0.26), and the model ran a verification tool only ~5% of the time — <em>the same as baseline</em>. The model ignores "go verify." So this release does the verifying mechanically:</p> <ul> <li><strong><code>api-guard</code></strong> (new) — a <code>PreToolUse</code> hook on <code>Write</code>/<code>Edit</code>/<code>MultiEdit</code>. It extracts <code>module.attribute</code> references from the code about to be written and resolves them against the <strong>installed</strong> runtime (<code>python3</code> / <code>node</code>): Python stdlib <code>mod.attr</code> and <code>from mod import Name; Name.attr</code> (via <code>hasattr</code>), Node builtins <code>require('mod').attr</code>, and JS global builtins incl. <code>.prototype.attr</code> (via <code>typeof</code>). If a real module/object is missing the referenced attribute, the write is <strong>blocked</strong> — the symbol was fabricated.</li> <li><strong>Substantiated: 100% in-scope catch, 0 false positives</strong> on a committed, reproducible bench — <code>node eval/api-guard-bench.js</code> (labeled corpus + a sweep of every <code>plugins/**.js</code>). Contrast the prompt's unproven ~18%.</li> <li><strong>Fail-open by construction:</strong> blocks ONLY on a positively-verified missing attribute; any uncertainty (no interpreter, import error, 3rd-party package, receiver-typed instance method, version skew) → allow. From-import bindings resolve before module names so <code>datetime.fromisoformat</code> (class method) is not mistaken for a module attribute.</li> <li><strong>Hardened by a double deadly-loop (2 Opus + 2 Codex) + empirical sweep.</strong> Security review found injection/ReDoS/resource genuinely closed. Correctness review found and fixed a P0 false-positive cluster: stdlib/global names used as <strong>local variables</strong> (<code>array = [1,2]; array.append(3)</code>, <code>const Math = myLib; Math.x</code>), require-vars <strong>reassigned</strong> to a 3rd-party (<code>fs = require('fs-extra')</code>), and from-import names rebound — now all resolve correctly (require a real <code>import</code>; exclude locally-bound names). Added <code>Buffer</code> to JS globals; <code>python</code> (3.x) fallback for non-<code>python3</code> systems; probe env pinned to PATH-only.</li> <li>Skip-hatch via <code>~/.anti-hall/skip.json</code> (<code>api-guard</code>); 27 tests (block/allow/fail-open/ scope/skip/regression/shadowing/portability), <code>python3</code>-gated for Windows CI. Full suite <strong>159/159</strong>.</li> <li><strong>eval/</strong> harness added: <code>claude -p</code> subscription backend (no API key), naive-baseline + tools-on knobs, 122 execution-verified traps, <code>analyze.js</code> (McNemar exact + bootstrap CI). Documents the honest finding that prompt-only fabrication reduction is unproven.</li> </ul> <h2 id="0211">0.21.1<a class="headerlink" href="#0211" title="Permanent link">¶</a></h2> <p>Refresh marketplace.json plugin description to current capability set (tasklist-guard, skip-guard, deadly-loop-multi, speculation guards, rule K, escape hatch) for the public listing; add assets/demo/ (VHS .tape + storyboard) to generate a demo GIF.</p> <h2 id="0210">0.21.0<a class="headerlink" href="#0210" title="Permanent link">¶</a></h2> <p>Pre-publish <strong>triple-deadly-loop</strong> hardening pass (3 Opus reviewers + 3 Codex critics, two re-converge passes) before the first public release. Closes a cluster of deliberate-evasion gaps in the always-on guards and the reflected-text sanitizers:</p> <ul> <li><strong>git-guard</strong> — now blocks AI/assistant self-credit trailers slipped in via <code>git commit --trailer</code>, including the <code>key=value</code> separator form (<code>Co-Authored-By=Claude <…></code>) alongside the <code>key: value</code> form, and a <code>-c trailer.<name>.key=<self-credit></code> remap that would emit a <code>Co-Authored-By</code> / <code>Generated-with</code> trailer from a benign-looking custom token. Benign trailers (<code>Reviewed-by=Alice</code>, <code>-c trailer.sob.key=Signed-off-by</code>) still pass.</li> <li><strong>tasklist-guard</strong> — fail-open when the state dir is unwritable instead of wedging.</li> <li><strong>swarm-guard</strong> — steals a zero-byte / corrupt lock instead of deadlocking, and resolves <code>vm_stat</code> by absolute path.</li> <li><strong>graphify-reminder</strong> — bounds the <code>git rev-parse</code> probe with a timeout.</li> <li><strong>sanitizers</strong> — reflected-text + terminal-control + <strong>Unicode bidi</strong> hygiene across <code>graphify-guard</code>, <code>graphify-session</code>, <code>speculation-judge</code>, and both statuslines: bidi overrides (U+202A–U+202E) and isolates (U+2066–U+2069) are now stripped so a crafted path/claim can't visually reorder terminal/model output. Regexes stay linear (no ReDoS); legit UTF-8 and branch names are preserved.</li> <li><strong>docs</strong> — accuracy fixes (<code>KB.md</code> → <code>0.21.0</code>, <code>TASK-WORK</code> >1/checks wording, README statusline + doctor list, E2E +24.x); README badges (tests / version / license / node / plugin).</li> </ul> <p>Accepted residuals (deliberate evasion of a safety net, documented not closed): <code>base64 | sh</code>, pipe-to-shell, process-substitution, deeply-nested <code>eval</code>, <code>commit -F <file></code> / editor commits, and Unicode-confusable token swaps. These require an adversary actively defeating their own guardrail and are out of scope for a fail-open static hook.</p> <h2 id="0203">0.20.3<a class="headerlink" href="#0203" title="Permanent link">¶</a></h2> <p>Add RELEASING.md — the ordered release + doc-currency checklist the agent follows on every ship (manual tagging by agent, no CD); pointer from AGENTS.md.</p> <h2 id="0202">0.20.2<a class="headerlink" href="#0202" title="Permanent link">¶</a></h2> <p>Doc currency sync — <code>llms.txt</code> + <code>plugins/anti-hall/README.md</code> were stale (predated tasklist-guard / skip-guard / rule K); added the missing hooks (<code>tasklist-guard</code>, <code>skip-guard</code>, <code>command-guard</code>, <code>swarm-guard</code>, <code>phase-tracker</code>, <code>agent-watchdog</code>), the user-override escape hatch, rule K output-presentation, the E2E test suite + CI, and the new docs (<code>CONTEXT-PRESERVATION-KB</code>, <code>TASK-WORK</code>, <code>TASKLIST-GUARD</code>, <code>KB</code>, <code>E2E-TESTING</code>). Fixed "four Stop hooks" → five. CHANGELOG + tests were already current.</p> <h2 id="0201">0.20.1<a class="headerlink" href="#0201" title="Permanent link">¶</a></h2> <ul> <li>orchestration skill gains a concise "References & context guardrails" section — points to <code>docs/CONTEXT-PRESERVATION-KB.md</code> (the swarm-researched context-discipline KB) and the findings-discipline (externalize durable findings to memory, case findings to <code>.anti-hall-progress.md</code>; compact early once externalized; context-rot sweet spot is a cadence not a length). Outcome of a deadly-loop on 3 candidate context enhancements: the other two (an always-on findings protocol line, a statusline pressure cue) were DROPPED as footprint regression / redundant with the existing color gradient + native Auto Memory; only the skill-reference survived.</li> </ul> <h2 id="0200">0.20.0<a class="headerlink" href="#0200" title="Permanent link">¶</a></h2> <p>New <code>tasklist-guard.js</code> Stop hook + a per-turn freshness note (in <code>task-tracker.js</code>) that enforce live task-list + fresh <code>.anti-hall-progress.md</code> discipline for non-trivial work, so real work is tracked and never silently dropped or declared-done by a later agent. It <strong>coexists with <code>task-guard</code></strong> (which drains declared tasks); each keeps an independent block cap so the two never compound.</p> <ul> <li><strong>When it blocks (Stop):</strong> ≥ <code>ANTIHALL_TASKLIST_WORK_THRESHOLD</code> (default 3) file-mutating actions (<code>Edit</code>/<code>Write</code>/<code>MultiEdit</code>/<code>NotebookEdit</code> + mutating <code>Bash</code>) AND (no task activity for it, OR more than one task <code>in_progress</code>, OR no fresh <code>.anti-hall-progress.md</code>). Blocks via <code>{decision:"block"}</code> (never exit 2 — plugin Stop hooks don't reliably continue on it), capped at <code>MAX_BLOCKS=3</code> cumulative per session, fully fail-open.</li> <li><strong>Progress file:</strong> <code><cwd>/.anti-hall-progress.md</code> (done/in-progress/next), must be updated this session to count fresh (window <code>ANTIHALL_PROGRESS_FRESH_MS</code>, default 30 min). Gitignored, never ships; the hook never creates it.</li> <li><strong>Per-turn freshness note:</strong> when open/stale tasks exist, <code>task-tracker</code> injects a one-line reminder (open-task count + oldest <code>in_progress</code> subject).</li> <li><strong>Escape hatch:</strong> <code>~/.anti-hall/skip.json</code> <code>{"tasklist-guard": <unix-ms expiry>}</code> (or <code>"all"</code>) via the shared skip-guard, or ask the agent to skip.</li> <li><strong>Hardened via 2 deadly-loop passes (Opus + Codex):</strong> multi-<code>in_progress</code>-only staleness (a single <code>in_progress</code> — the healthy flow — never false-blocks), emit-before-persist (a state-write failure can't retract an emitted block), <code>lstat</code> + <code>isFile</code> progress check (rejects a dir/symlink named like the file), cwd-missing fail-open, command-position-aware + broadened <code>Bash</code> work regex (ReDoS-safe), task-subject injection neutralized via <code>JSON.stringify</code>.</li> <li><strong>Coverage:</strong> 120-test E2E suite incl. 21 <code>Bash</code>-regex cases; <code>doctor</code> gains a tasklist-guard self-test (4 edits, no tasks, no progress file ⇒ block).</li> <li><strong>Accepted deferrals (fail-open by design):</strong> cumulative work counters vs the 512 KB transcript tail-clip (work before the window is unseen, can only <em>suppress</em> a block — the safe direction); sync-I/O stall on a network-mounted cwd (the 30 s hook timeout makes a non-block the outcome).</li> <li>New doc <a href="../TASKLIST-GUARD/"><code>docs/TASKLIST-GUARD.md</code></a> (usage); README + <code>docs/KB.md</code> pointers added.</li> </ul> <h2 id="0190">0.19.0<a class="headerlink" href="#0190" title="Permanent link">¶</a></h2> <p>Strengthened deadly-loop auditor instructions with explicit depth requirements to prevent surface-skimming and ensure genuine issue discovery. Both <code>deadly-loop</code> and <code>deadly-loop-multi</code> skills now mandate:</p> <ul> <li><strong>DIG DEEP.</strong> Read full implementations end-to-end, trace control and data flow across files, follow every branch and error path. Never judge from names, signatures, or diff hunks. Cite file:line ranges actually read.</li> <li><strong>ENUMERATE & SIMULATE EDGE CASES.</strong> Systematically exercise boundary conditions, empty/malformed/oversized/unicode/concurrent inputs, missing files, permission denials, clock skew, truncation, and injection-shaped data. Mentally execute or write throwaway harnesses; report predicted outcomes.</li> <li><strong>REAFFIRMED 3-TIER SEVERITY.</strong> P0/P1/P2 categories (plus EASY-WIN) ranked by heat — correctness > reliability > ergonomics.</li> <li><strong>CARRY-FORWARD DISCIPLINE.</strong> Verify each prior finding's fix resolved without regression before hunting genuinely NEW issues. Distinguish new discoveries from re-reported findings.</li> </ul> <p>For <code>deadly-loop/SKILL.md</code>: added new "B0. Auditor depth requirements" section with B1/B2 forward pointers. For <code>deadly-loop-multi/SKILL.md</code>: expanded the mandatory brief-footer and step-6 reconverge to enforce carry-forward discipline across iterations. Auditor reviews now drill into implementation to surface root causes instead of polish-layer issues.</p> <h2 id="0180">0.18.0<a class="headerlink" href="#0180" title="Permanent link">¶</a></h2> <p>Performance: trimmed the always-on SessionStart protocol (<code>verify-first-full.js</code>) footprint by condensing VERBOSE PROSE ONLY — every rule, every rationalization trigger phrase, all orchestration labels A-K, the USER OVERRIDE mechanism, and the DISCIPLINES section are preserved verbatim. The SessionStart injection dropped from 8074 B to 7474 B (~7.4 KB, ~160-180 tokens saved) with zero efficacy loss. Hardened by a new 77-test zero-dep E2E net that asserts every marker is present (so a dropped rule fails the build), plus TWO deadly-loop passes (Opus reviewer + Codex critic) that caught 8 semantic nuances the trim over-cut and restored them. Also fixed a stale "compact-matcher" comment in <code>verify-first.js</code> to match the verified no-matcher SessionStart mechanism. Net: smaller one-time injection, identical protocol.</p> <h2 id="0171">0.17.1<a class="headerlink" href="#0171" title="Permanent link">¶</a></h2> <p>Bug fix: task-tracker self-heals corrupted or future throttle state. A future or non-finite <code>lastFull</code> timestamp (from clock skew, manual edit, or corruption) previously left the throttle window permanently "within window", so the full task directive never re-showed. The throttle now detects a future-beyond-tolerance (5 min) or non-finite value as window-expired and rewrites the state to now, allowing the full task output to surface again on the next turn.</p> <h2 id="0170">0.17.0<a class="headerlink" href="#0170" title="Permanent link">¶</a></h2> <p>Tier-B guard hardening. Tightens the static command analysis so a quoted data literal is no longer mistaken for a heavy command, and unwraps a single <code>eval</code> payload so a guard sees the real command it would run — in both command-guard and git-guard.</p> <ul> <li><strong>command-guard is quote-aware.</strong> Heavy-pattern matching now distinguishes a heavy command from a heavy-looking string literal: <code>echo "npm run build"</code> / <code>printf 'go test ./...'</code> pass a quoted data argument and are no longer false-blocked in the coordinator, while an unquoted heavy command (<code>npm run build</code>), a command-substitution payload (<code>echo "$(npm run build)"</code>), and a <code>bash -c "npm run build"</code> wrapper still block. The intent is to stop flagging strings that merely <em>contain</em> a heavy verb without losing any real execution path.</li> <li><strong><code>eval</code> payload unwrapping in BOTH command-guard and git-guard.</strong> A single <code>eval "…"</code> is now unwrapped and its inner command re-scanned, so <code>eval "npm test"</code> blocks in the coordinator and <code>eval "git push -f"</code> is caught by git-guard, instead of slipping past as an opaque <code>eval</code> argument. <code>eval "echo hi"</code> / <code>eval "git status"</code> still pass.</li> <li><strong>Known residual gaps (accepted, by design).</strong> <code>base64 | sh</code>, process-substitution (<code><(…)</code>), and <code>git commit -F <file></code> trailer smuggling remain accepted defense-in-depth gaps: they cannot be closed by static inspection without over-blocking legitimate use. The guards are a safety net against the common slip, not a sandbox — a determined evasion is out of scope.</li> </ul> <h2 id="0160">0.16.0<a class="headerlink" href="#0160" title="Permanent link">¶</a></h2> <p>Round-2 Tier-A guard hardening, surfaced by the double deadly-loop on the round-1 changes. Closes the fail-closed and OOM cases that could wedge or crash a guard, and seals two guard evasion paths the round-1 hardening missed.</p> <ul> <li><strong>swarm-guard memory gate fails OPEN, not closed.</strong> The free-memory parser previously fell back to <code>os.freemem()</code> on a parse failure, which on some platforms reports near-zero and could block every spawn (fail-closed safety guard left silently disabled is the opposite of what a coordinator wants). It now SKIPS the memory gate entirely when the platform memory read can't be parsed, so a spawn is never blocked on bad telemetry.</li> <li><strong>swarm-guard lock uses an owner token — no blind <code>unlink</code>.</strong> The advisory spawn lock now writes an owner token and only releases a lock it actually owns, instead of blindly <code>unlink</code>-ing whatever lock file is present (which could stomp a concurrent holder).</li> <li><strong>Bounded transcript tail-reads (512 KB) — OOM guard.</strong> <code>speculation-guard.js</code>, <code>speculation-judge.js</code>, <code>task-guard.js</code>, and <code>graphify-reminder.js</code> previously <code>readFileSync</code>'d the entire transcript; a long session could grow that to hundreds of MB and OOM the hook. They now tail-read only the last 512 KB, which is more than enough for the recent-turn inspection each performs.</li> <li><strong>task-guard subject sanitization.</strong> The task subject is now sanitized before it is echoed into the block reason, closing a reflected-content path.</li> <li><strong>graphify rev-parse timeouts (2000 ms) + non-echoing block reason.</strong> <code>graphify-session.js</code> and <code>graphify-guard.js</code> now bound their <code>git rev-parse</code> calls at 2000 ms so a wedged git can't hang the hook; <code>graphify-guard.js</code> no longer echoes the raw command into its block reason (label only).</li> <li><strong>git-guard seals two evasion paths.</strong> It now catches a <code>+refspec</code> force-push smuggled AFTER a <code>--</code> end-of-options marker (<code>git push origin -- +main:main</code>), and blocks <code>git --config-env alias.*</code> config smuggling that could alias a benign verb to <code>push --force</code>. Normal pushes and literal <code>--</code>-separated operands are unaffected.</li> <li><strong>command-guard light exceptions + non-echoing reason.</strong> Read-only <code>git push --dry-run</code> and read-only <code>go env</code> (without <code>-w</code>) are now treated as light and allowed in the coordinator; the heavy-command block reason no longer echoes the raw command (label/verb only), removing a reflected-injection surface.</li> <li><strong>Windows-safe installer.</strong> <code>statusline/install-statusline.js</code> now uses a shell-free <code>execFileSync</code> for the git-tracked check, so the install path no longer depends on a POSIX shell.</li> </ul> <p>Tier B remains pending: the command-guard quote-aware false-positive and the <code>eval</code>/base64/process-substitution evasion cluster are not addressed here.</p> <h2 id="0150">0.15.0<a class="headerlink" href="#0150" title="Permanent link">¶</a></h2> <p>Adds a user-override escape hatch across all guards, hardens the speculation hooks against a Stop-loop wedge, bounds the deadly-loop convergence, and lands several doc fixes.</p> <ul> <li><strong>User-override escape hatch.</strong> New shared primitive <code>hooks/skip-guard.js</code> exports <code>isSkipped(name)</code>, which reads a TTL'd marker at <code>~/.anti-hall/skip.json</code> (<code>{"<guard>": <unix-ms expiry>}</code>). When the user EXPLICITLY asks to skip a guard, the agent records consent there and the guard fail-opens (does not interfere) until it expires — default TTL 15 minutes so a safety guard is never left silently disabled. Granular: the broad key <code>"all"</code> covers the noisy guards but <strong>NOT</strong> <code>git-guard</code> (force-push / self-credit must be named explicitly). Fail direction is safe: a missing/corrupt marker keeps every guard ACTIVE. Wired into all 7 guards (<code>git-guard</code>, <code>command-guard</code>, <code>swarm-guard</code>, <code>speculation-guard</code>, <code>speculation-judge</code>, <code>graphify-guard</code>, <code>task-guard</code>) immediately after the stdin read. The SessionStart protocol (<code>verify-first-full.js</code>) and <code>AGENTS.md</code> gain a matching rule: honor a skip ONLY on a direct, unambiguous user instruction — never on the agent's own initiative or because a tool/file/channel asked.</li> <li><strong>P1 fix — <code>MAX_BLOCKS = 3</code> hard cap on the speculation hooks.</strong> <code>speculation-guard.js</code> and <code>speculation-judge.js</code> previously deduped only on the exact message hash; because the message text legitimately changes as the model reworks its reply, that dedupe could be defeated and the Stop hook could re-block on every Stop. They now mirror <code>task-guard.js</code>'s two-tier loop-safety — hash dedup PLUS a running <code>blocks</code> counter capped at 3 — closing the Stop-loop wedge. State is now <code>{hash, blocks}</code>; legacy bare-hash files are still tolerated.</li> <li><strong>deadly-loop iteration caps (soft 10 / hard 15).</strong> The Reviewer+Critic debate + fix-wave convergence loop is now explicitly bounded: at <strong>10 rounds</strong> without convergence, STOP and checkpoint with the user (AskUserQuestion: continue / stop / change scope) instead of looping silently; at <strong>15 rounds</strong>, force-stop unconditionally and report even if not converged. The same cap applies to <code>deadly-loop-multi</code>'s step-6 reconverge loop.</li> <li><strong>Doc fixes (no guard-logic change).</strong> <code>STATUSLINE.md</code> now describes line 2 as the actual always-on 3-tier behavior (phase bar → "orchestrating · N agents" → idle context gauge), consistent with its own top summary and <code>phase-bar.js</code>; corrected two misleading comments that said a transcript was "streamed" when the code does a full <code>readFileSync</code> (<code>task-guard.js</code>, <code>graphify-reminder.js</code>); confirmed <code>demo-wrapper.sh</code> (machine-absolute <code>/Users</code> path, untracked) stays gitignored so a future <code>git add -A</code> cannot ship it.</li> </ul> <h2 id="0141">0.14.1<a class="headerlink" href="#0141" title="Permanent link">¶</a></h2> <p>Refines the 0.14.0 email chip: show-when-available with an opt-OUT, instead of opt-in.</p> <ul> <li><strong>The <code>✉ <email></code> chip now renders whenever the account email is available</strong> (read from <code>~/.claude.json</code>), rather than requiring <code>ANTIHALL_STATUSLINE_EMAIL=1</code>. Since a signed-in user almost always has an email, the opt-in gate was effectively "never shows unless you know the flag." Inverted to a privacy opt-OUT: set <strong><code>ANTIHALL_STATUSLINE_NO_EMAIL=1</code></strong> (or any non-<code>0</code>/<code>false</code> value) to hide it — e.g. for screenshots or screen-shares. Fail-open: no chip when the email can't be read.</li> </ul> <h2 id="0140">0.14.0<a class="headerlink" href="#0140" title="Permanent link">¶</a></h2> <p>Adds an opt-in account-email chip to the rich statusline.</p> <ul> <li><strong><code>statusline-rich.js</code> can now show a <code>✉ <email></code> chip</strong> as the last line-1 segment, reading the signed-in Claude account email from <code>~/.claude.json</code> (<code>oauthAccount.emailAddress</code>). <strong>OFF by default</strong> and gated behind the <code>ANTIHALL_STATUSLINE_EMAIL</code> env var (set to <code>1</code>/any non-<code>0</code>/<code>false</code> value to enable) — the plugin must never surface a user's email on their statusline without explicit consent. New <code>getClaudeEmail()</code> helper; pure file read, fail-open (no chip on any error). This brings the plugin's own renderer to parity with custom <code>.claude/helpers/</code> statuslines that already show the email, without making it a privacy-leaking default.</li> </ul> <h2 id="0130">0.13.0<a class="headerlink" href="#0130" title="Permanent link">¶</a></h2> <p>Adds an always-on output-presentation discipline so chat output is structured and scannable without becoming noisy.</p> <ul> <li><strong>New SessionStart rule K — "PRESENT FOR SCANNABILITY (do not overdo it)".</strong> Appended to the orchestration block in <code>verify-first-full.js</code> (and mirrored as a bullet in <code>AGENTS.md</code>). Encodes the conservative, renderer-verified subset of GitHub-flavored markdown that Claude Code's terminal actually renders: tables for comparisons/status, <strong>bold</strong> verdicts, <em>italic</em> caveats, <code>code</code> for flags/paths/commands, fenced blocks for output, at most a leading status glyph (emoji = signal, not decoration). Explicitly steers AWAY from syntax the terminal renderer drops or mangles — strikethrough, <code>[label](url)</code> link labels (paste the bare URL), nested blockquotes, task-list checkboxes — and notes that underline and per-word color do not exist in the renderer. Subset confirmed against Claude Code terminal-rendering issue reports + docs. Styling organizes, never pads: rule H (concise) still governs. Appended as rule K so existing letters A-J do not renumber. SessionStart footprint grows by ~0.5 KB.</li> <li><strong>Doc accuracy:</strong> corrected the stale "5 nudges" comments in <code>verify-first.js</code> (the <code>NUDGES</code> array has 12 entries, not 5) and refreshed the root README SessionStart footprint figure to the doctor-measured value.</li> </ul> <h2 id="0121">0.12.1<a class="headerlink" href="#0121" title="Permanent link">¶</a></h2> <p>Fixes AUDIT-REPORT-2 item #7(a): the shared global statusline base was deleted on every uninstall.</p> <ul> <li><strong><code>uninstall-statusline</code> no longer deletes the shared global base by default.</strong> <code>~/.anti-hall/base-statusline.json</code> is GLOBAL — every project whose statusLine points at the dispatcher wraps it as line 1. The uninstaller's Strategy A unconditionally <code>unlink</code>ed it after restoring the original command, so uninstalling in ONE project (even <code>--project</code> scope) silently stripped line 1 for EVERY other project still relying on it. A reference count is infeasible (no way to enumerate all projects' settings files), so the safe default is now: restore this scope's original command and LEAVE the shared base in place. An orphaned JSON is harmless; a deleted shared one is not. New opt-in <code>--purge-base</code> flag explicitly removes it for the "done with anti-hall on this whole machine" case (use only after uninstalling everywhere). Verified with a fake-HOME behavioral test: default keeps base + restores original command; <code>--purge-base</code> deletes it.</li> </ul> <h2 id="0120">0.12.0<a class="headerlink" href="#0120" title="Permanent link">¶</a></h2> <p>Closes three deferred AUDIT-REPORT-2 needs-review gaps (external reviewer + codex).</p> <ul> <li><strong>Recursive shell parsing (<code>command-guard</code> + <code>graphify-guard</code>).</strong> The segment splitter previously treated <code>$(...)</code> / backticks as plain boundaries and never inspected their CONTENTS, and never unwrapped <code>bash -c '...'</code> / <code>sh -c</code> / <code>zsh -c</code> payloads — so <code>echo "$(npm run build)"</code> and <code>bash -c "npm run build"</code> bypassed the coordinator block (and the graph-first nudge). Both guards now extract nested commands from command substitution and from shell <code>-c</code> payloads and re-apply the full heuristic (HEAVY_VERBS/HEAVY_PATTERNS/LIGHT_EXCEPTIONS for command-guard, the code-nav check for graphify-guard) to the inner commands, depth-bounded to 3 to avoid pathological input. Benign substitutions stay allowed (<code>echo "$(date)"</code> — <code>date</code> is not heavy). Fail-open preserved.</li> <li><strong>Atomic swarm-guard spawn cap.</strong> The prune→count→cap-check→append was a non-atomic read-modify-write, so concurrent spawns each read a stale pre-cap log and raced past the ceiling. It now runs inside a best-effort cross-process O_EXCL lock (<code>~/.anti-hall/swarm-spawns.lock</code>) with stale-lock steal (mtime > ~5s), bounded spin, re-read + cap-check INSIDE the lock, and release in <code>finally</code>. FAIL-OPEN if the lock can't be acquired — never deadlocks a spawn. The cap-before-append fix is retained.</li> <li><strong>Doctor 2-line assertion + new tests.</strong> The statusline self-test now requires<blockquote> <p>= 2 lines for the sample payload (which carries <code>context_window</code>, so the live context gauge on line 2 MUST render) and reports a FAILURE if only line 1 renders. Added self-tests proving the recursive-parse fix: command-guard now BLOCKS (in coordinator) <code>echo "$(npm run build)"</code> and <code>bash -c "npm run build"</code>, and still ALLOWS the benign <code>echo "$(date)"</code>.</p> </blockquote> </li> </ul> <h2 id="0113">0.11.3<a class="headerlink" href="#0113" title="Permanent link">¶</a></h2> <p>Precedence-aware install-statusline. Per-project install now writes <code>.claude/settings.local.json</code> (highest precedence + gitignored) instead of <code>settings.json</code>, so a committed project statusLine can no longer shadow it. The installer checks the statusLine across user/project/local scopes and reports shadowing or already-installed; resolves a STABLE marketplace dispatcher path (never the versioned cache path, which breaks on update); and auto-gitignores <code>.claude/settings.local.json</code>.</p> <h2 id="0112">0.11.2<a class="headerlink" href="#0112" title="Permanent link">¶</a></h2> <p>Double-deadly-loop (4-auditor) final-gate fixes. <code>git-guard</code> and <code>command-guard</code>: <code>sudo</code> with option flags no longer leaks — <code>sudo -u deploy git push --force</code> and <code>sudo -u deploy npm install</code> previously resolved their effective verb to <code>-u</code> and slipped past the guards; both now block (flag/operand-skip mirrors the env/timeout/nice handling). <code>git-guard</code>: a force form baked into an inline alias BODY (<code>git -c alias.p='push --force origin main' p</code>) was dropped — now the alias body's tokens are force-checked. <code>graphify-guard</code>: <code>segmentVerb</code> now skips wrapper words so <code>sudo rg secret</code> is still detected for the (non-blocking) graph- first nudge. Docs: corrected the skill count (llms.txt "five" → "seven workflow skills"; plugins README primer "7 skills" → "core 4 skills" to match verify-first-full.js; Features table notes deadly-loop-multi/install-statusline/doctor). See docs/AUDIT-REPORT-2.md for the reconciled findings, rejected false-positives, and deferred needs-review items. Fail-open and subagent-allow preserved; doctor green (36 checks).</p> <h2 id="0111">0.11.1<a class="headerlink" href="#0111" title="Permanent link">¶</a></h2> <p>Fix <code>deadly-loop-multi</code> SKILL.md YAML frontmatter: the <code>description</code> contained an unquoted inner <code>Multiplier:</code> (colon-space), which YAML parses as a mapping value — the skill failed to load ("mapping values are not allowed in this context"). Replaced the colon with a dash.</p> <h2 id="0110">0.11.0<a class="headerlink" href="#0110" title="Permanent link">¶</a></h2> <p>Cut the plugin's OWN context footprint (it was growing the conversation every turn — the exact thing the plugin warns against). Root cause (researched): Claude Code injects <code>UserPromptSubmit</code> additionalContext into the transcript EVERY turn and it accumulates (see anthropics/claude-code#40216), so a long, repeated per-turn directive is a real token drain.</p> <ul> <li><strong><code>task-tracker.js</code> throttled:</strong> injects the FULL task-discipline directive only on the first turn of a session (and once per ~6h window), then a SHORT one-line reminder after, via <code>~/.anti-hall/task-tracker-<session>.json</code> state. Fail-open to full on any state error. Steady-state per-turn injection dropped ~68% (≈693 B → ≈223 B).</li> <li><strong><code>verify-first-full.js</code> (SessionStart) tightened</strong> ~13% with no rule removed (Iron Law, full rationalization table, orchestration A–J, anti-speculation tiers, anti-sycophancy all intact). <code>verify-first.js</code> per-turn nudge was already one short line.</li> <li><strong><code>doctor.js</code> adds a "Context footprint" section</strong> reporting the SessionStart / per-turn / per-Stop injection sizes in bytes + estimated tokens, so the cost is measurable.</li> </ul> <p>Future levers (noted, not yet done): move the static protocol to CLAUDE.md / SessionStart-only and merge the four Stop hooks into one to further shrink per-turn overhead.</p> <h2 id="0100">0.10.0<a class="headerlink" href="#0100" title="Permanent link">¶</a></h2> <p>Audit-fix batch (from a 2-Opus + 2-Codex review) + context-protection discipline.</p> <p>Guards: - <strong>command-guard</strong> now evaluates PER SEGMENT (quote-aware split on <code>; && || |</code>, env-prefix + wrapper skip, cross-platform basename). Fixes real false-negatives where heavy commands bypassed: <code>cd app && npm test</code>, <code>git status && npm run build</code>, <code>FOO=1 docker build .</code>. - <strong>swarm-guard</strong> checks the spawn-rate cap BEFORE appending/persisting the timestamp, so blocked retries no longer extend the block window; state moved to <code>~/.anti-hall/</code>. - <strong>speculation-guard</strong> regex now catches "should be fine" (was excluded by a lookahead). - <strong>graphify-guard</strong> <code>/graphify</code> exemption is now segment/verb-aware (a substring like <code>echo /graphify && rg secret</code> no longer exempts the search); state → <code>~/.anti-hall/</code>. - <strong>git-guard</strong> resolves inline <code>-c alias.x=push</code> so aliased force-push is caught. - <strong>task-guard / graphify-reminder</strong> state relocated to <code>~/.anti-hall/</code> for cross-runner consistency.</p> <p>Doctor: added a Graphify health section + self-tests proving the command-guard per-segment fix and the speculation-guard "should be fine" catch.</p> <p>Discipline (protect the orchestrator's context — a bloated main thread degrades the model and induces the very hallucination this plugin prevents): - Delegate not just heavy commands but <strong>broad reads / Grep / Glob / code-nav searches</strong> to subagents; inline only a specific known-file read. Added to orchestration, AGENTS.md, verify-first-full.js. - <strong>Graphify-first:</strong> ensure the graph is fresh then QUERY it before raw search and before feature-launch analysis. - <strong>AFK goal-anchor (drift watcher):</strong> re-check work against the locked goal each cycle and course-correct on drift; only deviate when the user explicitly redirects.</p> <p>Docs synced (version drift, autoUpdate wording, 7-skill count, STATUSLINE 3-tier line 2, "latest OpenAI Codex" not a pinned version, broken MODEL-POLICY links) and a consolidated <code>docs/AUDIT-REPORT.md</code> written (includes the reconciled demo-wrapper.sh false-positive).</p> <h2 id="090">0.9.0<a class="headerlink" href="#090" title="Permanent link">¶</a></h2> <p>New <code>deadly-loop-multi</code> skill — double / triple / quadruple deadly loop.</p> <p>Scales the standard 1+1 deadly-loop to N parallel reviewers + N parallel critics with diversified lenses, then reconciles + validates into ONE consolidated report.</p> <ul> <li><strong>double</strong> = 2 Opus + 2 Codex, <strong>triple</strong> = 3+3, <strong>quadruple</strong> = 4+4. The tier is named by the user or auto-selected by job complexity × sensitivity (higher tier for security/schema/cross-repo/release work).</li> <li><strong>Always half-and-half cross-model:</strong> half the auditors are the latest Codex (a different model finds different bugs); if Codex is unavailable, substitute the latest Opus with a divergent persona — never drop below the full 2N. Model versions are deliberately NOT pinned ("latest Opus / latest Codex") so the skill survives new releases.</li> <li><strong>Runs as a swarm</strong> via the Opus Dynamic-Workflow primitives (KB §11 / docs/opus-4-8-swarm.md): a parallel fan-out of the 2N auditors feeding a reconcile + validate synthesis stage. The coordinator validates each finding against the code itself (agreement raises confidence, but evidence decides) and reconciles conflicts — so a single agent's false positive does not make it into the report.</li> </ul> <h2 id="081">0.8.1<a class="headerlink" href="#081" title="Permanent link">¶</a></h2> <p>Doctor: add a behavioral statusline check. Beyond confirming a statusLine is configured, the doctor now spawns the dispatcher with a sample payload and asserts it actually RENDERS (reports the line count — line 1 + live line 2), and validates that <code>statusline-rich.js</code> (the line-1 renderer) is present and syntax-valid. So a broken or missing renderer is caught, not just a missing setting.</p> <h2 id="080">0.8.0<a class="headerlink" href="#080" title="Permanent link">¶</a></h2> <p>Doctor (health check + live guard self-tests) + documentation refresh.</p> <p>Addresses external-review feedback: a guardrail plugin needs a way to prove it is actually running and that the guards actually fire — not just that files exist.</p> <ul> <li><strong>New <code>hooks/doctor.js</code> + <code>/anti-hall:doctor</code> skill.</strong> Reports Node version (flags < 18, which makes the hooks silently no-op), plugin version, every registered hook's presence + syntax, and the statusline install status. Crucially it runs LIVE behavioral self-tests: it spawns the real guards with crafted payloads and asserts exit codes — git-guard blocks force-push + AI self-credit and allows <code>git status</code>; command-guard blocks heavy commands in the coordinator but ALLOWS them in a subagent (payload <code>agent_id</code>); swarm-guard allows a normal spawn. Exits non-zero on any critical failure (CI-friendly); <code>--quiet</code> prints just the verdict.</li> <li><strong>Docs synced to the current feature set.</strong> README rewritten (modern, explained, with the two-layer model, the statusline, and AFK mode), <code>llms.txt</code> updated to list ALL hooks/skills (it and the README previously undercounted — e.g. omitted command-guard, swarm-guard, graphify-guard, phase-tracker), and AGENTS.md gained the AFK autonomy contract.</li> </ul> <h2 id="070">0.7.0<a class="headerlink" href="#070" title="Permanent link">¶</a></h2> <p>Automatic swarm-progress tracking + AFK-mode autonomous driver.</p> <p>Problem: the phase bar only updated if the coordinator manually called <code>phase.js</code>, and the autonomous-driver template never called it — so a running swarm/feature-launch showed no progress (verified gap; <code>CONTINUE-HERE.md</code> listed it as a TODO).</p> <ul> <li><strong>New <code>phase-tracker.js</code> hook</strong> (PreToolUse Agent/Task) — records every subagent spawn to a HOMEDIR log (<code>~/.anti-hall/agent-spawns.log</code>), never blocks, fail-open. Registered after swarm-guard so it logs only real spawns.</li> <li><strong><code>phase-bar.js</code> now has 3 tiers</strong> for line 2: (1) coordinator-set semantic phase -> rich phase bar; (2) recent spawns (auto) -> animated <code>orchestrating · N agents active</code> bar; (3) idle -> context-window gauge. So a swarm is visible with ZERO coordinator effort. (A registered hook needs a session restart to activate.)</li> <li><strong>AFK mode</strong> (<code>AUTONOMOUS-DRIVER-PROMPT.md</code>): wired the <code>phase.js set/step/agents/ advance/clear</code> calls into the per-phase loop, and added the AFK autonomy contract — the driver never returns to the owner or stops except for an ABSOLUTELY-DESTRUCTIVE hard gate; it collects data instead of pausing, and resolves confusion with a deadly-loop rather than asking the (away) owner.</li> </ul> <h2 id="060">0.6.0<a class="headerlink" href="#060" title="Permanent link">¶</a></h2> <p>Always-on line 2 (hybrid bar). The plugin's statusline now ships BOTH lines as a complete two-line statusline, and line 2 is always present (never blank): - During an active orchestration run -> the live phase progress bar (as before). - When idle -> a context-window usage bar rendered from the session JSON: <code>[███████████◐────────] 56% context</code> (· <code>used/max tokens</code> when the harness provides counts), color-coded green/yellow/red at <=70/70-89/>=90.</p> <p><code>statusline.js</code> now passes the session stdin through to the line-2 renderer (<code>phase-bar.js</code>) so the context bar has real data; <code>phase-bar.js</code> renders the phase bar when a fresh phase-state exists and the context bar otherwise. Fail-open: if neither source is available, line 2 is simply omitted.</p> <h2 id="051">0.5.1<a class="headerlink" href="#051" title="Permanent link">¶</a></h2> <p>Phase bar auto-hides stale state. The line-2 phase bar reads <code>~/.anti-hall/phase-state.json</code>; an orchestration run that ended without calling <code>phase.js clear</code> left an ORPHAN state file, so the bar showed a frozen, stale phase indefinitely (e.g. a "22h" phase that never moved). <code>phase-bar.js</code> now treats a state file whose mtime is older than 30 minutes as absent — active runs rewrite the file on every set/advance/step/agents call (well under the window), so live runs always render, but orphaned leftovers no longer linger. Fail-open preserved.</p> <h2 id="050">0.5.0<a class="headerlink" href="#050" title="Permanent link">¶</a></h2> <p>Rich statusline + on-demand install skill.</p> <ul> <li>New <code>statusline/statusline-rich.js</code> — a generic, project-agnostic rich line-1 renderer (project name from cwd, git branch/worktree/stash/staged-modified-untracked, ahead/behind, model, effort, subagent count, session duration, context-window %, cost, and the GSD <code>.planning</code> phase when present). Pure Node, fail-open, no project/user specifics. The dispatcher (<code>statusline.js</code>) now uses it as the primary own-dispatch line-1 renderer, falling back to the monorepo/simple renderers if it yields nothing.</li> <li>New <code>install-statusline</code> skill — installs the statusLine entry on demand (user scope by default for a global bar, <code>--project</code> for the current repo only), with a reminder that Claude Code reads <code>statusLine</code> only at startup so a restart is required.</li> <li><code>install-statusline.js</code> no longer clobbers an existing GLOBAL <code>base-statusline.json</code> (overwriting it changed line 1 for OTHER projects that rely on it). Existing base is kept; repos without their own helper fall through to the rich renderer.</li> </ul> <h2 id="047">0.4.7<a class="headerlink" href="#047" title="Permanent link">¶</a></h2> <p>Fix swarm-guard false-positive memory-pressure block. The memory check used <code>os.freemem() / os.totalmem() < 4%</code>, but on macOS and Linux <code>os.freemem()</code> reports only truly-free pages and EXCLUDES reclaimable cache (inactive / speculative / file-backed). On a healthy 64 GB Mac it read ~2 GB "free" (< 4%) while ~24 GB was actually available and memory pressure was green with zero swap — so legitimate agent spawns were blocked, defeating delegation just like the 0.4.6 command-guard bug.</p> <p>Fix: compute REAL available memory per-platform — macOS via <code>vm_stat</code> (free + inactive + speculative pages, honoring the actual page size: 16384 on Apple Silicon, not a hardcoded 4096), Linux via <code>/proc/meminfo</code> MemAvailable, and <code>os.freemem()</code> as the fallback where it is accurate (Windows) or on any parse error. Verified on a 64 GB Apple Silicon Mac: old calc 6.5% (would block), corrected 36.5% (no block), matching Activity Monitor. Fail-open preserved.</p> <h2 id="046">0.4.6<a class="headerlink" href="#046" title="Permanent link">¶</a></h2> <p>Fix command-guard blocking SUBAGENTS (not just the coordinator) under cmux and other launchers that wrap <code>claude</code> — which crippled the orchestration plugin's entire purpose: if subagents are also blocked from running heavy commands, there is nothing left to delegate TO, and the swarm deadlocks.</p> <p>Root cause (verified empirically this session by capturing real PreToolUse payloads): the old <code>isCoordinator()</code> relied solely on <code>CLAUDE_CODE_ENTRYPOINT === "agent_tool"</code> to detect subagents. That env var is only set on the subagent PROCESS in a vanilla <code>claude</code> CLI. Under cmux, subagents inherit the parent's exact environment (same <code>CLAUDE_CODE_ENTRYPOINT=cli</code>, same <code>CLAUDE_CODE_SESSION_ID</code>, same PID), so every subagent looked like the coordinator and got blocked.</p> <p>Fix: detect subagents from the hook PAYLOAD instead of the process env. Claude Code injects <code>agent_id</code> and <code>agent_type</code> into the PreToolUse payload for Task-tool subagents; the top-level coordinator's payload has neither. <code>isCoordinator(payload)</code> now treats a command as a subagent (allow) if the payload carries <code>agent_id</code>/<code>agent_type</code> OR the entrypoint is <code>agent_tool</code> (vanilla-CLI fallback). <code>main()</code> parses the payload before the context check. This works in BOTH environments (cmux and vanilla CLI). Fail-open preserved on any ambiguity.</p> <p>Verified: coordinator <code>node x.js</code> -> still blocked; the SAME command from a Task-tool subagent -> runs (payload carried <code>agent_id</code>/<code>agent_type</code>, matching the spawned agent id).</p> <h2 id="045">0.4.5<a class="headerlink" href="#045" title="Permanent link">¶</a></h2> <p>Fix task-guard false-block: completed tasks were over-counted as pending, causing the Stop hook to block on sessions with zero genuinely-open tasks.</p> <p>Root cause (verified against real transcripts): two interacting key-mismatch bugs:</p> <ol> <li> <p><code>TaskCreate</code> input in the real harness has no <code>id</code> or <code>task_id</code> field — the harness assigns a sequential numeric id (1, 2, 3...) but returns it only in the tool_result string <code>"Task #N created successfully: <subject>"</code>. The old code fell back to <code>Date.now() + random</code>, generating a different random key for each create.</p> </li> <li> <p><code>TaskUpdate</code> input uses the field <code>taskId</code> (camelCase) — the old code read <code>inp.id || inp.task_id</code>, both always <code>null</code>, so no update ever matched any create. All 34 task creates stayed at status <code>pending</code> forever.</p> </li> </ol> <p>Fix: parse tool_result strings to extract the harness-assigned numeric id, store TaskCreate provisionally under the tool_use wire id, then remap to the numeric id when the result is seen. Read <code>inp.taskId</code> first in TaskUpdate (fallback to <code>inp.id</code> and <code>inp.task_id</code> for alternate harnesses). Fail-open on any parse error. A cleared or all-completed task list yields zero open tasks and no block.</p> <p>Verified by 5 functional tests: 3-create-all-complete -> no block; 1-pending -> block with correct task name; TodoWrite all-completed -> no block; empty transcript -> no block; 34-creates-all-completed (exact real-scenario replay) -> no block.</p> <h2 id="044">0.4.4<a class="headerlink" href="#044" title="Permanent link">¶</a></h2> <p>Fix plugin load error — removed redundant manifest <code>hooks</code> reference; hooks/hooks.json is auto-loaded by Claude Code, so referencing it explicitly caused a duplicate-hooks-file load error (per /doctor). Hooks unchanged and still active.</p> <h2 id="043">0.4.3<a class="headerlink" href="#043" title="Permanent link">¶</a></h2> <p>Real statusline installer (wraps existing user statusline + scope-aware + uninstall). statusline.js now supports base-command wrap.</p> <ul> <li><strong><code>install-statusline.js</code> (new):</strong> Interactive installer (--user/--project scope, base-statusline.json config, backup + restore, dedup safety). Wraps the user's EXISTING statusline.json / statusline command (if present) as line 1, adds anti-hall phase bar as line 2. Scope-aware: --user writes to ~/.anti-hall/; --project writes to ./.anti-hall/. Preserves .ai-generated-index, .claude-plugin, and other dotfiles on uninstall.</li> <li><strong><code>uninstall-statusline.js</code> (new):</strong> Restores original statusline from backup, removes anti-hall phase state.</li> <li><strong><code>statusline.js</code> (updated):</strong> base-command wrap mode: reads ~/.anti-hall/base-statusline.json (schema: <code>{command: "..."}</code>) and dispatches it to shell, captures line 1, appends phase bar as line 2. Falls back to own dispatch (Claude | branch | repo) when config absent.</li> <li><strong><code>STATUSLINE.md</code> (updated):</strong> usage examples for install/uninstall; wrap behavior and base-statusline.json config schema.</li> </ul> <h2 id="042">0.4.2<a class="headerlink" href="#042" title="Permanent link">¶</a></h2> <p>Opt-in semantic speculation judge (LLM-evaluated Stop hook, off by default) covering the confident-inference-as-fact gap that the lexical Tier 2 guard cannot catch.</p> <ul> <li><strong><code>speculation-judge.js</code> (Stop hook, new, OPT-IN):</strong> semantic judgment tier that calls <code>claude-haiku-4-5</code> via the Anthropic API to evaluate whether the last assistant message asserts an unverified factual claim with no hedge word and no acknowledgment. Covers the gap left by <code>speculation-guard.js</code> (Tier 2), which catches hedged speculation but cannot catch a confidently-stated inference-as-fact that uses no hedge word at all (e.g., "The cause is the old build artifact." with zero hedging). Enabled ONLY when <code>ANTIHALL_SEMANTIC_JUDGE=1</code> is set in the environment; exits 0 immediately (zero cost, zero latency, zero network activity) when unset — the default. Also requires <code>ANTHROPIC_API_KEY</code>; fails-open (exits 0) when the key is absent or the API call fails for any reason. Judge prompt instructs conservative evaluation: allows honest hedging, quoted text, hypotheticals, plans, and general software knowledge; blocks only definitive unverified factual assertions. Loop-safe: hashes message text with a <code>":judge"</code> namespace suffix (separate from Tier 2's hash space); blocks at most once per distinct message, never wedges. Fail-open on any parse/read/write/API error. Cost/latency when enabled: ~$0.0001-0.001 per turn at Haiku rates + ~1-3 s per Stop. Misfire caveat: LLM judges can false-positive; the conservative prompt reduces this but does not eliminate it — disable <code>ANTIHALL_SEMANTIC_JUDGE</code> if misfires are disruptive.</li> <li><strong><code>hooks.json</code>:</strong> <code>speculation-judge.js</code> registered as the fourth Stop hook (timeout 30 s). It is a behavioral no-op in the default configuration (env var unset = exits 0 immediately). Stop hook order: <code>task-guard</code> (1st, highest-stakes), <code>graphify-reminder</code> (2nd), <code>speculation-guard</code> (3rd, Tier 2 lexical), <code>speculation-judge</code> (4th, Tier 3 semantic opt-in).</li> <li><strong><code>plugin.json</code>:</strong> version bumped 0.4.1 -> 0.4.2.</li> <li><strong><code>README.md</code>:</strong> speculation-guard entry in features table updated to label it Tier 2; <code>speculation-judge</code> entry added (OPT-IN). "Three Stop hooks coexist" note updated to four. "Known limit" paragraph replaced with a three-tier enforcement table (Tier 1: protocol/always-on/zero-cost; Tier 2: lexical/on-by-default/zero-cost; Tier 3: semantic/opt-in/LLM-cost). New "speculation-judge (Tier 3, OPT-IN)" section documenting enable instructions (shell profile and settings.json env block), what it catches, the fail-open/loop-safe/misfire/cost-latency properties.</li> <li><strong><code>AGENTS.md</code>:</strong> "speculation-guard (Stop hook)" section replaced with an "Anti- speculation enforcement: three tiers" section covering all three tiers, then dedicated subsections for Tier 2 (lexical, always-on) and Tier 3 (semantic, OPT-IN) with the same enable/cost/misfire/fail-open/loop-safe notes as the README.</li> </ul> <h2 id="041">0.4.1<a class="headerlink" href="#041" title="Permanent link">¶</a></h2> <p>Strengthened no-speculation Iron Law with inference-as-fact ban + hedged-speculation ban; added per-turn nudges; added <code>speculation-guard.js</code> Stop hook (lexical enforcement, block-once, fail-open).</p> <ul> <li><strong><code>verify-first-full.js</code> rationalization table (Iron Law hardening):</strong> added two explicit entries that were absent from prior versions: (a) <code>"X is happening because Y" / a clean causal story assembled from a couple of real facts -> that is an INFERENCE presented as fact</code> — bans confident inference-as-fact even when no hedge word is used; (b) <code>"very plausibly" / "likely" / "presumably" / "I suspect" / "I think" / "my guess is" / "it must be" -> hedging does NOT make a guess safe; it just disguises it</code> — explicitly bans hedged speculation. Both entries already existed in the RATIONALIZATION TABLE in 0.4.0; this entry documents them as rationale for the speculation-guard hook boundary.</li> <li><strong><code>verify-first.js</code> (per-turn nudge):</strong> the nudge at index 2 now explicitly names the hedge-word ban: <code>'likely' / 'plausibly' / 'I suspect' / 'I think' / 'it must be' = a guess in disguise. Hedging doesn't make it safe. Pull the data, or say 'I don't know - here's what I'd check'.</code> This was already present in the NUDGES array; documented here as the per-turn enforcement layer that pairs with the new Stop hook.</li> <li><strong><code>speculation-guard.js</code> (Stop hook, new):</strong> lexical speculation guard. Extracts the last assistant message from <code>transcript_path</code> (JSONL), scans for 15 speculation markers (hedge words: <code>very plausibly</code>, <code>plausibly</code>, <code>presumably</code>, <code>I suspect</code>, <code>my guess</code>, <code>I'd guess</code>, <code>I bet</code>, <code>likely</code>, <code>probably</code>, <code>must be</code>, <code>should be</code> [not <code>should I</code>], <code>seems to be</code>, <code>appears to be</code>, <code>I think it's</code>, <code>my hunch</code>), suppresses the block if the same message contains an evidence/uncertainty acknowledgment (<code>verified</code>, <code>I don't know</code>, <code>haven't checked</code>, <code>not verified</code>, <code>unverified</code>, <code>let me verify</code>, <code>I'll check</code>, <code>I will check</code>, <code>need to confirm</code>, <code>to confirm</code>, <code>file.ext:line</code> citation, <code>running</code>, <code>per the data</code>, <code>the data shows</code>). Block-once: hashes the message text and stores the hash in <code>~/.anti-hall/speculation-guard-state-<session>.json</code>; if the same hash was already blocked, exits 0 (nudge fires once, never wedges). Fail-open: any parse/read/write error exits 0 silently. Block reason names the matched marker and instructs the model to verify or explicitly flag as unverified. Registered in <code>hooks.json</code> as the third Stop hook (after <code>task-guard</code> and <code>graphify-reminder</code>). <strong>Known limit:</strong> catches hedged speculation (hedge word present) but not confident inference-as-fact with no hedge word. A semantic LLM-judge tier is architecturally possible but not shipped by default (cost/latency tradeoff; documented in README + AGENTS.md as opt-in design path).</li> <li><strong><code>hooks.json</code>:</strong> <code>speculation-guard.js</code> registered under <code>Stop</code> as third entry (timeout 30 s). Stop now has 3 hooks: <code>task-guard</code> (highest-stakes, first), <code>graphify-reminder</code> (second), <code>speculation-guard</code> (third).</li> <li><strong><code>plugin.json</code>:</strong> version bumped 0.4.0 -> 0.4.1.</li> <li><strong><code>README.md</code> + <code>AGENTS.md</code>:</strong> <code>speculation-guard</code> added to features table, new "speculation-guard" how-it-works section (markers list, acknowledgment suppression, loop-safe mechanism, known limit, opt-in semantic tier note); three-Stop-hook coexistence note updated from two to three hooks.</li> </ul> <h2 id="040">0.4.0<a class="headerlink" href="#040" title="Permanent link">¶</a></h2> <p>Rewritten feature-launch workflow + agent-watchdog + statusline phase.js wiring + KB merge.</p> <ul> <li><strong><code>feature-launch</code> workflow rewritten:</strong> plan-mode → deadly-loop-the-plan → loopback → self-heal; simpler phase loop; analyze-work fan-out; 4.8-fanout/synthesis/gate integration.</li> <li><strong><code>agent-watchdog.js</code>:</strong> new hook for heartbeat enforcement, babysit/backoff/kill logic. Polls <code>~/.anti-hall/agents/<id>.json</code> every 20 min; kills idle/hung agents; integrates with phase.js.</li> <li><strong><code>orchestration/SKILL.md</code> + <code>AGENTS.md</code>:</strong> updated with new agent-watchdog semantics, heartbeat convention (mtime update = alive signal), and agent supervisor responsibilities.</li> <li><strong><code>statusline/phase.js</code> wiring:</strong> orchestrator-main integration; calls phase.js (set/advance/step/agents/clear) from feature-launch as phases progress; terminal bar reflects real run state.</li> <li><strong><code>docs/KB-claude-codex.md</code>:</strong> merged knowledge base (67 sources): GSD methodology, superpowers, deadly-loop, 4.8-swarm patterns, orchestration discipline, graphify workflow. Consolidated from <code>.gsd/</code> and <code>superpowers/</code> skill refs.</li> <li><strong>Version bump 0.3.11 -> 0.4.0.</strong></li> </ul> <h2 id="0311_1">0.3.11<a class="headerlink" href="#0311_1" title="Permanent link">¶</a></h2> <p>Phase.js writer (real data source for the phase bar) + colored line-2 palette + full agent label.</p> <ul> <li><strong><code>phase.js</code> (statusline writer):</strong> new executable that writes phase-state.json as the orchestrator / feature-launch skill progresses. Commands: <code>set</code> (start phase), <code>advance</code>/<code>step</code>/<code>agents</code> (update in-flight), <code>clear</code> (hide bar). Writes to <code>~/.anti-hall/phase-state.json</code> (home dir, consistent across all processes). Fail-open: any error exits 0 without throwing.</li> <li><strong><code>statusline/STATUSLINE.md</code>:</strong> new "phase.js — Phase state writer" section documenting the data source, all 6 commands (set, advance, step, agents, update, clear), usage examples, and state-file schema (required + optional fields).</li> <li><strong><code>phase-bar.js</code> (colored palette):</strong> code (magenta/cyan), description (white), count (cyan), timer (yellow >20m, else dim), agents (blue "N agents"), step (dim). Renders as: <code>[bar] NN% | CODE - Desc done/total | timer agents step</code>.</li> <li><strong>Version bump 0.3.10 -> 0.3.11.</strong></li> </ul> <h2 id="0310_1">0.3.10<a class="headerlink" href="#0310_1" title="Permanent link">¶</a></h2> <p>Spinner repositioned inside the progress bar at the frontier. Phase-bar now uses box-drawing glyphs (█ for filled, ─ for empty) and a rotating half-disc spinner (◐◓◑◒) positioned at the progress frontier. Layout: <code>[████████◐────────────] 40% | P2 - Desc done/total | extras</code>.</p> <ul> <li><strong><code>phase-bar.js</code> (statusline):</strong> spinner (◐◓◑◒) now rendered at the progress frontier INSIDE the bar, not after. Filled segment uses █ (U+2588), empty segment uses ─ (U+2500). Spinner is a rotating half-disc (◐◓◑◒ U+25D0-D2) that advances every 125ms. Layout remains <code>[bar] NN% | CODE - Desc done/total | extras</code> with all styling preserved.</li> </ul> <h2 id="039">0.3.9<a class="headerlink" href="#039" title="Permanent link">¶</a></h2> <p>Phase-bar statusline enrichment: wider bar (20 chars), live percentage, longer description (cap 32 chars), and optional extras (elapsed time in yellow when >20m, active agent count, current step) rendered when present in phase-state.json.</p> <ul> <li><strong><code>phase-bar.js</code> (statusline):</strong> expanded bar width from 16 to 20 chars; added percentage render (e.g., "40%") in yellow after the bar; description cap increased from 16 to 32 chars with ellipsis on overflow; optional extras appended in dim text when phase-state includes <code>elapsed</code>, <code>agents</code>, and <code>step</code> keys (e.g., "3m 3ag rendering bar"). Elapsed times >20m highlighted in yellow (e.g., "[33m23m[0m") to signal long-running phases. Backward-compatible: phase-state without these keys renders as before.</li> </ul> <h2 id="037">0.3.7<a class="headerlink" href="#037" title="Permanent link">¶</a></h2> <p>Consolidated enforcement wave: command-delegation as the top always-on rule, active task-draining, swarm-guard, graphify-guard, concise-communication note, and .graphifyignore.</p> <ul> <li><strong><code>command-guard.js</code> (PreToolUse Bash, coordinator-only):</strong> new hook that BLOCKS heavy/long/state-changing commands (build, test, deploy, push, pull, install, migrate, dumps, bulk scripts) when running in a coordinator context, requiring the model to delegate them to a subagent instead. Detection uses <code>CLAUDE_CODE_ENTRYPOINT</code>: <code>agent_tool</code> = subagent (pass-through); <code>cli</code>/<code>vscode</code>/<code>jetbrains</code>/etc. = coordinator (block). Fail-open on absent/unknown entrypoint. Registered in hooks.json. COORDINATOR-VS-SUBAGENT INVESTIGATION: <code>CLAUDE_CODE_ENTRYPOINT</code> is a documented Claude Code env var set to <code>agent_tool</code> when spawned via the Task tool, and inherited by hook child processes — this is a reliable signal. The hook is SHIPPED.</li> <li><strong><code>swarm-guard.js</code> (PreToolUse Agent/Task):</strong> new hook, OS-agnostic pure Node. Tracks spawn timestamps under <code>os.tmpdir()/anti-hall/swarm-spawns.log</code>; prunes entries older than 60s; blocks if spawns in last 60s >= 20 (CAP). Secondary check: blocks if <code>os.freemem()/os.totalmem() < 4%</code> (critical memory pressure). Both thresholds are conservative. Fail-open on any error. Registered for <code>Agent</code> and <code>Task</code> matchers in hooks.json.</li> <li><strong><code>graphify-guard.js</code> (PreToolUse Grep/Glob/Bash):</strong> new hook. If a graphify graph exists (<code>graphify-out/</code> or <code>.planning/graphs/</code> at cwd or git toplevel), blocks the FIRST code-navigation search of the session (Grep tool, Glob tool, or Bash with grep/rg/ag/find/git-grep as first verb) and redirects to <code>/graphify query</code>. Blocks ONCE per session per project (loop-safe via <code>os.tmpdir</code> marker); second call is always allowed. Graphify-query Bash commands (<code>/graphify</code>) are explicitly excluded. No-op when no graph is present. Fail-open. Registered for <code>Grep</code>, <code>Glob</code>, <code>Bash</code>.</li> <li><strong><code>.graphifyignore</code> (repo root):</strong> new file — excludes <code>graphify-out/</code>, <code>.planning/graphs/</code>, <code>node_modules/</code>, <code>dist/</code>, lock files, <code>.git/</code>, and common generated/build patterns from graphify indexing.</li> <li><strong><code>verify-first-full.js</code> (SessionStart):</strong> ORCHESTRATION DISCIPLINE block restructured. Command-delegation moved to item A as the TOP RULE with explicit wording: "NEVER run verbose/long/state-changing commands inline... ALWAYS delegate to a subagent... never fill the main context with raw command output — the most counterproductive thing a coordinator can do." Active task-draining added as item C: "pick up pending tasks and dispatch subagents to finalize them; run INDEPENDENT tasks in parallel (up to the concurrency cap, ~min(16, cores-2)); never spawn unbounded agents... a runaway swarm can make the OS unusable." Concise-communication added as item H: "Communicate concisely: enough to convey meaning, not pages; offer to expand if the user wants more detail." Always-apply disciplines summary updated to reflect command-delegation top rule, task-draining, concurrency cap, concise note.</li> <li><strong><code>verify-first.js</code> (per-turn nudge):</strong> added two new nudges to the rotation — explicit command-delegation top rule; concise-communication note. Concurrency-cap language added to the task-list nudge. Hash-mod rotation auto-scales.</li> <li><strong><code>task-guard.js</code> (Stop):</strong> block reason now includes active task-draining instruction: "pick up pending tasks... run independent tasks in parallel (up to the concurrency cap, ~min(16, cores-2)); do not let tasks sit neglected."</li> <li><strong><code>AGENTS.md</code> + plugin <code>README.md</code>:</strong> command-delegation top rule added to orchestration section; concurrency cap added; active task-draining added; concise- communication added; "Recommended companion: graphify" section added (soft note — hooks no-op without it, no hard dependency).</li> <li><strong>Version bump 0.3.6 -> 0.3.7.</strong></li> </ul> <h2 id="036">0.3.6<a class="headerlink" href="#036" title="Permanent link">¶</a></h2> <p>Promote TWO disciplines to always-on ENFORCED via the hook layer, while keeping TWO skills conditional. Root-cause and orchestration now fire every session/turn; deadly-loop and feature-launch remain invoked-on-match.</p> <ul> <li><strong><code>verify-first-full.js</code> (SessionStart):</strong> added an always-apply ORCHESTRATION DISCIPLINE block framed as a BIAS TOWARD DELEGATION — non-blocking main thread; priority-sorted task list capturing every request and interruption; default to delegating any work that touches files/tools/commands/search/build/test or could balloon (avoid the eager "I'll just do it inline" trap), handle inline only genuinely atomic things (a direct answer, a single known-line read, the coordinator's own synthesis/decisions), delegate immediately if a quick inline task balloons; parallel agents when independent; noisy commands via a cheap-model subagent (Haiku/Codex) off-thread; report/synthesize. Reframed the skill primer into "ALWAYS APPLY (enforced): root-cause + orchestration + anti-sycophancy disciplines" vs "INVOKE WHEN IT MATCHES: /anti-hall:deadly-loop, /anti-hall:feature-launch (plus the root-cause/orchestration full playbooks on demand)". Anti-sycophancy named explicitly. Iron Law + rationalization table kept intact.</li> <li><strong><code>verify-first.js</code> (per-turn nudge):</strong> added three orchestration/anti-sycophancy one-liners to the rotation (bias-toward-delegation default-to-subagent; noisy commands via Haiku off-thread; capture every request/interruption in a priority-sorted list + parallel independent agents). The hash-mod rotation auto-scales to the new count; fail-open unchanged.</li> <li><strong>READMEs + AGENTS.md:</strong> note that root-cause + orchestration are enforced always-on via hooks while deadly-loop + feature-launch are conditional skills invoked on match; documented the bias-toward-delegation default, capture-every-request, and anti-sycophancy.</li> <li><strong>Version bump 0.3.5 -> 0.3.6.</strong></li> </ul> <h2 id="035">0.3.5<a class="headerlink" href="#035" title="Permanent link">¶</a></h2> <p>Production-doc finalization + doc-vs-code reconciliation: rewrite the READMEs, add <code>llms.txt</code>, and remove the never-registered PreCompact claims.</p> <ul> <li><strong>Production README rewrite:</strong> both the top-level <code>README.md</code> and the plugin <code>README.md</code> were rewritten for readability and structure — tagline, the four failure modes, quickstart, requirements (with the honest Node-on-PATH no-op caveat), a features table, plain-English how-it-works, the <code>/anti-hall:*</code> skills, statusline install, configuration/tuning, troubleshooting/FAQ, contributing (the 3 MODEL-POLICY copies must stay in sync), and license. The top README is the short overview; depth lives in the plugin README. Accurate to the real code (7 Node hooks, 4 skills, Node statusline).</li> <li><strong><code>llms.txt</code> added:</strong> an LLM-oriented index at the repo root (llms.txt standard) — H1 title, blockquote summary, and linked sections for the README, plugin README, CHANGELOG, AGENTS.md, KB doc, each skill SKILL.md, and each hook.</li> <li><strong>Doc-accuracy fixes (4 × P2):</strong></li> <li>Top README "What's inside" now states the FULL protocol injects at SessionStart (re-firing on compaction via <code>source=compact</code>) and a SHORT varying nudge per turn at UserPromptSubmit — not "the full protocol every turn".</li> <li><code>feature-launch/references/PRE-TOOL-USE-HOOK.md</code> no longer cites non-existent "Phase A.5.5" / "A.5.3"; it now points to "Phase A.5 (AFK readiness gate)" to match the prose (no sub-numbers) in <code>feature-launch/SKILL.md</code>.</li> <li><code>verify-first.js</code> comments + plugin README now state the per-turn nudge is hashed from the FULL stdin envelope (varies by session/cwd), not "reproducible for a given prompt". Runtime behavior unchanged.</li> <li><code>PRE-TOOL-USE-HOOK.md</code> example sentinel: added an explicit false-block/bypass tradeoff caveat and a pointer to the shipped quote-aware <code>git-guard.js</code> tokenizer; <code>--no-verify</code> / <code>--no-gpg-sign</code> patterns anchored to a flag boundary.</li> <li><strong>PreCompact placeholder claims removed (P1):</strong> <code>hooks.json</code> registers <code>verify-first-full.js</code> ONLY on <code>SessionStart</code> — there is no <code>PreCompact</code> registration block and never was in the shipped manifest. Every doc that described an "inert PreCompact placeholder registration" was inaccurate. Removed those claims from <code>verify-first-full.js</code> (banner + header + the dead <code>name === 'PreCompact'</code> echo branch), the plugin <code>README.md</code>, and <code>AGENTS.md</code>. Compaction survival is unchanged: it relies solely on the no-matcher <code>SessionStart</code> re-fire with <code>source="compact"</code>, which IS registered. The earlier 0.3.0/0.3.1 notes below describing a kept PreCompact placeholder are superseded by this entry.</li> <li><strong>AGENTS.md scope clarified (P2):</strong> <code>AGENTS.md</code> lives at the marketplace repo root, not under <code>plugins/anti-hall/</code>, so it is NOT bundled by <code>/plugin install</code> (the plugin <code>source</code> is <code>./plugins/anti-hall</code>). <code>plugin.json</code> description and the plugin <code>README.md</code> now state it is a repo-root Codex mirror for clone-based use that installed users must copy manually.</li> <li><strong>0.3.0 marketplace-source note corrected (P2):</strong> the 0.3.0 entry claimed the marketplace plugin <code>source</code> switched to a GitHub source object; the file actually uses the relative path <code>./plugins/anti-hall</code>. Corrected the 0.3.0 note to match the file (the relative path resolves because <code>marketplace add talas9/anti-hall</code> clones the whole repo).</li> </ul> <h2 id="034">0.3.4<a class="headerlink" href="#034" title="Permanent link">¶</a></h2> <p>Close the quoted-flag force-push bypass in git-guard.</p> <ul> <li><strong>git-guard: quoted force flags / refspecs / subcommand now BLOCK (P1):</strong> the force-push guard previously skipped any token that came entirely from inside quotes (<code>quotedOnly</code>). For argument-level flag/refspec/subcommand detection that was wrong — the POSIX shell strips quotes before git runs, so <code>git push "--force"</code>, <code>git push '--force'</code>, <code>git push "-f"</code>, <code>git push origin '+main'</code>, and <code>git "push" --force origin main</code> are byte-for-byte equivalent to their unquoted forms and DO rewrite published history, yet all five were reported ALLOW. <code>isForcePush()</code> now matches <code>--force</code>/<code>--force-with-lease</code>/bundled <code>-f</code>/<code>+refspec</code> regardless of quoting, and <code>gitSubcommand()</code> resolves a quoted subcommand token (e.g. <code>"push"</code>) instead of bailing to <code>sub=null</code> and leaving the command uninspected. Quoting still only changes meaning for commit-message CONTENT (a <code>--force</code>/<code>+main</code> inside an <code>-m</code> value), which is inspected separately on the <code>commit</code> path — never in a push arg list — so <code>git commit -m "fix --force bug"</code> is not false-blocked. Verified against the block/allow matrix (5 bypass cases now block; all prior blocks and legitimate pushes unchanged).</li> </ul> <h2 id="033">0.3.3<a class="headerlink" href="#033" title="Permanent link">¶</a></h2> <p>Portability finalization pass — restore the all-pure-Node guarantee.</p> <ul> <li><strong>Removed <code>node-preflight.sh</code> (P1):</strong> the POSIX-shell preflight added in 0.3.2 could not run on the platform it targeted. Claude Code executes a hook <code>command</code> via the system shell, which on a stock Windows box is cmd.exe/PowerShell — neither has <code>sh</code> on <code>PATH</code>. On the exact "fresh Windows box with no Node" case the preflight existed to warn about, <code>sh</code> is also absent, so <code>sh node-preflight.sh</code> failed to launch and the "anti-hall hooks are INACTIVE" warning never fired — the precise silent-off failure it was meant to prevent. The plugin's single non-Node component was also its least portable. Rather than ship a <code>.sh</code>/<code>.cmd</code>/<code>.ps1</code> matrix, the missing-Node case is now handled the same way as git-guard's other fail-open boundaries: documented as a hard, verify-before-relying prerequisite (README "Requirements" + install steps, "verify with <code>node --version</code>"). With the <code>.sh</code> gone, the manifest/README "all hooks pure Node, run unchanged on Windows/macOS/Linux given Node" claim is once again true.</li> <li><strong>Manifest/CHANGELOG vs README reconciled (P2):</strong> plugin.json's "All hooks and the statusline are pure Node" description and the 0.3.0 "no <code>.sh</code>, no bash" note no longer contradict the shipped hooks — there is no shell hook to contradict them. README, manifest, and CHANGELOG now agree.</li> </ul> <h2 id="032">0.3.2<a class="headerlink" href="#032" title="Permanent link">¶</a></h2> <p>Deadly-loop finalization pass (round 1 + round 2 findings; 0 P0, converged).</p> <ul> <li><strong>git-guard <code>--force-if-includes</code> false-block (P1):</strong> <code>--force-if-includes</code> / <code>--no-force-if-includes</code> is a safety modifier (a no-op on its own, only meaningful alongside <code>--force-with-lease</code>), not a force push. It is no longer a force trigger, so a bare <code>git push --force-if-includes origin main</code> is ALLOWED. <code>--force</code>, <code>-f</code>, <code>--force-with-lease</code>, and <code>+refspec</code> still BLOCK.</li> <li><strong>statusline installer cache-path discovery (P1):</strong> the one-command installer glob now covers the real install layout <code>~/.claude/plugins/cache/{marketplace}/{plugin}/{version}/</code> (verified against an actual install) plus the KB's shallower documented forms, keeping the existing marketplace globs and the <code>.claude-plugin/plugin.json</code> existence guard, so it no longer prints "not found" on a correctly installed plugin.</li> <li><strong>node-absent honest documentation (P1):</strong> every hook launches as <code>node <hook>.js</code>; if <code>node</code> is absent from the hook shell's <code>PATH</code>, all hooks (including the git-guard safety) silently do not run. A POSIX-shell <code>node-preflight.sh</code> SessionStart hook now emits a loud one-time warning when <code>node</code> is missing (a node-based preflight cannot detect its own missing interpreter), and the README install steps state the Node >= 18 requirement prominently and tell users to verify <code>node --version</code> before relying on the protections.</li> <li><strong>git-guard ANSI-C inline self-credit bypass (P2):</strong> a <code>git commit -m $'fix\n\nCo-authored-by: ...'</code> kept its <code>\n</code> literal after tokenizing, slipping past the line-anchored trailer regexes. Self-credit detection now also tests an escape-normalized copy of each inline message (<code>\n</code>/<code>\r</code>/<code>\t</code> interpreted), so the ANSI-C inline form is blocked too.</li> <li><strong>marketplace.json capability claim (P2):</strong> changed "PreCompact re-injection" to "SessionStart re-injection (fires with source=compact)" to match the code — PreCompact context injection is inert (forward-compat placeholder), SessionStart <code>source=compact</code> is the sole compaction-survival mechanism.</li> <li><strong>install-statusline.js ungrounded field (P2):</strong> dropped the undocumented <code>padding: 0</code> field; the installer writes only the doc-grounded <code>{ type: "command", command }</code>.</li> <li><strong>MODEL-POLICY triplication sync note (P2):</strong> the three byte-identical MODEL-POLICY.md copies (needed because skill bundling carries each skill's own <code>references/</code> copy and symlinks are stripped on install) now each carry a SYNC NOTE header, plus a README maintainer note, instructing that all three be updated together.</li> <li><strong>Stop-hook precedence (P2):</strong> <code>task-guard</code> is now registered before <code>graphify-reminder</code> on <code>Stop</code> so the higher-stakes open-task discipline reason wins when both fire (Claude Code does not merge Stop reasons); README caveat updated.</li> <li><strong>statusline blink removed (P2):</strong> the >=80% context tier in <code>statusline-monorepo.js</code> no longer uses SGR 5 (blink, inconsistently supported); it uses bold bright-red 256-color, consistent with the other tiers.</li> <li><strong>per-turn nudge wording (P2):</strong> README clarified that the verify-first nudge varies <strong>by prompt</strong> (deterministic SHA-1 of the prompt), not strictly per turn — an identical repeated prompt reproduces the same facet.</li> </ul> <h2 id="031">0.3.1<a class="headerlink" href="#031" title="Permanent link">¶</a></h2> <p>Deadly-loop finalization pass — applies the remaining open findings, all verified against the official Claude Code hooks docs.</p> <ul> <li><strong>graphify-reminder now actually reaches the model (P1):</strong> a Stop hook does NOT inject <code>additionalContext</code> (only UserPromptSubmit / UserPromptExpansion / SessionStart do, per the official docs), so the previous stdout-<code>additionalContext</code> emission was a silent no-op. The reminder is now surfaced via the only Stop channel that reaches the model — a <strong>one-time soft block</strong> (<code>{"decision":"block","reason":...}</code>), capped + deduped via <code>os.tmpdir</code> state so it nudges at most once per session and never loops. The "keep the graph updated" guidance also lives in the SessionStart primer, which IS a context-injection event.</li> <li><strong>git-guard emoji self-credit (P1):</strong> the <code>Generated with <AI></code> detector now catches the canonical footer even when prefixed by a leading glyph/emoji (e.g. the robot-emoji <code>Generated with [Claude Code]</code> footer) while still allowing prose ("we generated with care"). Verified against the full block/allow matrix.</li> <li><strong>Node prerequisite documented prominently (P1):</strong> added a <strong>Requirements</strong> section to the top-level README and the plugin description, and aligned the plugin README to Node >= 18. Hooks invoke a bare <code>node</code>; without a global Node on <code>PATH</code> every hook (including the git-guard safety) silently fails to launch.</li> <li><strong>git-guard scope caveat (P2):</strong> README now states the guard inspects only inline <code>-m</code>/<code>--message</code> trailers — <code>-F</code>/<code>--file</code>, editor commits, and interpreter wrappers (<code>sh -c</code>, <code>xargs</code>, aliases) are documented fail-open boundaries.</li> <li><strong>KB §1.4 corrected (P2):</strong> UserPromptSubmit context injection uses <strong>nested</strong> <code>hookSpecificOutput.additionalContext</code> (not flat); added an explicit note that context injection is event-gated to UserPromptSubmit / UserPromptExpansion / SessionStart, so <code>Stop</code>/<code>PreCompact</code> <code>additionalContext</code> is inert.</li> <li><strong>PreCompact framing corrected (P2):</strong> clarified that PreCompact never injects context (not "summarized away"); SessionStart <code>source="compact"</code> is the sole survive-compaction mechanism. (Superseded in 0.3.5: the PreCompact registration was never actually present in <code>hooks.json</code>, so all PreCompact-placeholder claims were removed rather than reworded.)</li> <li><strong>Two-Stop-hook coexistence noted (P2):</strong> <code>graphify-reminder</code> + <code>task-guard</code> both emit the top-level Stop <code>decision</code>/<code>reason</code> schema; Claude Code does not merge reasons, but each is capped so neither is lost (they sequence across Stops).</li> </ul> <h2 id="030">0.3.0<a class="headerlink" href="#030" title="Permanent link">¶</a></h2> <p>KB-driven effectiveness + portability revision (see <code>docs/PLUGIN-REVIEW.md</code>), plus a full OS-agnostic Node rewrite and the deadly-loop hardening pass.</p> <ul> <li><strong>OS-agnostic Node rewrite (portability):</strong> every hook and the statusline + its installer are now pure Node.js using only built-ins (<code>fs</code>, <code>path</code>, <code>os</code>, <code>crypto</code>, <code>child_process</code>) — no <code>.sh</code>, no bash/grep/sed/cksum/jq/python3, no <code>/dev/stdin</code>. The only spawned subprocess is <code>git</code> itself. They run unchanged on Windows, macOS, and Linux <strong>given Node.js on <code>PATH</code></strong> — Node is the one hard prerequisite (see README "Requirements"); without a global <code>node</code> the hooks cannot launch and the guards silently do not run. <code>hooks.json</code> invokes each as <code>node "${CLAUDE_PLUGIN_ROOT}/hooks/<name>.js"</code> (explicit <code>node</code>, not a shebang, since Windows ignores shebangs).</li> <li><strong>verify-first restructured (P0-1, P0-3):</strong> the FULL protocol moved to a SessionStart injection (<code>verify-first-full.js</code>), rewritten in the Superpowers "ONE Iron Law + rationalization/excuse table" form (names the specific bypass excuses: "probably", "should work", "seems to", "I'll just assume", "this looks done", "tests pass on first run"). The per-turn <code>UserPromptSubmit</code> (<code>verify-first.js</code>) is now a SHORT nudge that VARIES per turn (one of 5 one-liners, chosen deterministically by a SHA-1 hash (Node <code>crypto</code>) of the prompt) to fight habituation. JSON emitted via <code>JSON.stringify</code>, no jq.</li> <li><strong>Survive compaction (P0-2):</strong> the protocol persists across the compaction reset via the no-matcher <code>SessionStart</code> registration — Claude Code re-fires <code>SessionStart</code> after a compaction with <code>source="compact"</code>, and that injection is fresh post-reset context. (This note originally also described keeping a <code>PreCompact</code> registration as an inert placeholder; superseded in 0.3.5 — no PreCompact registration was ever present in <code>hooks.json</code>, and all such claims were removed. Per the official docs <code>additionalContext</code> is injected on exit 0 for UserPromptSubmit / UserPromptExpansion / SessionStart only, so a PreCompact hook would deliver nothing.) A duplicate matcher-<code>"compact"</code> SessionStart entry was removed so the protocol is not double-injected after a compaction.</li> <li><strong>git-guard force-push hardening (deadly-loop):</strong> force-push detection now resists prefix/wrapper/subshell bypasses and <code>+refspec</code> variants — env-prefix (<code>FOO=bar git push --force</code>), wrappers (<code>command</code>/<code>exec</code>/<code>sudo</code>/<code>env</code>/<code>time</code>/ <code>nohup</code>/<code>nice -n N</code>/<code>timeout 5</code>), subshell grouping (<code>(git push --force)</code>), global git options (<code>-c x=y push --force</code>), positional <code>+refspec</code> (with <code>--</code> end-of-options handling), and backslash-newline line continuations. Self-credit trailer detection is anchored to start-of-line trailer form so a prose mention ("docs: explain output generated with claude code") no longer false-blocks, while real <code>Co-Authored-By:</code> / <code>Generated with <AI></code> trailer lines still block. Verified against an 18-case block/allow matrix.</li> <li><strong>AGENTS.md mirror (P0-4):</strong> new repo-root <code>AGENTS.md</code> (<32 KiB) mirroring the verify-first Iron Law + commit/push hygiene + task discipline, so Codex subagents inherit the discipline (Codex <code>PreToolUse</code> cannot inject context).</li> <li><strong>Skill primer (P1-1):</strong> SessionStart now lists the plugin's skills (root-cause, deadly-loop, feature-launch, orchestration) + when to reach for each, folded into <code>verify-first-full.js</code>.</li> <li><strong>MODEL-POLICY (P1-2):</strong> documented <code>effort</code> default <code>high</code> / recommended max <code>xhigh</code> with fallback-to-<code>high</code> when unsupported (gpt-5.4-mini, Bedrock cmb); added the <code>codex</code> CLI-alias-in-subprocess caveat (detect in the executing shell, try an absolute path); added an anti-sycophancy clause (user agreement != correctness).</li> <li><strong>marketplace.json (P1-5, P1-6):</strong> plugin <code>source</code> is the relative path <code>"./plugins/anti-hall"</code>, which resolves because <code>/plugin marketplace add talas9/anti-hall</code> clones the whole repo (the relative path is taken from the marketplace root inside that clone); removed the per-plugin <code>version</code> duplication (version now lives only in <code>plugin.json</code>).</li> <li><strong>KB (<code>docs/KB-claude-codex.md</code>):</strong> added §9 "Anthropic Prompting 101" and §10 "Claude Opus 4.8 features relevant to this plugin", plus their source URLs.</li> <li><strong>CHANGELOG header:</strong> corrected "bump both manifests" to "bump <code>plugin.json</code> only (the authority)".</li> </ul> <h2 id="021">0.2.1<a class="headerlink" href="#021" title="Permanent link">¶</a></h2> <ul> <li>Fix <code>git-guard</code> self-credit regex: removed the bare <code>ai</code> alternation that false-blocked legitimate human trailers (e.g. <code>Co-authored-by: Ai ...</code>). Now matches specific AI/assistant signatures only.</li> <li>Add this CHANGELOG.</li> </ul> <h2 id="020">0.2.0<a class="headerlink" href="#020" title="Permanent link">¶</a></h2> <ul> <li>Add skills: <code>root-cause</code> (evidence-driven debugging), <code>orchestration</code> (non-blocking swarm; Claude+Codex load split; commands via Haiku), <code>feature-launch</code> (plan-first, deadly-loop-hardened, edge-case/scenario simulated), <code>deadly-loop</code> (iterative Reviewer+Critic debate), and shared <code>MODEL-POLICY.md</code> (Opus + Codex roster).</li> <li>Add graphify hooks: <code>graphify-session</code> (SessionStart, query-graph-first) and <code>graphify-reminder</code> (Stop, keep-graph-updated).</li> <li>Add <code>git-guard</code> (PreToolUse/Bash): block self-credit commit trailers and force push.</li> <li>Add conditional statusline (rich for monorepos, simple otherwise) + installer.</li> <li>Strengthen <code>verify-first</code> injection: no-jumping-to-conclusions, no-cause-no-fix, instrument-don't-guess, no-fake-completion, label claims.</li> </ul> <h2 id="010">0.1.0<a class="headerlink" href="#010" title="Permanent link">¶</a></h2> <ul> <li>Initial release: <code>verify-first</code> UserPromptSubmit hook + marketplace scaffold.</li> </ul> <h2 id="appendix-scriptsdevswarmjs-version-history">Appendix: <code>scripts/devswarm.js</code> version history<a class="headerlink" href="#appendix-scriptsdevswarmjs-version-history" title="Permanent link">¶</a></h2> <p>Moved here verbatim from <code>llms.txt</code> in 0.108.0 (llms.txt now carries only the current behaviour). Per-release detail is also in each version section above.</p> <ul> <li><a href="https://github.com/talas9/anti-hall/blob/main/plugins/anti-hall/scripts/devswarm.js">devswarm.js</a>: THE structured CLI (CLI over MCP, owner preference) — stable JSON on stdout, pure Node built-ins. Subcommands: register/ensure (write a workspace descriptor + populate sessionId), register-primary (register the CURRENT worktree's Primary under its per-worktree id <code>primary-<hash></code>, never a shared <code>'primary'</code>; as of v0.71.0 its <code>--session</code> defaults to <code>CLAUDE_CODE_SESSION_ID</code> instead of the workspace hash, so the registered row resolves its real transcript for liveness reads; as of v0.98.0 it also refuses, rather than plain-upserting, when a DIFFERENT currently-live session already holds this worktree's Primary row — <code>ok:false, reason:'live-primary-conflict'</code>, exit 2 — unless <code>--force</code> is passed; a same-session restart is never refused), heartbeat (turn-authored), inbox pull (child-side bounded native-queue drain into the durable inbox), inbox count/read/ack (the durable-inbox cursor primitive; ack advances the cursor = the parent-gate's non-skip clear path), inbox messages/read-primary (Primary/store non-destructive read of message bodies straight from the store — no descriptor needed, never touches the native queue; <code>--unread</code>/<code>--ack</code> against a separate <code>cursors/<id>.json</code> ACK cursor), workspaces list (derive+emit summary.json), gate --set/--clear (mark completion gates), nudge (poke-or-escalate), archive (archive-by-absence on anti-hall's OWN registry — hivecontrol has NO teardown command, so it SURFACES a manual "remove workspace in the DevSwarm app" step and never deletes), archive-request (REVISED v0.58 — see below; was v0.56.0 PARENT-side send-only via hivecontrol), archive-ignore/archive-unignore, migrate (<code>ANTIHALL_DEVSWARM_MIGRATE_MARK_READ=1</code> marks an imported legacy backlog as already-read, avoiding a false unread wall). <strong>v0.57 mesh (shipped in v0.58.0; Claude-side only):</strong> <code>send --to <meshId>|--broadcast --message TEXT [--urgency low|normal|high|urgent]</code> (daemon-independent, writes the shared per-project <code>store/<repoKey>/</code> directly; spoof-resistant <code>--from</code> via <code>callerIdentity</code>; fail-closed <code>--to</code> against the registry), <code>roster [--ack]</code> (project-scoped registry + <code>working_on</code> + <code>recent[]</code> projection), <code>mesh read</code> (alias of <code>roster --ack</code> — the only surface that clears <code>broadcastUnread</code>; as of v0.98.0, <code>--peek</code> reads without advancing the cursor and <code>--seq N</code> reads from an explicit historical seq instead of the caller's own cursor, always implying peek — neither flag changes the no-flag default behavior), and <code>heartbeat --summary TEXT</code> (also broadcasts a mesh status ping). Every mesh send/<code>inbox pull</code>/<code>archive-request</code> call runs a cooldown-bounded send-time self-heal of the per-project ingest daemon first. <strong>v0.58 "mesh-only messaging" (shipped in v0.58.0):</strong> <code>send --to-primary</code> (a third target mode, resolves the registry entry for THIS project's main worktree, fail-closed if unregistered); <code>roster</code> (plain, no <code>--ack</code>) now also folds in a read-only <code>hivecontrol workspace list children</code> view so an unregistered native child is still visible; <code>archive-request <childId> [--reason TEXT]</code> REVISED to a direct mesh-store write (zero <code>hivecontrol</code> calls, <code>--child-branch</code> removed — the old branch-resolution + <code>message-child</code> spawn is deleted); <code>reconcile</code> (one-shot per-worktree subprocess drain of every registered inbox in the project — never in-process, which would drain the wrong queue; auto-run since v0.58.1 by <code>doctor --fix</code> GATED and by <code>update</code> DevSwarm-session-only, manual verb still available); <code>spawn <branch> [...]</code> (thin pass-through wrap of <code>hivecontrol workspace create</code>, then best-effort auto-registers the worktree in the store); <code>merge [...]</code> (thin wrap of <code>check-merge</code> + <code>merge-into-source</code>, then broadcasts the outcome to the mesh). Full reference with source-line citations: <code>docs/KB-devswarm-hivecontrol.md</code> §8.8. <code>command-guard.js</code> carries a root-anchored LIGHT_EXCEPTION for it so the guard doesn't block its own wrapper, and (v0.58) a SEPARATE guard branch now blocks the native <code>message-child</code>/<code>message-parent</code> sends this CLI replaces (see <code>command-guard.js</code> above). <strong>v0.61.0 mesh self-heal:</strong> <code>send</code>/mesh delivery now resolves to the partition a child is ACTUALLY DRAINING (<code>pickSurvivor</code> — greatest registry <code>updatedAt</code> among live rows, cursor-value tiebreak), plus a phantom-only rescue on the child's first mechanical self-register that forwards any backlog stranded in a pre-registration phantom partition; new read-only <code>diagnose</code> (mesh-health detail: split/duplicate detection, orphans, stale partitions — never writes <code>summary.json</code>) and <code>healthcheck [--json]</code> (pass/fail over the same data, exit 0/2, for monitors/CI/the daemon; no <code>--json</code> prints one compact human line) verbs; register-time dedup (<code>retireWorktreeDuplicates</code>) now filters forwarded backlog through a new <code>isForwardable(msg)</code> noise filter — forwards only a real actionable direct, skipping broadcast/heartbeat and stale native poke/hash-mirror rows; <code>foldMeshDuplicates(home, ctx)</code> (MIGRATION) generalizes register-time dedup to the WHOLE registry, grouping by <code>canonicalMeshId</code> (git-toplevel-resolved, so a legacy subdir-split registration folds onto its toplevel) via the shared <code>groupRegistryByMeshId</code>, forward-then-tombstone, idempotent, non-destructive, fail-open — wired into both <code>update.js</code> (post-update, DevSwarm-session-gated) and <code>doctor</code>'s new AUTO-SAFE <code>fold-mesh-duplicates</code> repair (dry-run detect doubles as the read-only mesh-shape check under <code>--check</code>); <code>roster</code>/<code>workspaces list</code>/<code>diagnose</code> no longer call <code>deriveSummary</code>'s write path — pure reads, zero <code>summary.json</code> side-effect. <strong>v0.62.0:</strong> <code>unarchive <id></code> (reverses <code>archive</code> — restores an archived descriptor + store registry row, rejecting on an ownerKey/project mismatch or a conflicting live descriptor); <code>migrate-owner-keys</code> (forward-migration, idempotent/fail-open/no-delete — backfills a missing descriptor <code>ownerKey</code> and re-homes an ACTIVE descriptor stranded under a stale hash-keyed bucket into its fresh <code>repoKey</code>-keyed bucket; wired into <code>update.js</code>/<code>doctor</code>); <code>reap-stale [--yes|--confirm]</code> (project-scoped, dry-run by default — archives descriptors verdicted stale/escalated, gated by a fresh-heartbeat/recent-worktree-git-activity safety check); <code>reconcile-active [--active id,...] [--allow-empty] [--stdin] [--yes|--confirm]</code> (project-scoped, dry-run by default — archives every current workspace NOT in an explicit active set, refusing an empty set unless <code>--allow-empty</code>). <strong>v0.66.0:</strong> <code>heartbeat</code>'s reported <code>ok</code> and <code>reconcile</code>'s aggregate <code>ok</code> no longer mask a nested failure (a failed mesh broadcast, a crashed/timed-out drain target) — both now reflect what actually happened, treating a genuinely absent hivecontrol as a benign skip rather than a failure; <code>logs</code> now reads rotated log history, not only the live file. <strong>v0.67.0 human-readable workspace names:</strong> after a successful <code>hivecontrol workspace create</code>, <code>spawn</code> sets a title via a SEPARATE best-effort <code>hivecontrol workspace update-title -b <branch> "<title>"</code> call — derived from the <code>-p</code> brief (first non-empty line, one leading markdown marker stripped, whitespace collapsed, full line kept — no length cap since v0.108.0) unless the caller already passed <code>-t/--title</code>. <code>spawn</code>'s pass-through of the original argv to <code>hivecontrol workspace create</code> is untouched. <code>reconcile</code> caches whatever label hivecontrol already has for pre-existing workspaces but never invents one for a workspace with no brief on record. See <code>companion/lib/devswarm-names.js</code> below for the read-side cache. <strong>v0.67.1:</strong> <code>fetchNativeChildren</code>'s roster fold gained a cross-repo hijack guard — <code>hivecontrol workspace list children</code> resolves its scope entirely from <code>DEVSWARM_REPO_ID</code>/<code>DEVSWARM_BUILDER_ID</code>, never cwd, so a process holding a foreign repo's env got that repo's children back with exit 0 and valid JSON; each record's <code>repositoryId</code> is now cross-checked against a separate cwd-anchored, env-stripped <code>list all</code> ground truth (<code>fetchTrustedRepositoryId</code>), with mismatches dropped + logged and the fold failing open unfiltered when no ground truth is available. This is a NEW guard, not a repair — <code>fetchNativeChildren</code> passed <code>env</code> unmodified in every prior shipped release. Also new: a <code>skip <guard> [--ttl N]</code> CLI verb (<code>cmdSkip</code>, <code>devswarm.js:2637</code>), the CLI-side half of edit-guard's own skip-hint. <strong>v0.70.0 mesh/store hardening (mesh/store message-loss fix):</strong> <code>foldArchivedRegistryRows(home, ctx0)</code> (new) is a forward migration for registries split by the pre-fix archive bug — <code>cmdArchive</code> used to tombstone exactly one id per archive, so every registry that saw an archive under the old code could still hold live rows for worktrees whose workspace is now archived, and a live row is what made <code>computeSummary</code>/roster project that workspace as ACTIVE and, more seriously, is where a message could be forwarded even though that partition is dead. The sweep applies the SAME forward-then-tombstone + safety gate <code>cmdArchive</code> now applies at archive time (<code>retireArchivedWorktreeGroup</code>), retroactively, matching by canonical worktree real path (form-agnostic across id shapes) and sweeping every per-project store bucket (both <code>store/<repoKey></code> and the legacy <code>store/<8hex></code> bucket). Four load-bearing properties: IDEMPOTENT (a retired row is gone, a second run is a no-op), FAIL-OPEN HONESTLY (never throws into <code>update</code>/<code>doctor</code>, but a run that raised reports <code>ok:false</code> with the error, never a clean no-op), NO-DELETE (message rows are never deleted — unread directs are forwarded into the survivor's partition first, only registry rows are tombstoned), and SAFETY-GATED (<code>foldGroupIntoSurvivor</code> leaves any row with its own live descriptor alone). <code>ctx0.dryRun</code> classifies without writing (doctor's <code>--check</code>); the apply path runs each archived id's work under <code>withIdLock(id)</code>, surfacing (never silently dropping) a lock-busy id for retry next run. Wired into both <code>skills/update/scripts/update.js</code> (<code>fold-archived-rows</code> step) and <code>hooks/lib/doctor-repair.js</code>'s <code>migrationFix('fold-archived-rows', ...)</code>. Separately, <code>cmdArchive</code> gained <code>archivedTombstoneIsOrphaned(home, archivedStat)</code> — decides by INODE whether a leftover <code>archived/<id>.json</code> from a prior archive generation is orphaned (no live descriptor shares its <code>(dev, ino)</code>), fail-CLOSED on any incomplete scan; replaces the old unlink-then-link sequence with link-to-temp + atomic same-directory rename so <code>archivedPath</code> is never observably missing. <strong>v0.74.0:</strong> <code>gate --set merged</code> now also runs a best-effort git-ancestry check (<code>devswarm-git-truth.js</code>'s <code>gitMergedInto</code>, HEAD vs. the resolved default branch) and persists the verdict as a separate <code>merged_verified</code> gate row alongside <code>merged</code> — REPORT-ONLY, <code>merged</code> is set regardless of the verdict; a resolved-false verdict prints a stderr warning and shows as <code>merged (unverified)</code> on the parent roster, an unresolvable check omits <code>merged_verified</code> entirely. <strong>v0.75.0:</strong> <code>inbox peek-primary</code> (new) is the non-acking counterpart to <code>read-primary</code> — same read (message bodies straight from the store), <code>--ack:false</code> forced regardless of flags, so status can be checked without advancing the ACK cursor. <strong>v0.84.0 partition resolution follows the WORKSPACE, not the caller's cwd (closes defect <code>e586afdaa968</code>):</strong> <code>inbox read-primary</code>/<code>inbox count</code> resolved the store partition from the caller's working directory, so a Primary could be Stop-gated on mail it structurally could not see, and running the gate's own prescribed command from the wrong directory risked writing a cursor into another project's partition; resolution now comes from the workspace's REGISTERED project via the shared <code>registeredRepoKey</code> helper (<code>companion/lib/devswarm-repokey.js</code>; precedence fresh key → recorded <code>repoKey</code> → non-hash <code>ownerKey</code>), the same helper <code>devswarm-parent-gate.js</code> calls, so CLI and hook can no longer disagree about which workspaces a session owns. Same change, five more P0s: <code>gate</code>/<code>ensure</code>/<code>archive</code> re-homed a FOREIGN project's workspace — copying messages and registry rows and rewriting <code>ownerKey</code>, with <code>archive</code> additionally removing the live descriptor — BEFORE their own ownership guard ran, so even an invocation returning <code>ok:false</code> had already mutated another project's state (the guard now precedes every write; a refused call writes nothing); <code>inbox ack <foreign-id></code> advanced the NDJSON cursor after the resolver had already refused, permanently skipping that workspace's mail; and <code>inbox count</code>/<code>read</code> returned a silent <code>0</code> indistinguishable from "no mail", now <code>known:false</code> plus the NAMED <code>registeredRepoKey</code>/<code>callerRepoKey</code>. Full record: <code>docs/KB-devswarm-hivecontrol.md</code> §29. <strong>v0.85.0 archive retires the whole identity family:</strong> <code>cmdArchive</code> tombstoned by <code><id></code> ONLY, but an identity family can be cross-linked by <code>sessionId</code> (one row's <code>sessionId</code> IS the other row's <code>id</code>), so the twin stayed live in <code>workspaces/</code> forever and <code>devswarm-parent-gate.js</code> nagged un-clearably about its impossible inbox. <code>cmdArchive</code> now retires the whole family at archive time, and <code>foldArchivedFamilyDescriptors(home, ctx0)</code> (new, exported) is the FORWARD MIGRATION for sets already split by the pre-fix code — the DESCRIPTOR-file counterpart of <code>foldArchivedRegistryRows</code> (which covers only the registry half), wired into BOTH <code>update.js</code> and doctor's AUTO-SAFE <code>migrationFix('fold-archived-family-descriptors')</code>, scoped to genuinely-archived ids (<code>archived/<id>.json</code> present AND <code>workspaces/<id>.json</code> absent — a mid-archive/crashed state has BOTH and belongs to <code>applyRecoveryIntents</code>), IDEMPOTENT, NO-DELETE (bytes are hardlinked into <code>archived/</code> and the active path unlinked only after a fresh lstat proves the same inode; a tombstone already holding DIFFERENT bytes is never clobbered, and no message row is touched), SAFETY-GATED (grouping is <code>identityFamilyTwins</code>' cross-link ONLY, never bare worktree equality, so two legitimately-live tabs on one worktree are never retired), and FAIL-OPEN HONESTLY (<code>ok:false</code> with the error rather than a clean no-op); <code>ctx0.dryRun</code> classifies without writing and takes no lock. WRITE AUTHORITY MUST BE PROVEN: every descriptor writer publishes via <code>rename</code>, which allocates a NEW INODE at the same pathname, so a scan-time classification acted on later could unlink a freshly-registered LIVE descriptor. <code>descriptorFileGeneration(p)</code> (one coherent lstat+read — dev/ino/size/mtimeMs/bytes; deliberately NOT named <code>descriptorFingerprint</code>, which is already the recovery-intent sha256-of-an-object helper, since a duplicate declaration would silently repoint every recovery-intent call site) is re-read INSIDE the per-id lock and compared via <code>sameDescriptorGeneration</code> (FAIL CLOSED on either side absent), and <code>worktreeIsProvablyGone</code> gates the migration path (ABSOLUTE path + real ENOENT lstat only — never a relative path, dangling symlink, or unresolvable stat). Any generation mismatch or unproven gone-ness REFUSES the retire; refusals land in <code>left[]</code> with a reason and are surfaced by <code>update</code>'s summary and doctor's <code>notice</code> (rendered on BOTH the pending and not-pending path) rather than reading as "nothing to migrate". A one-way historical identity link is NOT by itself write authority. <code>worktreePath</code> is now persisted ABSOLUTE at build time; legacy relative values fail closed (treated as NOT gone) on every reader. <strong>v0.94.0 bounded reconcile + resume (defect f3c1bc827d89):</strong> <code>cmdReconcile</code> previously spawned one serial 30s-timeout child PER registry row with no total budget, so a project with dozens of stale rows made <code>update.js</code> (which awaits it synchronously) hang for minutes with zero progress output. A total wall-clock budget (<code>ANTIHALL_RECONCILE_BUDGET_MS</code>, default 60000ms; <code>--budget-ms</code> for a direct CLI call; <code>0</code> = unlimited) now bounds one sweep; a row whose <code>worktreePath</code> no longer exists is skipped BEFORE spawning (zero budget cost), and whatever is left when the budget runs out is deferred to <code>reconcile-resume.json</code> (per-repoKey, fail-open) and prioritized FIRST on the next sweep (FIFO rotation across runs). The result now also reports <code>budgetMs</code>, <code>processed</code>, <code>skippedMissingWorktree</code>, <code>deferred</code>, and <code>elapsedMs</code>. <strong>v0.95.0:</strong> <code>diagnose</code> resolves <code>sessionId</code> through the descriptor when the registry is stale and reports a <code>descriptorSessionId</code> field on disagreement; <code>unclaimed:</code> promotion derives the caller's real session id from <code>--session</code>, <code>CLAUDE_CODE_SESSION_ID</code>, or — only for a row still carrying the marker or lacking a sessionId — the harness's own session file found by walking the parent-pid chain with a cwd-in-worktree check and a pid-reuse/staleness liveness guard; descriptor/registry divergence is repaired in both directions, and a registry write failure during promotion now surfaces as <code>registryWriteError</code> instead of being swallowed. <strong>v0.96.0 (D11-A, defect f56dcc08f048):</strong> <code>resolveMeshTarget</code>/<code>pickSurvivor</code>'s target-selection gate is now <code>isRoutingLiveRowStrict</code> — a bare, no-descriptor-fallback <code>isSiblingPartitionLive</code> call (the SAME heartbeat-freshness + harness-session-dormancy predicate <code>companion/lib/devswarm-liveness-select.js</code>'s <code>pickFreshestLive</code> composes its ranking around) — replacing the old bare <code>isLiveSessionId</code> shape test, so a real-but-dormant sessionId no longer outranks a genuinely live sibling for a <code>send</code>/fold target; <code>groupRegistryByMeshId</code>'s REPORTING-only <code>liveRows</code> counter uses the fallback-inclusive <code>isRoutingLiveRow</code> (adds a descriptor-existence check for a just-registered row with no heartbeat yet) instead. <code>rehomeMiskeyedRow</code>/<code>retireWorktreeDuplicates</code>/<code>foldGroupIntoSurvivor</code> are DELIBERATELY untouched, still gating on bare <code>isLiveSessionId</code> (an identity-match question, not a drain/routing decision — see <code>docs/KB-devswarm-hivecontrol.md</code> §26). <code>callerOwnsRow</code>'s clause 3 ("sole row on the caller's own worktree") now additionally requires that sole row be unclaimed before granting ownership — a lone CLAIMED foreign row sharing the caller's worktree no longer lets the caller stamp its own sessionId over it. The sibling-watermark write and its read-then-conditional-unlink now serialize under one per-<code>(callerId,siblingId)</code> lock (<code>withWatermarkLock</code>), closing a TOCTOU window where a concurrent write landing between the read and the unlink was silently discarded (fail-open on contention: runs unlocked rather than dropping the op). <code>send</code>/<code>heartbeat</code> results and every ownership-refusal shape now carry an additive <code>identity:{id,kind}</code> (<code>resolved</code>/<code>declared</code>/<code>unresolvable</code>) alongside the existing bare identity string. <strong>v0.96.0 (D11-C):</strong> <code>diagnose</code> rows carry an additive <code>archivedInApp</code> field and force <code>live:false</code> whenever it is true, even against a fresh heartbeat; <code>reconcile</code> gained a fifth benign pre-spawn skip, <code>skippedNotGitRoot</code> (a worktree that exists on disk but fails <code>git rev-parse --show-toplevel</code>), and (D11-C2) the git-root probe itself is now bounded by <code>reconcile</code>'s own wall-clock budget — checked BEFORE the probe runs, not just before the resulting spawn, so N broken worktrees can no longer each burn a full probe timeout unaccounted-for before the first row defers; <code>inbox ack</code> now refuses the whole verb on a POSITIVE, resolvable ownership mismatch (the caller's own cwd/env resolves to a REAL, different registered row) instead of the previous half-ack, but an unresolvable-caller-identity or unregistered-caller shape still fails open exactly as before (<code>--ack-as-owner</code> still overrides either way). <strong>v0.98.0:</strong> new <code>wake-directive <id></code> verb — on-demand REPRINT of the full SessionStart idle-wake directive for <code><id></code> (placeholder substituted for the concrete id), what the trimmed Stop-gate <code>MAILBOX WAKE CHECK</code> re-verify (<code>wakeReassert</code>) now points an agent at instead of re-issuing <code>CronCreate</code> inline. Also: every read verb's <code>storeUnavailable</code> field is now a plain BOOLEAN (was, on <code>count</code>/<code>read</code>/<code>ack</code>, sometimes the full refusal detail object) with the underlying fs/sqlite error code surfaced separately as top-level <code>storeUnavailableReason</code> (string|null); <code>count</code>/<code>read</code>/<code>ack</code> additionally keep that fuller detail under a separate <code>storeUnavailableDetail</code> key so nothing is lost. The <code>read-primary</code>/<code>inbox messages --ack</code> ownership check and the ghost-id existence guard both now re-probe the store's read error AFTER their own <code>listRegistry()</code> call, so a genuinely unreadable registry reports <code>store-unavailable</code>/<code>storeUnavailableReason</code> instead of the misleading <code>caller-not-registered</code>/<code>unregistered-workspace</code>.</li> </ul> </article> </div> <script>var tabs=__md_get("__tabs");if(Array.isArray(tabs))e:for(var set of document.querySelectorAll(".tabbed-set")){var labels=set.querySelector(".tabbed-labels");for(var tab of tabs)for(var label of labels.getElementsByTagName("label"))if(label.innerText.trim()===tab){var input=document.getElementById(label.htmlFor);input.checked=!0;continue e}}</script> <script>var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))</script> </div> <button type="button" class="md-top md-icon" data-md-component="top" hidden> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"><path d="M13 20h-2V8l-5.5 5.5-1.42-1.42L12 4.16l7.92 7.92-1.42 1.42L13 8z"/></svg> Back to top </button> </main> <footer class="md-footer"> <nav class="md-footer__inner md-grid" aria-label="Footer" > <a href="../REPO-PIPELINES/" class="md-footer__link md-footer__link--prev" aria-label="Previous: Repository pipelines"> <div class="md-footer__button md-icon"> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"><path d="M20 11v2H8l5.5 5.5-1.42 1.42L4.16 12l7.92-7.92L13.5 5.5 8 11z"/></svg> </div> <div class="md-footer__title"> <span class="md-footer__direction"> Previous </span> <div class="md-ellipsis"> Repository pipelines </div> </div> </a> </nav> <div class="md-footer-meta md-typeset"> <div class="md-footer-meta__inner md-grid"> <div class="md-copyright"> <div class="md-copyright__highlight"> MIT © Mohammed Talas </div> Made with <a href="https://squidfunk.github.io/mkdocs-material/" target="_blank" rel="noopener"> Material for MkDocs </a> </div> </div> </div> </footer> </div> <div class="md-dialog" data-md-component="dialog"> <div class="md-dialog__inner md-typeset"></div> </div> <div class="md-progress" data-md-component="progress" role="progressbar"></div> <script id="__config" type="application/json">{"annotate": null, "base": "..", "features": ["navigation.instant", "navigation.instant.progress", "navigation.tracking", "navigation.tabs", "navigation.top", "navigation.footer", "navigation.indexes", "toc.follow", "search.suggest", "search.highlight", "search.share", "content.code.copy", "content.action.edit", "content.tabs.link"], "search": "../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}</script> <script src="../assets/javascripts/bundle.d7400e89.min.js"></script> </body> </html>