Skip to content

Claude Code harness feature surface vs anti-hall usage (audited 2026-08-01)

Research note

This page is background research kept for the project's own reference. It is not user documentation, it is not kept up to date with every release, and model or product details in it may be out of date. For how anti-hall works today, start at the home page.

1. Purpose

A living map of what the Claude Code CLI/harness offers a plugin, what anti-hall currently uses, and the gaps — so feature-adoption decisions are evidence-based instead of vibes-based. Companion doc: docs/archive/superpowers/specs/2026-08-01-harness-feature-adoption.md turns the gaps below into a phased adoption plan.

Provenance: every feature claim below was verified against the official code.claude.com/docs tree on 2026-08-01. Docs move — re-verify before acting, especially on exact hook payload contracts. Provenance (2026-08-21): the cross-session/agent-to-agent messaging, agent teams, and §4 DevSwarm-implication material below was verified against the same docs tree on 2026-08-21 (local Claude Code version 2.1.238) and is additive to the 2026-08-01 audit — none of the earlier claims were changed.


2. anti-hall CURRENT usage

Feature Used? Evidence file Notes
SessionStart hook USED plugins/anti-hall/hooks/hooks.json 7 handlers: verify-first-full.js, verify-first-orch.js, devswarm-child-role.js, version-alert.js, fable-availability.js, codex-availability.js, progress-prune.js. Also covers source=compact re-injection (session resume after compaction).
UserPromptSubmit hook USED same 5 handlers: verify-first.js, task-tracker.js, limit-conserve-inject.js, devswarm-parent-inbox.js, devswarm-child-turn.js.
PreToolUse hook USED same Bash: git-guard, command-guard, merge-gate, scan-throttle. Write/Edit/MultiEdit: api-guard, ship-it-guard. Write/Edit/MultiEdit/NotebookEdit: edit-guard. Read: inbox-read-guard. Agent+Task: model-routing-guard, swarm-guard, phase-tracker. Agent
SubagentStart hook USED hooks/verify-first-subagent.js Claude plugin only — not present in the Codex port (Codex has no equivalent lifecycle hook for sub-sessions today).
PostToolUse hook USED hooks/hooks.json:370,380,390 3 Bash-matcher handlers: output-verify-guard.js (:370), devswarm-parent-reply-tracker.js (:380), devswarm-child-drain.js (:390). [UPDATE 2026-08-21: this row was missing — the same fact was already documented at KB.md:71 as a v0.69.0 addition; this file was never updated to match. See §2's "NOT USED" row below, corrected in the same pass.]
PostToolUseFailure hook USED hooks/hooks.json:402 1 Bash-matcher handler: failure-root-cause-nudge.js. [UPDATE 2026-08-21: see note above — also shipped since v0.69.0, previously mis-listed as NOT USED here.]
Stop hook USED same task-guard, tasklist-guard, speculation-guard, speculation-judge, claim-ledger, codex-nudge, devswarm-parent-gate, devswarm-child-gate.
Monitor tool USED (opt-in) plugins/anti-hall/monitors/monitors.json → companion/lib/devswarm-wake-watch.js Watches a DevSwarm workspace's own mailbox count delta, edge-triggered wake. Gated to DevSwarm sessions (DEVSWARM_REPO_ID); dormant otherwise.
run_in_background USED orchestration guidance + the wake-watch process Standard background-Bash pattern for long ops; also underlies the wake-watch companion.
CronCreate / CronList USED (opt-in) hooks/lib/devswarm-wake.js DevSwarm self-wake fallback: an injected agent directive tells the agent to CronCreate its own wake job (Cron is a Claude tool, not a hook the plugin can register directly).
Skills USED plugins/anti-hall/skills/ (17, + MODEL-POLICY.md, not itself a skill) + plugins/anti-hall/codex/skills/ (20) [UPDATE 2026-09-24, v0.108.0] Claude: activate, deadly-loop, deadly-loop-multi, debt, defects, devswarm, doctor, handover, install-statusline, jev, orchestration, root-cause, settings, ship-it, simplify, system-briefing, update. Codex: the same set as anti-hall-* plus context-conserve, omc, omx (Codex-specific bridges), model-policy.
Statusline USED plugins/anti-hall/statusline/ Rich / simple / monorepo renderers + phase bar; installed via the install-statusline skill.
Subagent guards (Agent/Task matchers) USED hooks.json PreToolUse model-routing-guard, swarm-guard, phase-tracker fire on both Agent and Task tool calls.
Workflow tool USED plugins/anti-hall/skills/{deadly-loop,ship-it}/references/*.workflow.js Delivered as user-saved workflow templates — a plugin cannot ship a workflow directly as an installable command, so these ship as reference files a skill instructs the user/agent to save.
SendMessage GUIDANCE-ONLY hooks/verify-first-subagent.js:56, skills/orchestration/SKILL.md:111 [APPENDED 2026-08-21] Mentioned only in prompt guidance text — no actual tool-call site in the codebase. ListAgents: 0 code hits repo-wide (docs-only). See the owner-constraint note appended below §2 for why cross-machine SendMessage/ListAgents addressing should NOT be recommended for this owner.
Plugin marketplace USED .claude-plugin/marketplace.json, plugins/anti-hall/.claude-plugin/plugin.json (v0.68.2 [UPDATE 2026-08-21: stale — committed version is now 0.75.1; the working tree carries 0.76.0 with a release in flight (uncommitted, per release-prep convention — do not assert 0.76.0 has shipped)]) Codex mirror at plugins/anti-hall/codex/.codex-plugin/.
NOT USED — — PostToolBatch, SubagentStop, PreCompact, PostCompact, SessionEnd, Setup, Notification, TaskCreated/TaskCompleted hooks, ConfigChange, PermissionRequest/PermissionDenied, MessageDisplay, WorktreeCreate/WorktreeRemove, ScheduleWakeup (referenced only in a code comment, not invoked), CronDelete, LSP servers (.lsp.json), output styles, Agent SDK / headless mode, MCP server ship-or-consume (deliberate CLI-over-MCP posture), sandboxing, checkpointing awareness. [UPDATE 2026-08-21: PostToolUse and PostToolUseFailure REMOVED from this row — both are now USED (see the two new rows above); this row previously listed them as not-used, which was stale since v0.69.0.] SendMessage = GUIDANCE-ONLY (see §2 usage-table appendix below); ListAgents = NOT USED.

[APPENDED 2026-08-21] Owner constraint — do not recommend cross-machine SendMessage/ ListAgents addressing for this owner. The owner rotates Claude accounts whenever a usage limit is exhausted. Remote Control is account-scoped, so RC/cloud-session addressing is not a usable transport across a rotation: a session addressed under one account disappears from ListAgents/SendMessage reach once the owner rotates to another. The file-based DevSwarm mesh (~/.anti-hall/devswarm/store/) is account-agnostic and survives rotation — it has no equivalent blind spot. Consequence for any future recommendation in this doc or its adoption plan: score Claude-native messaging on its LOCAL/same-machine capability only; do not propose it as a cross-machine coordination layer for this owner. UNVERIFIED: whether same-machine, same-account local-session addressing (the "Other LOCAL Claude Code sessions" row in §3's cross-session-messaging table) is itself account-scoped was not tested this session — flagged, not asserted either way.

[APPENDED 2026-08-21] This repo's own tools/eval/ harness (distinct from evals/). tools/eval/ (singular) exists at the repo root: run.js (433 lines) + 4 pilot graders (ship-it, scope-fidelity, rule-behavior, false-done, 986 lines total) + README.md (296 lines). evals/ (plural — the claude plugin eval CLI convention scaffolded by eval init) does not exist in this repo; the two names are easy to conflate and refer to different, non-interoperating things.

Capability boundary — honest, not equivalent to claude plugin eval: the tools/eval/ pilots score what the agent STATES it would do, not what it actually did — see tools/eval/…/ship-it-pilot.mjs:7-11, which documents this directly. claude plugin eval run with --allow-tools plus a tool_used: Skill grader is categorically different: it proves the skill actually fired as a real tool call in a sandboxed execution, not merely that the model claimed it would. Neither harness subsumes the other — the in-repo pilots are cheap self-report checks; claude plugin eval is the mechanically verified alternative, with actual case execution still UNVERIFIED this session (see the claude plugin eval entry in §3's Additions table above).

[APPENDED 2026-08-21] CI gap. package.json:7's "test" script is node --test only — CI runs no tools/eval/ pilot and no claude plugin eval suite. Neither harness's result gates a merge or release today.


3. Full Claude Code harness feature surface (2026-08)

Hook events

Source: /docs/en/hooks.

Event One-line description
SessionStart Fires when a session starts or resumes (incl. after /compact, source=compact).
Setup Fires on first-time plugin/project setup.
SessionEnd Fires when a session terminates.
UserPromptSubmit Fires when the user submits a prompt, before the model sees it.
UserPromptExpansion Fires when a prompt is expanded (e.g. slash-command/skill substitution).
PreToolUse Fires before a tool call executes; can block/modify.
PostToolUse Fires after a tool call completes successfully.
PostToolUseFailure Fires after a tool call fails.
PostToolBatch Fires after a batch of parallel tool calls completes.
PermissionRequest Fires when a permission prompt would be shown.
PermissionDenied Fires when a permission request is denied.
Stop Fires when the agent is about to stop responding this turn.
SubagentStart Fires when a spawned subagent session starts.
SubagentStop Fires when a spawned subagent session stops.
StopFailure Fires when a stop/completion attempt itself fails.
TeammateIdle Fires when a teammate/agent in a multi-agent session goes idle.
PreCompact Fires before context compaction runs.
PostCompact Fires after context compaction completes.
TaskCreated Fires when a task is added to the task list.
TaskCompleted Fires when a task is marked complete.
InstructionsLoaded Fires when CLAUDE.md/AGENTS.md/instruction files are loaded.
ConfigChange Fires when settings/config files change.
CwdChanged Fires when the working directory changes.
FileChanged Fires when a watched file changes on disk.
Notification Fires on harness notifications (e.g. permission-needed, idle).
MessageDisplay Fires when a message is rendered to the user.
WorktreeCreate / WorktreeRemove Fire on git worktree lifecycle events.
Elicitation Fires when the harness elicits structured input (e.g. from MCP).

Long-running / background

Source: /docs/en/tools-reference.

Feature One-line description
Monitor Registers a background watcher process that can emit wake events.
Background Bash Runs a shell command detached, polled/notified on completion.
Task / Agent tools Spawn subagents/sub-sessions for delegated work.
SendMessage Sends a message to another agent/teammate session.
PushNotification Pushes a notification to the user outside the transcript.
SendUserFile Delivers a file artifact to the user.

Scheduling

Source: /docs/en/scheduled-tasks, /docs/en/routines.

Feature One-line description
/loop Re-runs a prompt/command on a recurring interval within a session.
CronCreate / CronList / CronDelete Create/list/delete cron-style scheduled jobs that fire independent of an open REPL.
ScheduleWakeup Schedules a one-time future wake for the current session.
Routines Cloud-scheduled recurring agents (cron-driven, run headless).
RemoteTrigger Triggers a remote/cloud agent run externally.

Plugin components

Sources: /docs/en/plugins, /docs/en/plugins-reference, /docs/en/headless, /docs/en/settings, /docs/en/permissions, /docs/en/memory.

Feature One-line description
Plugin system (bin/, output-styles/, workflows/, monitors/, .mcp.json, .lsp.json) Declarative manifest surface a plugin ships components through.
Skills Reusable, invokable instruction packages (SKILL.md + assets).
Slash commands User-typed shortcuts that expand to a prompt/skill.
Subagents Named agent personas with scoped tools/model.
Output styles Alternate system-prompt presentation modes.
Statusline Persistent status bar rendered above the input.
MCP Model Context Protocol server integration (tools/resources).
LSP Language Server Protocol integration for diagnostics/navigation.
Workflow tool Programmatic multi-agent orchestration primitive.
Agent SDK / headless mode Programmatic, non-interactive Claude Code execution.
Settings permissions Allow/deny/ask rules for tools and commands.
Memory CLAUDE.md/project-memory persistence layer.

Newer 2026 features

Feature One-line description Doc
Checkpointing / rewind Save and roll back to a prior conversation/file state. /docs/en/checkpointing
Sandboxing Contain tool execution (filesystem/network) inside a restricted boundary. /docs/en/sandboxing
Permission modes Named presets (default/acceptEdits/plan/bypass/etc.) governing tool approval. /docs/en/permission-modes
Channels Structured multi-party communication surface. /docs/en/channels
Sessions: resume / branch / fork Session lifecycle operations beyond linear continuation. —
Remote Control Externally drive/observe a running session. —
Worktrees Git-worktree-scoped isolated session workspaces. —

Flagged / unverified

  • The /docs/en/settings summary listed hook names beforeBash/afterBash/beforeWrite/afterWrite/configChange that do not appear on /docs/en/hooks. Treat this as a summarization artifact — use the /docs/en/hooks event names (PreToolUse/PostToolUse/ConfigChange, etc.), not the settings-summary names.
  • agent-teams, agent-view, desktop-scheduled-tasks, and the Remote Control page were referenced during this audit but not fetched — their contract is unverified. [UPDATE 2026-08-21: not re-checked this session — mark UNVERIFIED-since-2026-08-01, left as-is.]
  • /docs/en/llms.txt returned 404 at audit time. [UPDATE 2026-08-21: not re-checked this session — mark UNVERIFIED-since-2026-08-01, left as-is.]

Cross-session / agent-to-agent messaging

Local Claude Code version observed: 2.1.238. Tools: SendMessage, ListAgents.

target class bidirectional durable wakes idle addressable when not running
In-process subagents (spawned this session) yes NO N/A (not idle-capable) session-scoped only
Agent-team teammates yes NO starts a turn when idle only while session active
Other LOCAL Claude Code sessions (same machine) yes NO starts a new turn if idle NO — message dropped if target not running
Cloud/web sessions one-way if sender not connected to Remote Control (receives, cannot reply) NO appears in conversation only while running
Remote Control sessions (other machines) yes, if sender is on Remote Control NO appears in conversation only while running

Documented behavior: "The receiving Claude reads the message between tool calls during an active turn. When the receiving session is idle, Claude Code starts a new turn with the message."

DURABILITY VERDICT: SendMessage is live-session IPC, NOT a durable message bus. Messages to a non-running session are dropped silently with no error to the sender. No ack, no read receipt, no replay, no queue-for-later. Does not survive the target's exit/crash, /clear, /compact, or a machine restart.

Nuance — do not overstate: durability for offline targets is UNDOCUMENTED, not documented-as-absent. Agent teams do persist a per-agent inbox JSON at ~/.claude/teams/{team}/inboxes/{agent}.json, so some on-disk state exists, but no offline-delivery mechanism is described anywhere in the docs. Mark as UNVERIFIED — needs an empirical test (see §4 below).

Source: /docs/en/cross-session-messaging, /docs/en/agent-teams

Agent teams (EXPERIMENTAL)

Opt-in via CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 (v2.1.178+). Multiple coordinated sessions, a shared task list, and direct messaging; split panes or in-process. Documented limitations: no session resumption with in-process teammates, task status can lag, no nested teams. Architecture: per-agent inbox JSON under ~/.claude/teams/{team}/inboxes/.

Flagged EXPERIMENTAL — unsuitable as the foundation for a shipped substrate today.

Source: /docs/en/agent-teams

Additions to the feature surface (absent from every prior doc in this repo — verified by grep)

[UPDATE 2026-08-21: the header claim above ("absent from every prior doc in this repo") is FALSE for two of the rows below — corrected per-row rather than rewritten, per this repo's own "never silently rewrite" rule (KB.md:37). The "absent" framing holds ONLY for the ListAgents row; the other rows below were not all independently re-checked, but isolation:"worktree" is confirmed already documented elsewhere in this repo (see its row).]

Feature Status One-line description
ListAgents Stable Enumerates addressable agents: in-process subagents, other local sessions, cloud sessions, and (when Remote Control is connected) the account's other sessions. Names are the address for SendMessage. Confirmed genuinely absent elsewhere in this repo: 0 code hits repo-wide (docs-only usage).
EnterWorktree / ExitWorktree Stable Worktree entry/exit tools.
Artifacts — capabilities Gated per-account Runtime capabilities for published interactive pages beyond static HTML.
claude plugin eval [UPDATE 2026-08-21: was "Org-flag gated" — corrected below, see new row] superseded, see below
Agent tool subagent_type: "fork" Stable (v2.1.212+) Spawns a fork that inherits the parent's full context.
Agent tool isolation: "worktree" Stable (v2.1.203+) Runs the spawned agent in a temporary git worktree. [UPDATE 2026-08-21: NOT absent from prior docs — already documented at skills/orchestration/SKILL.md:27 AND docs/KB-claude-workflow-orchestration.md:58,104 before this row was added. The "absent from every prior doc" framing in this section's header does not hold for this row.]
Monitor Stable, with a gap Unavailable on Bedrock/Vertex/Foundry.
CronCreate / CronList Stable Scheduled tasks within a session; restored on --resume.

[UPDATE 2026-08-21] claude plugin eval: corrected from "Org-flag gated" — the CLI is AVAILABLE LOCALLY (verified this session): CLI version 2.1.238, claude plugin eval --help renders successfully. Flags observed: --ablation with-without, --allow-tools, --threshold, --json, --report, --no-publish, --max-cost-usd, --judge-model, --case, --tag; eval init [--bare] scaffolds a suite. Suite convention: an <eval-dir>/**/case.yaml OR a prompt.md + graders/*.md pair; results land at <eval-dir>/results/<timestamp>/aggregate-result.json. Graders: regex, LLM judge, tool_used, file_exists, baseline. --json v1 shape (schemaVersion, cases[].arms.{with,without}[].graders, aggregates), --report HTML, --no-publish local. Exit codes: 0 pass (default threshold 1.0), 1 fail/error/empty, 2 partial (cost-ceiling / auth-fail). Still UNVERIFIED: whether a case actually EXECUTES — no run was attempted this session; only --help and flag enumeration were checked. Do not claim a passing/failing run without evidence.

Source: tools reference, sub-agents, hooks, plugin eval, docs index


4. Implication for anti-hall's DevSwarm mesh

Question asked: could Claude Code's native agent-to-agent messaging (SendMessage / ListAgents / agent teams) REPLACE anti-hall's DevSwarm mesh store (~/.anti-hall/devswarm/store/<hash>/devswarm.db, introduced v0.58.0)?

Answer: NO — they solve different problems. The mesh is a DURABLE system of record that survives session death and supports replay/audit; SendMessage addresses only LIVE agents and drops silently otherwise.

property DevSwarm mesh SendMessage
Survives target death ✅ ❌
Sender learns of non-delivery ✅ ❌
Wakes an idle session via cron + Monitor watcher ✅ starts a new turn — genuinely better
Cross-machine ❌ local SQLite ✅ via Remote Control
Replay + audit ✅ ❌

Recommended posture: LAYER, don't replace. Keep the mesh as the durable store of record; Claude's native messaging is a candidate LOW-LATENCY WAKE + live-coordination layer over it — conceptually what plugins/anti-hall/companion/lib/devswarm-wake-watch.js already does with a Monitor. The one genuinely NEW capability it unlocks is cross-machine reach, which a local SQLite mesh structurally cannot provide.

Open / unverified: an empirical test of what actually happens to a message sent to a dead session — is the inbox JSON written and later drained, or is it truly dropped? Until tested, treat durability as absent.

Assessed 2026-08-21 against Claude Code 2.1.238.


Audited 2026-08-01. Re-verify hook contracts against current docs before building against them — see the adoption plan's Fable-review step. The adoption plan was Fable-reviewed 2026-08-02, which corrected several hook contracts assumed above: PostCompact cannot inject additionalContext (side-effect-only), SubagentStop injects into the subagent's own turn rather than the parent, and ConfigChange does not watch ~/.anti-hall/skip.json.

Extended 2026-08-21 (Claude Code 2.1.238): added the cross-session/agent-to-agent messaging and agent-teams subsections plus §4's DevSwarm-mesh-vs-SendMessage comparison. The empirical test of message delivery to a dead session (inbox JSON written-and-drained vs. truly dropped) remains open — see §4.