Claude Code harness feature surface vs anti-hall usage (audited 2026-08-01)¶
Research note
This page is background research kept for the project's own reference. It is not user documentation, it is not kept up to date with every release, and model or product details in it may be out of date. For how anti-hall works today, start at the home page.
1. Purpose¶
A living map of what the Claude Code CLI/harness offers a plugin, what anti-hall
currently uses, and the gaps — so feature-adoption decisions are evidence-based
instead of vibes-based. Companion doc:
docs/archive/superpowers/specs/2026-08-01-harness-feature-adoption.md
turns the gaps below into a phased adoption plan.
Provenance: every feature claim below was verified against the official code.claude.com/docs tree on 2026-08-01. Docs move — re-verify before acting, especially on exact hook payload contracts. Provenance (2026-08-21): the cross-session/agent-to-agent messaging, agent teams, and §4 DevSwarm-implication material below was verified against the same docs tree on 2026-08-21 (local Claude Code version 2.1.238) and is additive to the 2026-08-01 audit — none of the earlier claims were changed.
2. anti-hall CURRENT usage¶
| Feature | Used? | Evidence file | Notes |
|---|---|---|---|
SessionStart hook |
USED | plugins/anti-hall/hooks/hooks.json |
7 handlers: verify-first-full.js, verify-first-orch.js, devswarm-child-role.js, version-alert.js, fable-availability.js, codex-availability.js, progress-prune.js. Also covers source=compact re-injection (session resume after compaction). |
UserPromptSubmit hook |
USED | same | 5 handlers: verify-first.js, task-tracker.js, limit-conserve-inject.js, devswarm-parent-inbox.js, devswarm-child-turn.js. |
PreToolUse hook |
USED | same | Bash: git-guard, command-guard, merge-gate, scan-throttle. Write/Edit/MultiEdit: api-guard, ship-it-guard. Write/Edit/MultiEdit/NotebookEdit: edit-guard. Read: inbox-read-guard. Agent+Task: model-routing-guard, swarm-guard, phase-tracker. Agent |
SubagentStart hook |
USED | hooks/verify-first-subagent.js |
Claude plugin only — not present in the Codex port (Codex has no equivalent lifecycle hook for sub-sessions today). |
PostToolUse hook |
USED | hooks/hooks.json:370,380,390 |
3 Bash-matcher handlers: output-verify-guard.js (:370), devswarm-parent-reply-tracker.js (:380), devswarm-child-drain.js (:390). [UPDATE 2026-08-21: this row was missing — the same fact was already documented at KB.md:71 as a v0.69.0 addition; this file was never updated to match. See §2's "NOT USED" row below, corrected in the same pass.] |
PostToolUseFailure hook |
USED | hooks/hooks.json:402 |
1 Bash-matcher handler: failure-root-cause-nudge.js. [UPDATE 2026-08-21: see note above — also shipped since v0.69.0, previously mis-listed as NOT USED here.] |
Stop hook |
USED | same | task-guard, tasklist-guard, speculation-guard, speculation-judge, claim-ledger, codex-nudge, devswarm-parent-gate, devswarm-child-gate. |
Monitor tool |
USED (opt-in) | plugins/anti-hall/monitors/monitors.json → companion/lib/devswarm-wake-watch.js |
Watches a DevSwarm workspace's own mailbox count delta, edge-triggered wake. Gated to DevSwarm sessions (DEVSWARM_REPO_ID); dormant otherwise. |
run_in_background |
USED | orchestration guidance + the wake-watch process | Standard background-Bash pattern for long ops; also underlies the wake-watch companion. |
CronCreate / CronList |
USED (opt-in) | hooks/lib/devswarm-wake.js |
DevSwarm self-wake fallback: an injected agent directive tells the agent to CronCreate its own wake job (Cron is a Claude tool, not a hook the plugin can register directly). |
| Skills | USED | plugins/anti-hall/skills/ (17, + MODEL-POLICY.md, not itself a skill) + plugins/anti-hall/codex/skills/ (20) [UPDATE 2026-09-24, v0.108.0] |
Claude: activate, deadly-loop, deadly-loop-multi, debt, defects, devswarm, doctor, handover, install-statusline, jev, orchestration, root-cause, settings, ship-it, simplify, system-briefing, update. Codex: the same set as anti-hall-* plus context-conserve, omc, omx (Codex-specific bridges), model-policy. |
| Statusline | USED | plugins/anti-hall/statusline/ |
Rich / simple / monorepo renderers + phase bar; installed via the install-statusline skill. |
Subagent guards (Agent/Task matchers) |
USED | hooks.json PreToolUse |
model-routing-guard, swarm-guard, phase-tracker fire on both Agent and Task tool calls. |
Workflow tool |
USED | plugins/anti-hall/skills/{deadly-loop,ship-it}/references/*.workflow.js |
Delivered as user-saved workflow templates — a plugin cannot ship a workflow directly as an installable command, so these ship as reference files a skill instructs the user/agent to save. |
SendMessage |
GUIDANCE-ONLY | hooks/verify-first-subagent.js:56, skills/orchestration/SKILL.md:111 |
[APPENDED 2026-08-21] Mentioned only in prompt guidance text — no actual tool-call site in the codebase. ListAgents: 0 code hits repo-wide (docs-only). See the owner-constraint note appended below §2 for why cross-machine SendMessage/ListAgents addressing should NOT be recommended for this owner. |
| Plugin marketplace | USED | .claude-plugin/marketplace.json, plugins/anti-hall/.claude-plugin/plugin.json (v0.68.2 [UPDATE 2026-08-21: stale — committed version is now 0.75.1; the working tree carries 0.76.0 with a release in flight (uncommitted, per release-prep convention — do not assert 0.76.0 has shipped)]) |
Codex mirror at plugins/anti-hall/codex/.codex-plugin/. |
| NOT USED | — | — | PostToolBatch, SubagentStop, PreCompact, PostCompact, SessionEnd, Setup, Notification, TaskCreated/TaskCompleted hooks, ConfigChange, PermissionRequest/PermissionDenied, MessageDisplay, WorktreeCreate/WorktreeRemove, ScheduleWakeup (referenced only in a code comment, not invoked), CronDelete, LSP servers (.lsp.json), output styles, Agent SDK / headless mode, MCP server ship-or-consume (deliberate CLI-over-MCP posture), sandboxing, checkpointing awareness. [UPDATE 2026-08-21: PostToolUse and PostToolUseFailure REMOVED from this row — both are now USED (see the two new rows above); this row previously listed them as not-used, which was stale since v0.69.0.] SendMessage = GUIDANCE-ONLY (see §2 usage-table appendix below); ListAgents = NOT USED. |
[APPENDED 2026-08-21] Owner constraint — do not recommend cross-machine SendMessage/
ListAgents addressing for this owner. The owner rotates Claude accounts whenever a
usage limit is exhausted. Remote Control is account-scoped, so RC/cloud-session addressing
is not a usable transport across a rotation: a session addressed under one account
disappears from ListAgents/SendMessage reach once the owner rotates to another. The
file-based DevSwarm mesh (~/.anti-hall/devswarm/store/) is account-agnostic and survives
rotation — it has no equivalent blind spot. Consequence for any future recommendation in
this doc or its adoption plan: score Claude-native messaging on its LOCAL/same-machine
capability only; do not propose it as a cross-machine coordination layer for this owner.
UNVERIFIED: whether same-machine, same-account local-session addressing (the
"Other LOCAL Claude Code sessions" row in §3's cross-session-messaging table) is itself
account-scoped was not tested this session — flagged, not asserted either way.
[APPENDED 2026-08-21] This repo's own tools/eval/ harness (distinct from evals/).
tools/eval/ (singular) exists at the repo root: run.js (433 lines) + 4 pilot graders
(ship-it, scope-fidelity, rule-behavior, false-done, 986 lines total) +
README.md (296 lines). evals/ (plural — the claude plugin eval CLI convention
scaffolded by eval init) does not exist in this repo; the two names are easy to
conflate and refer to different, non-interoperating things.
Capability boundary — honest, not equivalent to claude plugin eval: the tools/eval/
pilots score what the agent STATES it would do, not what it actually did — see
tools/eval/…/ship-it-pilot.mjs:7-11, which documents this directly. claude plugin eval run
with --allow-tools plus a tool_used: Skill grader is categorically different: it
proves the skill actually fired as a real tool call in a sandboxed execution, not
merely that the model claimed it would. Neither harness subsumes the other — the
in-repo pilots are cheap self-report checks; claude plugin eval is the mechanically
verified alternative, with actual case execution still UNVERIFIED this session (see the
claude plugin eval entry in §3's Additions table above).
[APPENDED 2026-08-21] CI gap. package.json:7's "test" script is node --test
only — CI runs no tools/eval/ pilot and no claude plugin eval suite. Neither harness's
result gates a merge or release today.
3. Full Claude Code harness feature surface (2026-08)¶
Hook events¶
Source: /docs/en/hooks.
| Event | One-line description |
|---|---|
SessionStart |
Fires when a session starts or resumes (incl. after /compact, source=compact). |
Setup |
Fires on first-time plugin/project setup. |
SessionEnd |
Fires when a session terminates. |
UserPromptSubmit |
Fires when the user submits a prompt, before the model sees it. |
UserPromptExpansion |
Fires when a prompt is expanded (e.g. slash-command/skill substitution). |
PreToolUse |
Fires before a tool call executes; can block/modify. |
PostToolUse |
Fires after a tool call completes successfully. |
PostToolUseFailure |
Fires after a tool call fails. |
PostToolBatch |
Fires after a batch of parallel tool calls completes. |
PermissionRequest |
Fires when a permission prompt would be shown. |
PermissionDenied |
Fires when a permission request is denied. |
Stop |
Fires when the agent is about to stop responding this turn. |
SubagentStart |
Fires when a spawned subagent session starts. |
SubagentStop |
Fires when a spawned subagent session stops. |
StopFailure |
Fires when a stop/completion attempt itself fails. |
TeammateIdle |
Fires when a teammate/agent in a multi-agent session goes idle. |
PreCompact |
Fires before context compaction runs. |
PostCompact |
Fires after context compaction completes. |
TaskCreated |
Fires when a task is added to the task list. |
TaskCompleted |
Fires when a task is marked complete. |
InstructionsLoaded |
Fires when CLAUDE.md/AGENTS.md/instruction files are loaded. |
ConfigChange |
Fires when settings/config files change. |
CwdChanged |
Fires when the working directory changes. |
FileChanged |
Fires when a watched file changes on disk. |
Notification |
Fires on harness notifications (e.g. permission-needed, idle). |
MessageDisplay |
Fires when a message is rendered to the user. |
WorktreeCreate / WorktreeRemove |
Fire on git worktree lifecycle events. |
Elicitation |
Fires when the harness elicits structured input (e.g. from MCP). |
Long-running / background¶
Source: /docs/en/tools-reference.
| Feature | One-line description |
|---|---|
Monitor |
Registers a background watcher process that can emit wake events. |
Background Bash |
Runs a shell command detached, polled/notified on completion. |
Task / Agent tools |
Spawn subagents/sub-sessions for delegated work. |
SendMessage |
Sends a message to another agent/teammate session. |
PushNotification |
Pushes a notification to the user outside the transcript. |
SendUserFile |
Delivers a file artifact to the user. |
Scheduling¶
Source: /docs/en/scheduled-tasks, /docs/en/routines.
| Feature | One-line description |
|---|---|
/loop |
Re-runs a prompt/command on a recurring interval within a session. |
CronCreate / CronList / CronDelete |
Create/list/delete cron-style scheduled jobs that fire independent of an open REPL. |
ScheduleWakeup |
Schedules a one-time future wake for the current session. |
| Routines | Cloud-scheduled recurring agents (cron-driven, run headless). |
RemoteTrigger |
Triggers a remote/cloud agent run externally. |
Plugin components¶
Sources: /docs/en/plugins, /docs/en/plugins-reference, /docs/en/headless, /docs/en/settings, /docs/en/permissions, /docs/en/memory.
| Feature | One-line description |
|---|---|
Plugin system (bin/, output-styles/, workflows/, monitors/, .mcp.json, .lsp.json) |
Declarative manifest surface a plugin ships components through. |
| Skills | Reusable, invokable instruction packages (SKILL.md + assets). |
| Slash commands | User-typed shortcuts that expand to a prompt/skill. |
| Subagents | Named agent personas with scoped tools/model. |
| Output styles | Alternate system-prompt presentation modes. |
| Statusline | Persistent status bar rendered above the input. |
| MCP | Model Context Protocol server integration (tools/resources). |
| LSP | Language Server Protocol integration for diagnostics/navigation. |
| Workflow tool | Programmatic multi-agent orchestration primitive. |
| Agent SDK / headless mode | Programmatic, non-interactive Claude Code execution. |
| Settings permissions | Allow/deny/ask rules for tools and commands. |
| Memory | CLAUDE.md/project-memory persistence layer. |
Newer 2026 features¶
| Feature | One-line description | Doc |
|---|---|---|
| Checkpointing / rewind | Save and roll back to a prior conversation/file state. | /docs/en/checkpointing |
| Sandboxing | Contain tool execution (filesystem/network) inside a restricted boundary. | /docs/en/sandboxing |
| Permission modes | Named presets (default/acceptEdits/plan/bypass/etc.) governing tool approval. | /docs/en/permission-modes |
| Channels | Structured multi-party communication surface. | /docs/en/channels |
| Sessions: resume / branch / fork | Session lifecycle operations beyond linear continuation. | — |
| Remote Control | Externally drive/observe a running session. | — |
| Worktrees | Git-worktree-scoped isolated session workspaces. | — |
Flagged / unverified¶
- The
/docs/en/settingssummary listed hook namesbeforeBash/afterBash/beforeWrite/afterWrite/configChangethat do not appear on/docs/en/hooks. Treat this as a summarization artifact — use the/docs/en/hooksevent names (PreToolUse/PostToolUse/ConfigChange, etc.), not the settings-summary names. agent-teams,agent-view,desktop-scheduled-tasks, and the Remote Control page were referenced during this audit but not fetched — their contract is unverified. [UPDATE 2026-08-21: not re-checked this session — mark UNVERIFIED-since-2026-08-01, left as-is.]/docs/en/llms.txtreturned 404 at audit time. [UPDATE 2026-08-21: not re-checked this session — mark UNVERIFIED-since-2026-08-01, left as-is.]
Cross-session / agent-to-agent messaging¶
Local Claude Code version observed: 2.1.238. Tools: SendMessage, ListAgents.
| target class | bidirectional | durable | wakes idle | addressable when not running |
|---|---|---|---|---|
| In-process subagents (spawned this session) | yes | NO | N/A (not idle-capable) | session-scoped only |
| Agent-team teammates | yes | NO | starts a turn when idle | only while session active |
| Other LOCAL Claude Code sessions (same machine) | yes | NO | starts a new turn if idle | NO — message dropped if target not running |
| Cloud/web sessions | one-way if sender not connected to Remote Control (receives, cannot reply) | NO | appears in conversation | only while running |
| Remote Control sessions (other machines) | yes, if sender is on Remote Control | NO | appears in conversation | only while running |
Documented behavior: "The receiving Claude reads the message between tool calls during an active turn. When the receiving session is idle, Claude Code starts a new turn with the message."
DURABILITY VERDICT: SendMessage is live-session IPC, NOT a durable message bus.
Messages to a non-running session are dropped silently with no error to the sender. No
ack, no read receipt, no replay, no queue-for-later. Does not survive the target's
exit/crash, /clear, /compact, or a machine restart.
Nuance — do not overstate: durability for offline targets is UNDOCUMENTED, not
documented-as-absent. Agent teams do persist a per-agent inbox JSON at
~/.claude/teams/{team}/inboxes/{agent}.json, so some on-disk state exists, but no
offline-delivery mechanism is described anywhere in the docs. Mark as UNVERIFIED — needs
an empirical test (see §4 below).
Source: /docs/en/cross-session-messaging, /docs/en/agent-teams
Agent teams (EXPERIMENTAL)¶
Opt-in via CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 (v2.1.178+). Multiple coordinated
sessions, a shared task list, and direct messaging; split panes or in-process. Documented
limitations: no session resumption with in-process teammates, task status can lag, no
nested teams. Architecture: per-agent inbox JSON under ~/.claude/teams/{team}/inboxes/.
Flagged EXPERIMENTAL — unsuitable as the foundation for a shipped substrate today.
Source: /docs/en/agent-teams
Additions to the feature surface (absent from every prior doc in this repo — verified by grep)¶
[UPDATE 2026-08-21: the header claim above ("absent from every prior doc in this repo")
is FALSE for two of the rows below — corrected per-row rather than rewritten, per this
repo's own "never silently rewrite" rule (KB.md:37). The "absent" framing holds ONLY
for the ListAgents row; the other rows below were not all independently re-checked, but
isolation:"worktree" is confirmed already documented elsewhere in this repo (see its row).]
| Feature | Status | One-line description |
|---|---|---|
ListAgents |
Stable | Enumerates addressable agents: in-process subagents, other local sessions, cloud sessions, and (when Remote Control is connected) the account's other sessions. Names are the address for SendMessage. Confirmed genuinely absent elsewhere in this repo: 0 code hits repo-wide (docs-only usage). |
EnterWorktree / ExitWorktree |
Stable | Worktree entry/exit tools. |
Artifacts — capabilities |
Gated per-account | Runtime capabilities for published interactive pages beyond static HTML. |
claude plugin eval |
[UPDATE 2026-08-21: was "Org-flag gated" — corrected below, see new row] | superseded, see below |
Agent tool subagent_type: "fork" |
Stable (v2.1.212+) | Spawns a fork that inherits the parent's full context. |
Agent tool isolation: "worktree" |
Stable (v2.1.203+) | Runs the spawned agent in a temporary git worktree. [UPDATE 2026-08-21: NOT absent from prior docs — already documented at skills/orchestration/SKILL.md:27 AND docs/KB-claude-workflow-orchestration.md:58,104 before this row was added. The "absent from every prior doc" framing in this section's header does not hold for this row.] |
Monitor |
Stable, with a gap | Unavailable on Bedrock/Vertex/Foundry. |
CronCreate / CronList |
Stable | Scheduled tasks within a session; restored on --resume. |
[UPDATE 2026-08-21] claude plugin eval: corrected from "Org-flag gated" — the CLI is
AVAILABLE LOCALLY (verified this session): CLI version 2.1.238, claude plugin eval
--help renders successfully. Flags observed: --ablation with-without, --allow-tools,
--threshold, --json, --report, --no-publish, --max-cost-usd, --judge-model,
--case, --tag; eval init [--bare] scaffolds a suite. Suite convention: an
<eval-dir>/**/case.yaml OR a prompt.md + graders/*.md pair; results land at
<eval-dir>/results/<timestamp>/aggregate-result.json. Graders: regex, LLM judge,
tool_used, file_exists, baseline. --json v1 shape (schemaVersion,
cases[].arms.{with,without}[].graders, aggregates), --report HTML, --no-publish
local. Exit codes: 0 pass (default threshold 1.0), 1 fail/error/empty, 2 partial
(cost-ceiling / auth-fail). Still UNVERIFIED: whether a case actually EXECUTES — no
run was attempted this session; only --help and flag enumeration were checked. Do not
claim a passing/failing run without evidence.
Source: tools reference, sub-agents, hooks, plugin eval, docs index
4. Implication for anti-hall's DevSwarm mesh¶
Question asked: could Claude Code's native agent-to-agent messaging (SendMessage /
ListAgents / agent teams) REPLACE anti-hall's DevSwarm mesh store
(~/.anti-hall/devswarm/store/<hash>/devswarm.db, introduced v0.58.0)?
Answer: NO — they solve different problems. The mesh is a DURABLE system of record that
survives session death and supports replay/audit; SendMessage addresses only LIVE agents
and drops silently otherwise.
| property | DevSwarm mesh | SendMessage |
|---|---|---|
| Survives target death | ✅ | ❌ |
| Sender learns of non-delivery | ✅ | ❌ |
| Wakes an idle session | via cron + Monitor watcher | ✅ starts a new turn — genuinely better |
| Cross-machine | ❌ local SQLite | ✅ via Remote Control |
| Replay + audit | ✅ | ❌ |
Recommended posture: LAYER, don't replace. Keep the mesh as the durable store of
record; Claude's native messaging is a candidate LOW-LATENCY WAKE + live-coordination
layer over it — conceptually what
plugins/anti-hall/companion/lib/devswarm-wake-watch.js already does with a Monitor.
The one genuinely NEW capability it unlocks is cross-machine reach, which a local
SQLite mesh structurally cannot provide.
Open / unverified: an empirical test of what actually happens to a message sent to a dead session — is the inbox JSON written and later drained, or is it truly dropped? Until tested, treat durability as absent.
Assessed 2026-08-21 against Claude Code 2.1.238.
Audited 2026-08-01. Re-verify hook contracts against current docs before building against them — see the adoption plan's Fable-review step. The
adoption plan was Fable-reviewed 2026-08-02, which corrected several hook
contracts assumed above: PostCompact cannot inject additionalContext
(side-effect-only), SubagentStop injects into the subagent's own turn
rather than the parent, and ConfigChange does not watch
~/.anti-hall/skip.json.
Extended 2026-08-21 (Claude Code 2.1.238): added the cross-session/agent-to-agent
messaging and agent-teams subsections plus §4's DevSwarm-mesh-vs-SendMessage
comparison. The empirical test of message delivery to a dead session (inbox JSON
written-and-drained vs. truly dropped) remains open — see §4.