TASK-WORK — Task discipline for agentic Claude work¶
Knowledge base for the
anti-halltasklist-guard feature: enforced, self-refreshing task discipline. Distills Claude Code's native task tooling, Anthropic's long-running-agent guidance, and the hook model into implementation-ready facts.External facts (Claude Code task tooling, hook behaviour, the linked issues) are as documented by Anthropic and the cited sources when this doc was written; they are not re-verified on every release — check the current docs before relying on one.
Overview¶
Claude Code has two generations of task tooling: the legacy TodoWrite tool
(one call rewrites the whole list) and the current Task tools
(TaskCreate / TaskUpdate / TaskList / TaskGet, one task per call, with
status, ownership, and dependencies). Task tools became the default in Claude
Code v2.1.142 / TS Agent SDK 0.3.142; set CLAUDE_CODE_ENABLE_TASKS=0 to fall
back to TodoWrite.
A guard that enforces task hygiene must (a) read the live list, (b) recognize the tool calls and results in the transcript, and (c) act only at the lifecycle points the hook model actually fires — there is no wall-clock timer, so "always checking" is approximated per turn and at stop time.
Claude task tooling¶
TodoWrite (legacy)¶
- Shape: one call carries the entire
todosarray. Each item is{ content, status, activeForm }.status∈pending | in_progress | completed. content= imperative ("Fix the auth bug");activeForm= present-continuous shown while in progress ("Fixing the auth bug").- Rewrites the whole list every call — no per-item identity, no dependencies, no ownership, no cross-session persistence.
- Transcript:
tool_useblock withname === "TodoWrite", full list inblock.input.todos. To track, replace your local copy on each call.
Task tools (current default)¶
TaskCreate splits one item out; TaskUpdate patches one item by id; TaskList
and TaskGet read back. Designed for long, multi-session, multi-subagent work.
| Tool | Input | Returns |
|---|---|---|
TaskCreate |
{ subject, description, activeForm?, metadata? } |
tool_result = { task: { id, subject } } |
TaskUpdate |
{ taskId, status?, subject?, description?, activeForm?, addBlocks?, addBlockedBy?, owner?, metadata? } |
updated task |
TaskList |
(filter) | records with only id, subject/title, status, owner, blockedBy |
TaskGet |
{ taskId } |
full record incl. description, metadata, activeForm |
subject= brief imperative title;description= the detail;activeForm= present-continuous for the in-progress spinner.- Status: every task is born
pending; move toin_progressbefore work,completedwhen done.status: "deleted"removes it. - Dependencies:
addBlockedBymakes this task wait on others;addBlocksis the inverse (this task gates those). A task is "available" whenstatus === "pending"ANDowneris empty ANDblockedByis empty (all deps resolved). - Ownership:
ownerkeys multi-agent coordination — an unowned available task is free to claim; an owned one is someone's in-flight work. - Owner-blocked marker (anti-hall
task-guardconvention, not a harness field): a task that is genuinely non-dispatchable because it is blocked on the OWNER — hardware, a decision only a human can make, anything no agent can resolve — should be marked honestly rather than given a fakeblockedBypointing at a nonexistent/unrelated task id (a real field incident: a Primary faked a dependency purely to silence the nag). Mark it withmetadata.blockedOn: 'owner'('user'/'human'/'external'also recognized, case-insensitive), or give the subject an"OWNER:"/"OWNER DECISION"prefix (case-insensitive).task-guard.js's IDLE NEGLECT check (isOwnerBlocked()) treats either as non-dispatchable — it is excluded from the ACTIONABLE-NOW set and never nagged, without needing a fabricated blocker. The generic "open tasks remain" nudge also leaves out marked tasks and tasks whoseblockedBynames a still-open task; if every open task is blocked, there is no nudge. Toggle:guards.taskGuardOwnerBlockedMarker(default on). - Parallel-dispatch demand: every turn with a pending, unblocked, unowned,
not-owner-blocked task that no in-flight agent covers,
task-trackerinjectsDISPATCH NOW in parallel — …: #7 "subject", #11 "subject", …; at Stop the same set drivestask-guard's IDLE NEGLECT block (capped). Coverage is per task from THIS session's transcript: a running background agent whose description names#<id>covers that task; an agent naming none is assumed to be on an in_progress task first, else on one pending task. The Stop block itself needs proof: it fires only when dispatchable tasks outnumber the running agents that name no task (guards.idleNeglectProvenOnly, default on). In that proof an unmapped agent does not count as cover once it has shown no sign of life forguards.idleNeglectAgentMaxAgeMin(default 30) minutes, or when it was launched before the uncovered task existed. The machine-global~/.anti-hall/agentsheartbeat no longer suppresses it (any session's spawn refreshed it). Toggle:guards.dispatchDemand(default on). Metrics (demandsShown,demandsFollowed,demandsIgnored, compliance rate,idleNeglectBlocks):node plugins/anti-hall/scripts/dispatch-report.jsand onedoctorline. - Persistence: written to disk immediately under
~/.claude/tasks/<TASK_LIST_ID>/(index.json+task-*.json). Survives compaction, restart, and multi-day gaps.CLAUDE_CODE_TASK_LIST_IDselects the active list; reuse it to resume across sessions.
Transcript shape (critical for a transcript-scanning hook)¶
- Match assistant
tool_useblocks byblock.name:"TaskCreate","TaskUpdate","TaskList","TaskGet"(or"TodoWrite"in legacy mode). - The new task id is NOT in the
TaskCreateinput. It comes back in the matchingtool_resultas{ task: { id, subject } }. A guard reconstructing the list must key its map off the result block, then apply laterTaskUpdateinputs bytaskId. - To get an authoritative snapshot rather than replaying deltas, watch for a
TaskListtool_result(its records carryid/subject/status/owner/blockedBy). TodoWriteis simpler to scan: every call has the full state inblock.input.todos; take the last one.
Live-list best practices (per Anthropic / system prompts)¶
- Use it for non-trivial work: 3+ distinct steps, multi-part requests, plan mode, or any explicit user task list. Skip for single trivial / purely conversational actions.
- Capture immediately: turn user requirements into tasks up front; add follow-ups the moment they're discovered. Don't lose work to memory.
- Exactly ONE
in_progressat a time — "not less, not more." Mark a taskin_progressbefore starting it. - Complete only when FULLY done. Never mark
completedif tests fail, implementation is partial, errors remain, or files/deps are missing — leave itin_progressor split out a blocker. - Don't batch completions — mark complete immediately after each finish so the list reflects reality in real time.
- Don't over-list: ~7+ steps tempts the model to batch them "for efficiency," defeating the checkpoints. Prefer specific, verifiable items ("navbar height 60→80px") over vague ones ("style the navbar").
- Prune: remove (or
deleted) tasks that are no longer relevant rather than leaving stale entries. - Keep it non-blocking: the main thread coordinates; delegate heavy/long work to subagents so list maintenance never stalls behind a build.
Anthropic's agentic loop = gather context → act → verify → repeat; the task list is the spine of that loop. Verification should be evidence-based (show the command + output), ideally by a fresh verifier, not self-asserted.
Dedup & smart-relate¶
The system prompt explicitly says: check TaskList before TaskCreate to
avoid duplication. Practical heuristics for an incoming request:
- Read the live list first (
TaskList, or the reconstructed transcript map). - Duplicate → same intended outcome as an existing open task. Signals:
high subject/description token overlap, same target file/symbol/endpoint, same
verb+object. Action: do not create; surface the existing task id, optionally
refine its
description. Supersede only bydeleted+ one replacement. - Related but distinct → shares scope but is a different deliverable
(sub-step, prerequisite, or follow-on). Action: create it and link via
addBlockedBy(prereq) oraddBlocks(this gates that). Usemetadatato tag a cluster/epic. - Genuinely new → no scope overlap. Create normally.
Heuristic ordering for a guard: exact-subject match → normalized-subject match
(lowercase, strip stopwords) → target-artifact match → semantic similarity over
subject + description. Above a high threshold = duplicate; mid threshold =
relate via dependency; low = new.
Staleness & freshness (and the event-driven, not-timer constraint)¶
Anthropic's long-running-agent guidance:
- Keep a progress/state file (their example:
claude-progress.txt) logging what's been done; new sessions read git logs + progress file to recover state before picking work. - Maintain a feature/task list initialized as "failing"/incomplete, and only
flip to "passing"/
completedafter careful self-verification. This directly defends the named failure mode: "a later agent instance would look around, see that progress had been made, and declare the job done." - "It is unacceptable to remove or edit tests" to make things look done.
Staleness signals for a guard to detect:
- a task in_progress with no related tool activity for many turns (untouched);
- more than one in_progress (invariant violation);
- pending tasks that are now available (deps cleared, unowned) but ignored;
- requests in the conversation never captured as tasks (forgotten / orphaned);
- a completed task with no verification evidence in the transcript.
The hard constraint — no native timer. Claude Code hooks are event-driven only; there is no cron/interval. "Always checking" must be approximated by firing on lifecycle events:
| Event | Fires | Useful for | Output channel |
|---|---|---|---|
UserPromptSubmit |
each user turn, before Claude sees it | per-turn freshness sweep; inject reminders | additionalContext (exit 0) injects text into context |
Stop |
when Claude finishes a response | enforce "don't stop with work undone" | {"decision":"block","reason":...} forces continuation |
TaskCreated / TaskCompleted |
native task lifecycle events | react to task changes directly | standard hook output |
PostToolUse / PostToolBatch |
after tool calls | observe TaskCreate/Update results live |
— |
Key field facts (for the guard script reading stdin JSON):
- Every event includes session_id, cwd, hook_event_name.
- The transcript path arrives as transcript_path (JSONL). Cold-start
caveat: a Stop hook may miss the very first turn or find the JSONL not yet
flushed — handle a missing/empty transcript gracefully.
- UserPromptSubmit also gets prompt; on exit 0 anything on stdout (or
additionalContext JSON) is injected into Claude's context. Its timeout is
lowered to 30s — keep the sweep fast.
- Stop (and PostToolUse) use top-level decision: "block" with a
reason; the reason is the next instruction Claude acts on, so make it
specific ("Task #3 is in_progress but untested — verify before stopping").
- Stop input carries stop_hook_active — when true you already forced one
continuation; bow out to avoid an infinite loop (one push-back per stretch).
- Exit codes: exit 2 blocks with stderr fed back to Claude; exit 0 = no
objection (and for UserPromptSubmit, stdout becomes context). Don't mix exit 2
with JSON — JSON is ignored when you exit 2.
- Plugin caveat: there is a known issue where Stop hooks installed via
plugins don't reliably continue on exit 2 — prefer the JSON decision:"block"
form and test under plugin packaging.
Design implications for anti-hall enforcement¶
Goal: enforce (not merely nudge) continuous, deduped, fresh task discipline, treated as a feature launch.
- Two coordinated hooks, no timer:
UserPromptSubmitfreshness sweep (≤30s, fast): read the live list (parsetranscript_pathJSONL forTask*/TodoWriteblocks + their results; key offtool_result { task: { id } }). InjectadditionalContextsummarizing open tasks, the single allowedin_progress, any newlyavailabletasks, and any uncaptured request from the just-submitted prompt — and run the dedup/relate check on that prompt against the live list.-
Stoptasklist-guard: if there's unverified or in-flight work (>1in_progress— multiple-in-progress at once is the smell; a singlein_progressis the healthy invariant and never triggers this cause — plus available-but-ignored tasks, or completed-without-evidence), return{"decision":"block","reason":"<specific next action>"}. Always short-circuit whenstop_hook_activeis true. -
Invariant enforcement: reject/flag transcripts that show >1
in_progress, acompletedwith no verification evidence, or a request that produced no task. -
Dedup/relate at capture time: before encouraging a
TaskCreate, diff the prompt againstTaskList(exact → normalized → artifact → semantic). Duplicate ⇒ point at existing id; related ⇒ suggestaddBlockedBy/addBlockslink; new ⇒ allow. -
State file as ground truth: keep a per-session progress file at
.anti-hall/progress/<date>/<session-id>.md(<date>= UTCYYYY-MM-DD,<session-id>= the sanitizedsession_id) refreshed so freshness survives compaction and cold starts (mirrors Anthropic'sclaude-progress.txt) — scoping the file per session (rather than one shared file) means two concurrent sessions on the same project never clobber each other's progress. NOTE: the hook only checks this file's freshness (its mtime, relative to the session cwd) — it never writes or refreshes it. The agent/user maintains the file each sweep; the guard just nudges when it's missing or stale. A running.anti-hall/progress/INDEX.mdlinks every session's file via idempotent, atomic single-line appends (never a read-modify-rewrite). Tasks themselves already persist under~/.claude/tasks/<TASK_LIST_ID>/— read that as the authoritative store whenCLAUDE_CODE_TASK_LIST_IDis known. -
Mode-agnostic scanning: support both
Task*andTodoWritetranscripts (the user may setCLAUDE_CODE_ENABLE_TASKS=0). ForTodoWrite, take the last fulltodosarray; for Task tools, fold create-results + update-inputs into a map, or trust the latestTaskListsnapshot. -
Non-blocking & resilient: keep the sweep cheap, tolerate missing transcript on cold start, never
rm/delete tasks, and prefer JSONdecision:"block"over exit 2 for portability across plugin packaging. -
Optional native events:
TaskCreated/TaskCompletedhooks let the guard react the instant a task changes, complementing the per-turn sweep.
Coordinator work window (report and baseline)¶
coordinator-work-guard counts the main thread's state-changing Bash calls (WORK) in a time window. See the "Coordinator work window" section of docs/GUIDE.md for what counts.
Report. node plugins/anti-hall/scripts/dispatch-report.js (and doctor) print one block:
- shown nudges, blocks, mean and max blocks per session, and the skipped-would-block counter (a would-be block passed by an explicit skip; it is not part of any share);
- per plugin version: sessions, the posted share
work / calls, and the attempted share(work + blocks) / (calls + blocks); - the known-gaps line.
Who applies patches. Integration agents apply patches. The coordinator has no git am or git apply exemption, apart from the recovery commands (--abort/--quit, stash pop|apply), which are counted but never blocked.
Baseline. To get a "before" number for sessions that ran without the guard, run
node plugins/anti-hall/scripts/coordinator-work-baseline.js ~/.claude/projects/<slug>/<session>.jsonl --json
on pre-release sessions. It prints {calls, work, share, attemptedShare, wouldNudge, wouldBlock}. share is as recorded (no enforcement): every logged call is counted. attemptedShare is with enforcement: a call the window would have blocked is not posted. Compare share against the per-version shares in the dispatch-report. The baseline resolves $VAR script paths from its own environment, not the original session's, cannot count direct-exec scripts that no longer exist on disk, and judges freshness against current file mtimes.
Sources¶
- Todo Lists — Claude Agent SDK docs — TodoWrite vs Task tools, exact input/result shapes,
{ task: { id, subject } }result, migration table. - Automate actions with hooks — Claude Code docs — event list, common input fields (
session_id,cwd,hook_event_name,transcript_path), exit-code semantics,additionalContext,decision:"block", event-driven (no timer). - Hooks reference — Claude Code docs — full event schemas, decision-control table.
- Effective harnesses for long-running agents — Anthropic Engineering — progress file, verify-before-passing, "declare the job done" failure mode.
- Best practices for Claude Code — Anthropic Engineering — agentic loop, evidence-based verification, plan/execute separation.
- TodoWrite tool description (Piebald-AI/claude-code-system-prompts) — one-in_progress rule, complete-only-when-done, no-batching, capture-new, prune.
- TaskCreate tool description (Piebald-AI/claude-code-system-prompts) — when to use, fields, check TaskList before creating to dedup.
- Task Operations and Lifecycle — DeepWiki — TaskList vs TaskGet field visibility, "available" definition, disk persistence path,
CLAUDE_CODE_TASK_LIST_ID. - Claude Code Stop Hook: force task completion — claudefa.st —
decision:"block"+reason,stop_hook_activeloop guard. - Stop hooks exit-2 plugin bug — anthropics/claude-code #10412 — plugin-packaged Stop hooks unreliable on exit 2.
- Stop/UserPromptSubmit cold-start timing — anthropics/claude-code #56631 — first-turn Stop miss; transcript not yet flushed.