How to Detect When Codex Exec or Claude Code Finishes
Prepared and checked using AI. Sources are linked in the article.
To detect when codex exec or a Claude Code headless run is finished, check both the child process and its final event. Codex’s --json stream includes turn.completed and turn.failed events, while Claude Code’s --output-format stream-json stream ends with a result message. A process exit tells you the CLI stopped; a successful event tells you its agent loop finished. Neither proves the requested change is correct, so review the diff and run your own checks before accepting it.
Use separate states for stopped, finished, and accepted
A useful run record has a process status, a final event, and a review status. Keep these separate when several agents are working at once:
| Evidence | State to show | Next action |
|---|---|---|
| Process running; events still arriving | Working | Keep monitoring. |
| Process running; no recent event | Possibly stalled | Inspect the latest event and error output before intervening. |
| Process exited nonzero | Failed or interrupted | Keep the log and inspect the error. |
| Process exited; expected final event missing | Incomplete record | Investigate the launcher, signal, or truncated output. |
| Process exited; successful final event present | Ready for review | Check files, tests, and the diff. |
| Review checks passed and change meets the task | Accepted | Proceed with your normal integration process. |
Silence alone is weak evidence. An agent can spend time on a command without emitting a new event, and a log can stop growing while the process is still alive. Conversely, a final answer in a log does not establish that the launcher has collected the process exit status. Record both before moving a run out of the working state.
Capture each process and its event stream
Bash wait returns the exit status of a specified background job. Save each process ID when launching agents, then wait for each ID separately. An unqualified wait tells you when all jobs have stopped but does not give you a separate result for every agent.
For example, suppose ../wt/codex-auth and ../wt/claude-auth are separate worktrees and run-logs/ already exists. Before launch, record each worktree’s starting commit with git rev-parse --verify HEAD so you can review changes even if an agent commits them. This Bash script also gives each run its own JSONL and error files:
#!/usr/bin/env bash
set -u
(
cd ../wt/codex-auth || exit
git rev-parse --verify HEAD
) >run-logs/codex-auth.base || exit
(
cd ../wt/claude-auth || exit
git rev-parse --verify HEAD
) >run-logs/claude-auth.base || exit
(
cd ../wt/codex-auth || exit
codex exec --sandbox workspace-write --json \
"Fix the authentication tests and report which checks ran"
) >run-logs/codex-auth.jsonl 2>run-logs/codex-auth.err &
codex_pid=$!
(
cd ../wt/claude-auth || exit
claude -p "Fix the authentication tests and report which checks ran" \
--permission-mode acceptEdits \
--output-format stream-json --verbose
) >run-logs/claude-auth.jsonl 2>run-logs/claude-auth.err &
claude_pid=$!
wait "$codex_pid"
codex_status=$?
wait "$claude_pid"
claude_status=$?
printf 'codex=%s claude=%s\n' "$codex_status" "$claude_status"
The explicit Codex sandbox matters: codex exec defaults to a read-only sandbox, and --sandbox workspace-write permits edits in the workspace. Claude Code’s acceptEdits mode permits file edits, but other shell commands may still require a permission rule. If a requested test cannot run, classify that check as unverified even if the agent produces a final response.
The script waits in launch order, so the first wait can block after the second agent has already stopped. The individual log files still let another terminal display progress. For a larger supervisor, retain a map from process ID to run ID and record each exit independently; avoid reducing all agents to one combined success flag.
Read the terminal event, not just the last line of prose
jq processes a stream of JSON values and -e reports a nonzero status when a filter produces no valid result. After the relevant process exits, these commands answer whether the expected success event exists:
jq -e 'select(.type == "turn.completed")' \
run-logs/codex-auth.jsonl >/dev/null
jq -e 'select(.type == "result" and .subtype == "success")' \
run-logs/claude-auth.jsonl >/dev/null
Run those checks separately from the agents. A missing match, malformed JSON, or nonzero process status needs investigation; do not turn every nonzero jq result into “the agent failed.” To inspect outcomes without hiding error events:
jq -c 'select(.type == "turn.completed" or
.type == "turn.failed" or
.type == "error")' run-logs/codex-auth.jsonl
jq -c 'select(.type == "result" or
.type == "system")' run-logs/claude-auth.jsonl
For Codex, turn.failed is a failed turn, while turn.completed means that turn ended. Inspect command execution items and the final message as well: a completed turn can still describe work the agent could not finish. --output-last-message is useful when you need the final prose in a separate file, but that prose is not a replacement for the event stream or process status.
For Claude Code, the final result’s subtype distinguishes success from limits and execution errors. Examples include error_max_turns, error_max_budget_usd, and error_during_execution. Treat these as stopped runs requiring attention, even if earlier messages contain useful partial work. Read the result text only for a success subtype; an error result can have no final text.
Handle runs that appear stuck
Set two clocks in your supervisor: elapsed time since launch and elapsed time since the last event. A quiet interval should trigger inspection, not an automatic declaration of failure. Check the most recent JSON event, the separate error file, and whether the child process still exists. A recent command item may explain the silence; a permission denial or repeated error may call for a change to the run configuration.
For unattended Claude Code runs, --permission-prompts none denies unresolved requests instead of waiting for a permission host. It requires Claude Code v2.1.259 or later. The stream can report denials as permission_denied system messages, and the final result lists permission_denials; use those fields to distinguish “finished with denied work” from “finished as requested.” Do not enable broader permissions merely to make a timer stop firing.
A wall-clock limit is a separate safeguard. With GNU timeout, status 124 means the duration expired unless --preserve-status was requested. For example, this gives a run 30 minutes, then sends a stronger signal after another 30 seconds if it has not stopped:
timeout --kill-after=30s 30m \
codex exec --sandbox workspace-write --json \
"Investigate the authentication failure" \
>run-logs/codex-limited.jsonl 2>run-logs/codex-limited.err
status=$?
Choose the limit for the task and environment; the example durations are a policy choice, not a measure of normal agent speed. GNU timeout can return 137 after a KILL signal, which does not by itself identify whether the managed command or timeout received that signal. Preserve the stream and error output so an interrupted run is not mistaken for a completed one. On a system without GNU timeout, use a supervisor with equivalent, documented signal and exit-status behavior.
Verify the work after a successful run
A successful terminal event means the agent stopped normally, not that its proposed fix passed your acceptance criteria. From the directory where the launcher ran, use the saved starting commit to review the Codex worktree:
base=$(cat run-logs/codex-auth.base)
(
cd ../wt/codex-auth || exit
git status --porcelain=v1 --untracked-files=all
git diff "$base"
git diff --check "$base"
)
Git’s status --porcelain=v1 --untracked-files=all lists individual non-ignored untracked files, even if status.showUntrackedFiles would otherwise hide them or summarize their directory. Status shows current uncommitted changes; comparing the working tree with the saved commit shows the net tracked changes since launch, including committed changes and any remaining uncommitted edits. git diff --check "$base" reports conflict markers and whitespace errors in that comparison and exits nonzero when it finds them. Git diff does not show untracked file contents, so inspect those files separately.
Run the project’s relevant tests yourself, inspect their exit status, then review the patch against the original request. If an agent says it ran tests, use the recorded command output or rerun the checks before marking them passed.
Keep a separate record for each agent: starting commit, process exit, final event, permission denials, changed files, check results, and review decision. That makes a parallel run easier to triage than a single “done” badge. If you are preparing to combine changes, the workflow in merging parallel agent branches starts after each candidate has been reviewed; reviewing AI-generated code covers the patch inspection itself.
Where Parallel Code fits
Parallel Code is our free, open-source desktop app for macOS and Linux. It runs agents such as Claude Code and Codex CLI in parallel, each in its own git worktree, and provides a diff-first review surface. Its mobile progress monitoring works via QR code over Wi-Fi or Tailscale, which can help you keep track of background work away from the desktop.
Frequently asked questions
Does exit code 0 mean the coding task succeeded?
It means the CLI process returned successfully to its launcher. Check the tool’s final event, then verify the requested files, tests, and diff before accepting the task.
Can I decide a run is stalled because its log stopped growing?
No. Treat inactivity as a prompt to inspect the process and its latest event. Use a task-specific wall-clock limit if an unattended run must eventually stop.
What if Claude Code reports error_max_turns?
The agent stopped at its turn limit before finishing. Keep the partial work for review, then decide whether to continue the session with a larger limit or narrow the task; do not count the result as completed.
What if the process exits but there is no final JSON event?
Mark the record incomplete and inspect the exit status and error output. A signal, startup failure, or truncated stream may have prevented a terminal event, so the final prose or partial edits alone should not be treated as completion.