Summary
Fix #531 added a 30-second post-kill grace period to allow the kernel to naturally EOF the stdout pipe after terminate_process_group(). However, agents using npx-based MCP servers still fail with "Execution completed without a result message" 502s after the update. The 30s wait always times out because MCP server subprocesses spawned via npx -y <package> are not in claude's process group — they survive terminate_process_group(pgid=claude_pgid) and keep the pipe write FD open indefinitely, preventing EOF from ever reaching the reader thread.
Observed as 100% failure rate on a specific schedule (hourly heartbeat) while a different schedule (work-loop, same agent, same container) succeeds. The work-loop doesn't invoke aistudio/moltbook MCP tools; the heartbeat does.
Component
Agent Runtime — docker/base-image/agent_server/utils/subprocess_pgroup.py
Priority
P1 — scheduled executions that use stdio MCP servers fail 100% of the time on affected agents
Error
Agent-server logs (new post-#531 sequence):
[Subprocess] Reader thread(s) still busy after process exit (pid=607, stuck_count=1)
— killing process group, then waiting 30s for natural drain
[Subprocess] Reader thread(s) still stuck after 30s post-kill grace
— force-closing pipes; some buffered data may be lost (pid=607, stuck_count=1)
[Subprocess] 1 reader thread(s) leaked for pid=607 after force-close; continuing anyway
[Headless Task] Error reading stdout: I/O operation on closed file.
[Headless Task] Execution completed without a result message after 0 tool calls / 22 turns
(raw_messages=40). Likely cause: a tool or child subprocess inherited stdout and prevented
the claude reader thread from capturing the final result block.
Backend logs:
[TaskExecService] Agent <name> responded: HTTP 502 (1648586ms)
[TaskExecService] Failed to execute task on <name>: Execution completed without a result
message after 0 tool calls / 22 turns (raw_messages=40). Likely cause: a tool or child
subprocess inherited stdout...
Circuit breaker fires continuously from ~1 minute into execution until the 502 is returned (~27 minutes), indicating the agent-server event loop is saturated during the execution.
Location
- File:
docker/base-image/agent_server/utils/subprocess_pgroup.py
- Function:
drain_reader_threads
- File:
docker/base-image/agent_server/services/claude_code.py
- Function:
execute_headless_task
Root Cause
Process group topology
When Claude Code starts, it launches stdio MCP servers defined in .mcp.json as child processes. For npx-based servers:
agent-server
└── claude (pgid=X)
├── [claude internals]
└── npm exec --yes -- <mcp-package> ← spawned by Claude Code's MCP manager
└── node <mcp-server-entry> ← spawned by npm, may be a new pgid
The npm → node chain typically starts in its own session/process group (npm calls setsid() or equivalent). Claude Code's MCP subprocess management may not set the child's process group explicitly either. Result: the MCP server process tree is outside claude's pgid.
What #531 fixed vs. what it didn't
#531 correctly reordered drain_reader_threads to:
- Kill
claude's process group
- Wait
post_kill_grace=30s for natural drain
- Force-close only as last resort
This fixes the original "buffer discarded before reader could drain it" scenario. But in this case, terminate_process_group(pgid=claude_pgid) only kills processes within claude's pgid. The npm/node MCP server processes — in a different pgid — survive, still holding an open write FD on claude's stdout pipe. The kernel cannot EOF the read end while any writer FD remains open. The 30s grace period expires without ever receiving EOF, so the reader thread stays blocked, the force-close path runs, and the result line (already in the buffer or waiting to be written) is lost.
Why 0 tool calls / 22 turns
The tool call count is zero because tool results never arrived: the MCP server processes may be slow to initialize (npm install on first run) or the agents actively invoke MCP tools whose responses never complete because the servers are hung. Claude runs through its 22 turns producing text but never receives tool results, then exits with return_code=0. The MCP server orphans are still running with the write FD open, so the reader blocks until forced.
Reproduction Steps
- Configure an agent with one or more stdio MCP servers using
npx -y <package> in .mcp.json
- Schedule a headless task on that agent that invokes those MCP tools
- Observe the task run: it will complete claude's execution (return code 0) but fail with HTTP 502
- Agent-server logs will show: "Reader thread(s) still stuck after 30s post-kill grace"
- Backend logs will show: "Execution completed without a result message after 0 tool calls"
Contrast: a different schedule on the same agent that does NOT invoke the npx MCP tools will succeed (work-loop vs. heartbeat pattern observed in production).
Suggested Fix
Two complementary approaches:
Option A — Kill orphan pipe-writers before waiting (Linux /proc scan)
After terminate_process_group(), scan /proc/*/fd for any process (not in our pgid) that has an open write FD pointing at the same pipe inode as claude's stdout, then SIGKILL those processes before starting the post_kill_grace wait:
import os, stat
def _kill_pipe_writers(pipe_write_fd: int, our_pgid: int) -> None:
"""Kill any process outside our_pgid that holds pipe_write_fd's inode open for writing."""
try:
target_ino = os.fstat(pipe_write_fd).st_ino
except OSError:
return
for pid_str in os.listdir("/proc"):
if not pid_str.isdigit():
continue
pid = int(pid_str)
try:
if os.getpgid(pid) == our_pgid:
continue # already killed above
except OSError:
continue
fd_dir = f"/proc/{pid}/fd"
try:
for fd_name in os.listdir(fd_dir):
fd_path = os.path.join(fd_dir, fd_name)
try:
st = os.stat(fd_path)
if stat.S_ISFIFO(st.st_mode) and st.st_ino == target_ino:
# Check it's writable (mode bit O_WRONLY or O_RDWR)
fdinfo = open(f"/proc/{pid}/fdinfo/{fd_name}").read()
if "flags:" in fdinfo:
flags = int(fdinfo.split("flags:")[1].split()[0], 8)
if flags & os.O_WRONLY or flags & os.O_RDWR:
os.kill(pid, signal.SIGKILL)
break
except OSError:
pass
except OSError:
pass
Call this after terminate_process_group() and before the post_kill_grace join loop in drain_reader_threads.
Option B — Prevent FD inheritance when Claude Code spawns MCP servers (preferred, upstream fix)
The cleaner fix is to ensure MCP server subprocesses never inherit the pipe FDs in the first place. Claude Code's MCP subprocess spawning should use close_fds=True (Python default on 3.2+, but subprocess spawning via shell may override) or explicitly set stdin=DEVNULL, stdout=PIPE, stderr=PIPE for MCP server processes so they get their own stdio.
This may need to be fixed in Claude Code's MCP client implementation rather than in Trinity's agent-server.
Option C — Track MCP server PIDs and kill them explicitly
If the agent-server starts MCP servers itself (or can enumerate them from Claude Code's process tree), record those PIDs at task start and add them to the kill list in drain_reader_threads.
Observed Behavior vs. Expected
|
Observed |
Expected |
| After #531 |
Reader thread blocked for full 30s grace, then force-closed → data lost |
Grandchild termination → EOF → reader drains → clean exit |
| npx MCP servers |
Survive terminate_process_group, hold pipe open indefinitely |
Should be killed or should not inherit the write FD |
| Task result |
HTTP 502, execution marked failed |
HTTP 200, execution marked success with result |
Environment
Related
Secondary Finding (non-blocking)
Several agent skills have allowed_tools formatted as a comma-separated string instead of a YAML list, causing pydantic parse warnings in the updated agent-server:
Failed to parse skill at .../SKILL.md: 1 validation error for SkillInfo
allowed_tools
Input should be a valid list [type=list_type, input_value='Bash, Read', input_type=str]
This appears to be a skill format validation tightening introduced in the v0.5.0 release. Skills with this format fail to load and their allowed_tools restrictions are not enforced. Tracked separately — skills still execute, the restriction is just not applied.
Summary
Fix #531 added a 30-second post-kill grace period to allow the kernel to naturally EOF the stdout pipe after
terminate_process_group(). However, agents usingnpx-based MCP servers still fail with "Execution completed without a result message" 502s after the update. The 30s wait always times out because MCP server subprocesses spawned vianpx -y <package>are not in claude's process group — they surviveterminate_process_group(pgid=claude_pgid)and keep the pipe write FD open indefinitely, preventing EOF from ever reaching the reader thread.Observed as 100% failure rate on a specific schedule (hourly heartbeat) while a different schedule (work-loop, same agent, same container) succeeds. The work-loop doesn't invoke aistudio/moltbook MCP tools; the heartbeat does.
Component
Agent Runtime —
docker/base-image/agent_server/utils/subprocess_pgroup.pyPriority
P1 — scheduled executions that use stdio MCP servers fail 100% of the time on affected agents
Error
Agent-server logs (new post-#531 sequence):
Backend logs:
Circuit breaker fires continuously from ~1 minute into execution until the 502 is returned (~27 minutes), indicating the agent-server event loop is saturated during the execution.
Location
docker/base-image/agent_server/utils/subprocess_pgroup.pydrain_reader_threadsdocker/base-image/agent_server/services/claude_code.pyexecute_headless_taskRoot Cause
Process group topology
When Claude Code starts, it launches stdio MCP servers defined in
.mcp.jsonas child processes. Fornpx-based servers:The
npm→nodechain typically starts in its own session/process group (npm callssetsid()or equivalent). Claude Code's MCP subprocess management may not set the child's process group explicitly either. Result: the MCP server process tree is outsideclaude's pgid.What #531 fixed vs. what it didn't
#531 correctly reordered
drain_reader_threadsto:claude's process grouppost_kill_grace=30sfor natural drainThis fixes the original "buffer discarded before reader could drain it" scenario. But in this case,
terminate_process_group(pgid=claude_pgid)only kills processes within claude's pgid. The npm/node MCP server processes — in a different pgid — survive, still holding an open write FD on claude's stdout pipe. The kernel cannot EOF the read end while any writer FD remains open. The 30s grace period expires without ever receiving EOF, so the reader thread stays blocked, the force-close path runs, and the result line (already in the buffer or waiting to be written) is lost.Why
0 tool calls / 22 turnsThe tool call count is zero because tool results never arrived: the MCP server processes may be slow to initialize (npm install on first run) or the agents actively invoke MCP tools whose responses never complete because the servers are hung. Claude runs through its 22 turns producing text but never receives tool results, then exits with return_code=0. The MCP server orphans are still running with the write FD open, so the reader blocks until forced.
Reproduction Steps
npx -y <package>in.mcp.jsonContrast: a different schedule on the same agent that does NOT invoke the npx MCP tools will succeed (work-loop vs. heartbeat pattern observed in production).
Suggested Fix
Two complementary approaches:
Option A — Kill orphan pipe-writers before waiting (Linux
/procscan)After
terminate_process_group(), scan/proc/*/fdfor any process (not in our pgid) that has an open write FD pointing at the same pipe inode as claude's stdout, then SIGKILL those processes before starting thepost_kill_gracewait:Call this after
terminate_process_group()and before thepost_kill_gracejoin loop indrain_reader_threads.Option B — Prevent FD inheritance when Claude Code spawns MCP servers (preferred, upstream fix)
The cleaner fix is to ensure MCP server subprocesses never inherit the pipe FDs in the first place. Claude Code's MCP subprocess spawning should use
close_fds=True(Python default on 3.2+, but subprocess spawning via shell may override) or explicitly setstdin=DEVNULL, stdout=PIPE, stderr=PIPEfor MCP server processes so they get their own stdio.This may need to be fixed in Claude Code's MCP client implementation rather than in Trinity's agent-server.
Option C — Track MCP server PIDs and kill them explicitly
If the agent-server starts MCP servers itself (or can enumerate them from Claude Code's process tree), record those PIDs at task start and add them to the kill list in
drain_reader_threads.Observed Behavior vs. Expected
terminate_process_group, hold pipe open indefinitelyEnvironment
d4bb2a1(post-bug: drain_reader_threads closes stdout pipe before reader can drain backlog — silently loses final result line on long agentic tasks #531, post-fix(agent): drain pipe before close to preserve final result line (#531) #532)npx -y <package>stdio servers in.mcp.jsonRelated
drain_reader_threadspipe-ordering fix (addresses different root cause)docker/base-image/agent_server/utils/subprocess_pgroup.py—drain_reader_threads,terminate_process_groupdocker/base-image/agent_server/services/claude_code.py—execute_headless_task.mcp.jsonstdio MCP server configurationSecondary Finding (non-blocking)
Several agent skills have
allowed_toolsformatted as a comma-separated string instead of a YAML list, causing pydantic parse warnings in the updated agent-server:This appears to be a skill format validation tightening introduced in the v0.5.0 release. Skills with this format fail to load and their
allowed_toolsrestrictions are not enforced. Tracked separately — skills still execute, the restriction is just not applied.