Skip to content

bug(agent-runtime): npx MCP server subprocesses outside claude pgid hold stdout pipe open — reader stuck after #531 grace period #618

Description

@vybe

Summary

Fix #531 added a 30-second post-kill grace period to allow the kernel to naturally EOF the stdout pipe after terminate_process_group(). However, agents using npx-based MCP servers still fail with "Execution completed without a result message" 502s after the update. The 30s wait always times out because MCP server subprocesses spawned via npx -y <package> are not in claude's process group — they survive terminate_process_group(pgid=claude_pgid) and keep the pipe write FD open indefinitely, preventing EOF from ever reaching the reader thread.

Observed as 100% failure rate on a specific schedule (hourly heartbeat) while a different schedule (work-loop, same agent, same container) succeeds. The work-loop doesn't invoke aistudio/moltbook MCP tools; the heartbeat does.

Component

Agent Runtime — docker/base-image/agent_server/utils/subprocess_pgroup.py

Priority

P1 — scheduled executions that use stdio MCP servers fail 100% of the time on affected agents

Error

Agent-server logs (new post-#531 sequence):

[Subprocess] Reader thread(s) still busy after process exit (pid=607, stuck_count=1)
  — killing process group, then waiting 30s for natural drain
[Subprocess] Reader thread(s) still stuck after 30s post-kill grace
  — force-closing pipes; some buffered data may be lost (pid=607, stuck_count=1)
[Subprocess] 1 reader thread(s) leaked for pid=607 after force-close; continuing anyway
[Headless Task] Error reading stdout: I/O operation on closed file.
[Headless Task] Execution completed without a result message after 0 tool calls / 22 turns
  (raw_messages=40). Likely cause: a tool or child subprocess inherited stdout and prevented
  the claude reader thread from capturing the final result block.

Backend logs:

[TaskExecService] Agent <name> responded: HTTP 502 (1648586ms)
[TaskExecService] Failed to execute task on <name>: Execution completed without a result
  message after 0 tool calls / 22 turns (raw_messages=40). Likely cause: a tool or child
  subprocess inherited stdout...

Circuit breaker fires continuously from ~1 minute into execution until the 502 is returned (~27 minutes), indicating the agent-server event loop is saturated during the execution.

Location

  • File: docker/base-image/agent_server/utils/subprocess_pgroup.py
  • Function: drain_reader_threads
  • File: docker/base-image/agent_server/services/claude_code.py
  • Function: execute_headless_task

Root Cause

Process group topology

When Claude Code starts, it launches stdio MCP servers defined in .mcp.json as child processes. For npx-based servers:

agent-server
  └── claude (pgid=X)
        ├── [claude internals]
        └── npm exec --yes -- <mcp-package>   ← spawned by Claude Code's MCP manager
              └── node <mcp-server-entry>      ← spawned by npm, may be a new pgid

The npm → node chain typically starts in its own session/process group (npm calls setsid() or equivalent). Claude Code's MCP subprocess management may not set the child's process group explicitly either. Result: the MCP server process tree is outside claude's pgid.

What #531 fixed vs. what it didn't

#531 correctly reordered drain_reader_threads to:

  1. Kill claude's process group
  2. Wait post_kill_grace=30s for natural drain
  3. Force-close only as last resort

This fixes the original "buffer discarded before reader could drain it" scenario. But in this case, terminate_process_group(pgid=claude_pgid) only kills processes within claude's pgid. The npm/node MCP server processes — in a different pgid — survive, still holding an open write FD on claude's stdout pipe. The kernel cannot EOF the read end while any writer FD remains open. The 30s grace period expires without ever receiving EOF, so the reader thread stays blocked, the force-close path runs, and the result line (already in the buffer or waiting to be written) is lost.

Why 0 tool calls / 22 turns

The tool call count is zero because tool results never arrived: the MCP server processes may be slow to initialize (npm install on first run) or the agents actively invoke MCP tools whose responses never complete because the servers are hung. Claude runs through its 22 turns producing text but never receives tool results, then exits with return_code=0. The MCP server orphans are still running with the write FD open, so the reader blocks until forced.

Reproduction Steps

  1. Configure an agent with one or more stdio MCP servers using npx -y <package> in .mcp.json
  2. Schedule a headless task on that agent that invokes those MCP tools
  3. Observe the task run: it will complete claude's execution (return code 0) but fail with HTTP 502
  4. Agent-server logs will show: "Reader thread(s) still stuck after 30s post-kill grace"
  5. Backend logs will show: "Execution completed without a result message after 0 tool calls"

Contrast: a different schedule on the same agent that does NOT invoke the npx MCP tools will succeed (work-loop vs. heartbeat pattern observed in production).

Suggested Fix

Two complementary approaches:

Option A — Kill orphan pipe-writers before waiting (Linux /proc scan)

After terminate_process_group(), scan /proc/*/fd for any process (not in our pgid) that has an open write FD pointing at the same pipe inode as claude's stdout, then SIGKILL those processes before starting the post_kill_grace wait:

import os, stat

def _kill_pipe_writers(pipe_write_fd: int, our_pgid: int) -> None:
    """Kill any process outside our_pgid that holds pipe_write_fd's inode open for writing."""
    try:
        target_ino = os.fstat(pipe_write_fd).st_ino
    except OSError:
        return
    for pid_str in os.listdir("/proc"):
        if not pid_str.isdigit():
            continue
        pid = int(pid_str)
        try:
            if os.getpgid(pid) == our_pgid:
                continue  # already killed above
        except OSError:
            continue
        fd_dir = f"/proc/{pid}/fd"
        try:
            for fd_name in os.listdir(fd_dir):
                fd_path = os.path.join(fd_dir, fd_name)
                try:
                    st = os.stat(fd_path)
                    if stat.S_ISFIFO(st.st_mode) and st.st_ino == target_ino:
                        # Check it's writable (mode bit O_WRONLY or O_RDWR)
                        fdinfo = open(f"/proc/{pid}/fdinfo/{fd_name}").read()
                        if "flags:" in fdinfo:
                            flags = int(fdinfo.split("flags:")[1].split()[0], 8)
                            if flags & os.O_WRONLY or flags & os.O_RDWR:
                                os.kill(pid, signal.SIGKILL)
                                break
                except OSError:
                    pass
        except OSError:
            pass

Call this after terminate_process_group() and before the post_kill_grace join loop in drain_reader_threads.

Option B — Prevent FD inheritance when Claude Code spawns MCP servers (preferred, upstream fix)

The cleaner fix is to ensure MCP server subprocesses never inherit the pipe FDs in the first place. Claude Code's MCP subprocess spawning should use close_fds=True (Python default on 3.2+, but subprocess spawning via shell may override) or explicitly set stdin=DEVNULL, stdout=PIPE, stderr=PIPE for MCP server processes so they get their own stdio.

This may need to be fixed in Claude Code's MCP client implementation rather than in Trinity's agent-server.

Option C — Track MCP server PIDs and kill them explicitly

If the agent-server starts MCP servers itself (or can enumerate them from Claude Code's process tree), record those PIDs at task start and add them to the kill list in drain_reader_threads.

Observed Behavior vs. Expected

Observed Expected
After #531 Reader thread blocked for full 30s grace, then force-closed → data lost Grandchild termination → EOF → reader drains → clean exit
npx MCP servers Survive terminate_process_group, hold pipe open indefinitely Should be killed or should not inherit the write FD
Task result HTTP 502, execution marked failed HTTP 200, execution marked success with result

Environment

Related

Secondary Finding (non-blocking)

Several agent skills have allowed_tools formatted as a comma-separated string instead of a YAML list, causing pydantic parse warnings in the updated agent-server:

Failed to parse skill at .../SKILL.md: 1 validation error for SkillInfo
allowed_tools
  Input should be a valid list [type=list_type, input_value='Bash, Read', input_type=str]

This appears to be a skill format validation tightening introduced in the v0.5.0 release. Skills with this format fail to load and their allowed_tools restrictions are not enforced. Tracked separately — skills still execute, the restriction is just not applied.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions