Skip to content

StdioServerParameters exposes no preexec_fn / rlimit / process-group hook for spawned MCP server subprocess #3457

Description

@dahai80

Context

StdioServerParameters (and the stdio_client context manager that spawns the server subprocess) expose no hook to control the spawned child process's resource limits or process-group assignment at fork time.

Problem

A client that spawns an untrusted or third-party MCP server via stdio cannot, from the SDK public API:

  • set CPU / memory / FD / address-space rlimits on the child,
  • place the child in its own process group (so a runaway child can be killed as a group without orphaning grandchildren),
  • set a preexec_fn (POSIX) or equivalent to run arbitrary setup between fork and exec.

StdioServerParameters signature (command, args, env, cwd, encoding, encoding_error_handler) has no slot for any of these. stdio_client owns the subprocess spawn internally, so a caller cannot inject a custom Popen either.

Impact

Downstream hosts (e.g. inference servers spawning MCP tool servers) cannot enforce hard resource caps or reliable teardown on a hung/misbehaving MCP server subprocess from the client side. The only mitigation available today is a bounded connect timeout around __aenter__, which does not cover a server that accepts the connection then later runs away.

Request

Expose at least one of:

  1. an optional preexec_fn / process_group / rlimit-style kwarg on StdioServerParameters (or stdio_client), passed through to the underlying subprocess.Popen, or
  2. an injection point for a custom Popen factory / spawn callable.

(1) mirrors subprocess.Popen(..., preexec_fn=..., start_new_session=...) and would let hosts enforce resource limits + process-group isolation without forking the SDK.

Environment: mcp python-sdk, macOS / Linux. Filed from fusion-mlx (local MLX inference host) where we need to cap spawned MCP server subprocesses.

Activity

  1. connerkup commented on Sep 6, 2026

    @connerkup

    Confirming this requirement from host process supervision and sandboxing in production environments. When orchestrating untrusted or long-running MCP tool servers via stdio, relying solely on high-level timeouts fails when a child process hangs in compute, allocates unbounded memory, or spawns detached worker subprocesses. Setting POSIX resource.setrlimit via preexec_fn and managing signal dissemination via process_group / process group isolation are standard patterns to guarantee deterministic host bounds.

    We've prepared and verified an implementation addressing option (1):

    • Extends StdioServerParameters with optional preexec_fn: Callable[[], Any] | None = None and process_group: int | None = None (allowing arbitrary types in Pydantic config).
    • Plumbs these parameters through stdio_client into anyio.open_process (with POSIX-safe handling).
    • Includes comprehensive unit tests in tests/client/test_stdio.py verifying child hook execution and process group propagation, maintaining 100.00% statement test coverage under strict coverage gates.

    Branch and clean patch are ready on connerkup:fix/3457-stdio-server-parameters-preexec-rlimit (previously drafted in PR #3461). If maintainers are open to this addition and can assign this issue, GitHub will automatically reopen PR #3461 for review.

  2. dahai80 commented on Sep 7, 2026

    @dahai80
    Author

    Thanks @connerkup — I reviewed the patch on connerkup:fix/3457-stdio-server-parameters-preexec-rlimit against current main and it looks clean and correctly scoped:

    • StdioServerParameters gains preexec_fn: Callable[[], Any] | None and process_group: int | None, with model_config = ConfigDict(arbitrary_types_allowed=True) so the callable type validates — minimal, additive, no breaking change to existing callers (both default to None).
    • stdio_client forwards them via **extra_spawn_kwargs into _create_platform_compatible_process, and from there into anyio.open_process only on the POSIX branch. The Windows path (create_windows_process) is untouched, which is correct — preexec_fn/process_group are POSIX-only concepts.
    • start_new_session=True is preserved as the default, and flips to False only when a custom process_group is supplied, avoiding the OS-level conflict of passing both. The TypeError fallback to anyio.lowlevel.get_async_backend().open_process handles older anyio versions that reject unknown kwargs.
    • The four unit tests (test_stdio_server_parameters_preexec_and_process_group, test_create_platform_compatible_process_forwards_preexec_and_process_group, test_preexec_fn_executes_in_child_process, test_process_group_sets_child_process_group) cover model init, kwarg forwarding, and the actual POSIX child-process semantics — the last two being the load-bearing ones, asserting the hook really runs in the child and the pgid really changes.

    This matches option (1) from the issue exactly and is the approach I would take myself. It covers the two concrete needs behind the original report from fusion-mlx: (a) resource.setrlimit caps via preexec_fn, and (b) killing a runaway server as a process group via process_group without orphaning grandchildren — neither of which the high-level connect timeout covers.

    To the maintainers: this is a well-scoped, non-AI-generated, tested contribution that addresses a real production-sandboxing gap. Assigning @connerkup to this issue would let PR #3461 reopen for review. I am the issue author and would rather see this one contributor-driven fix land than file a competing PR.

  3. memtomem commented on Sep 12, 2026

    @memtomem

    A second, distinct use case for the same seam, in case it helps scope the API: teardown reaping in an MCP proxy.

    memtomem-stm proxies several stdio upstreams through stdio_client. On shutdown it has to tell "child this process spawned via stdio_client" apart from "child the embedding host spawned", keyed by something that survives pid reuse. Today the only thing it can enumerate is pids (pgrep -P), which is exactly the identity that is not stable — and when stdio_client's own teardown gives up (the "still alive after the kill escalation; abandoning it" branch in stdio.py), that abandoned child is handed to the caller as a bare pid. The bookkeeping this forces on the caller is written up in memtomem/memtomem-stm#1009.

    Option (1) as implemented in the closed PRs (preexec_fn, process_group) helps with kill-as-group, but it yields no handle, so it does not cover this case. Option (2) does: either a spawn callable that returns the Popen, or the process exposed on the yielded value / context object. Would (2) be acceptable alongside (1)? Happy to help with a PR if a maintainer wants to assign it.

  4. memtomem commented on Sep 12, 2026

    @memtomem

    Correction to my earlier comment: I wrote that the SDK "hands the caller a bare pid" when it abandons a child. That is wrong, and the real situation is worse for this use case rather than better.

    stdio_client yields only read_stream, write_stream (mcp/client/stdio.py:204 in 2.0.0). The abandon branch at :261 writes the pid into a logger.warning and nothing else. So the caller receives neither a handle nor a pid, programmatically — a client that wants to reap an abandoned child has to rediscover it by enumerating its own direct children, which is the pid-shaped identity this request is about.

    The rest of the comment stands: option (1) gives kill-as-group but no ownership identity, and option (2) is what the reaping case needs.

  5. added
    enhancementRequest for a new feature that's not currently supported
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementRequest for a new feature that's not currently supported

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions