mcp-node-zig
mcp-node-zig
Give your AI agent a shell on any machine. One binary, no runtime, no SSH.

1.9 ms cold start · 0.62 MiB idle RSS · 8.70 MiB on disk · 0.7 ms p50 exec round-trip (benchmarks)
mcp-node-zig is a remote execution node that speaks MCP: put one static binary on a machine and your agent can run commands, drive long-running sessions and read or write files there. Since v0.2.0 the node can dial out to a hub instead of listening, so a box behind NAT or someone else's firewall is reachable with no inbound ports and no SSH. It answers its first MCP request 1.9 ms after start and idles at 0.62 MiB, which makes leaving one on every machine basically free.
Install
curl -fsSL https://raw.githubusercontent.com/alexchen-sys/mcp-node-zig/main/install.sh | shThe script detects the platform, downloads the latest release, verifies it against SHA256SUMS.txt and installs to /usr/local/bin (or ~/.local/bin). Pin a version with MCP_NODE_VERSION=0.2.0, change the target directory with PREFIX=/some/dir.
Manual install, Linux x86_64:
curl -L https://github.com/alexchen-sys/mcp-node-zig/releases/download/v0.2.1/mcp-node-v0.2.1-x86_64-linux.tar.gz | tar xzOther targets on the releases page: aarch64-linux, aarch64-macos, x86_64-windows. Each release ships SHA256SUMS.txt.
Related MCP server: interminal
Quickstart
openssl rand -hex 32 > token
./mcp-node-v0.2.1-x86_64-linux/mcp-node &
curl -sS http://127.0.0.1:8341/mcp \
-H 'Content-Type: application/json' \
-H "X-Node-Token: $(cat token)" \
--data '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"sys_info","arguments":{}}}'You get back hostname, OS, load, memory and uptime as JSON. Now attach your agent:
claude mcp add --transport http mcp-node http://127.0.0.1:8341/mcp \
--header "X-Node-Token: $(cat token)"curl.exe -L -o mcp-node.zip https://github.com/alexchen-sys/mcp-node-zig/releases/download/v0.2.1/mcp-node-v0.2.1-x86_64-windows.zip
Expand-Archive mcp-node.zip
[guid]::NewGuid().ToString('N') + [guid]::NewGuid().ToString('N') | Set-Content -NoNewline -Encoding Ascii token
.\mcp-node-v0.2.1-x86_64-windows\mcp-node.exeFrom a second window:
Invoke-RestMethod -Uri http://127.0.0.1:8341/mcp -Method Post -ContentType "application/json" `
-Headers @{ "X-Node-Token" = (Get-Content token) } `
-Body '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"sys_info","arguments":{}}}'Why
Nothing to provision. No Node, no Python, no OpenSSH, no keys to distribute. Copy one file and run it.
Processes outlive requests. Start a build, disconnect, come back and read the output by offset.
Cheap to keep running. Under 1 MiB idle. Each extra connected agent adds about 288 KiB, not a second ~190 MiB server process.
No shell unless you ask.
execpasses argv verbatim.exec_shellis the one explicit shell layer.Kills the whole tree. Process groups on POSIX, Job Objects on Windows. No orphaned children holding pipes.
Linux, macOS, Windows. All three are built and smoke-tested in CI on every push.
How is this different from SSH?
They solve different problems, and they work well together.
SSH is a secure transport and a terminal for people. It has encryption, key management, agent forwarding, PTYs, scp/sftp/rsync and decades of hardening. Keep using it for all of that.
mcp-node is an execution API for agents. Through SSH, an agent gets one string that a remote shell parses again, plus raw bytes and an exit status. Here it gets:
argv without a shell. No second round of quoting for a remote shell to reparse. A shell is used only when you ask for
exec_shell.Structured results.
stdout,stderr,exit_code,truncated_*and timings as JSON, with per-call timeouts.Sessions that outlive the request.
exec_start, thenexec_pollby byte offset,exec_writeto stdin,exec_waitup to 300 s,exec_killfor the whole tree. A dropped connection doesn't kill the build.Verified file writes.
write_filetakes base64 and returns the SHA-256 of what landed on disk.A self-describing interface. Any MCP client discovers the tools from
tools/list; nothing to teach the agent.
Speed isn't the argument: on an open connection both are fast. mcp-node has no encryption of its own, so for remote machines the usual setup is both together: the node listens on 127.0.0.1 and you reach it through an SSH tunnel (ssh -L 8341:127.0.0.1:8341 host), a VPN or a TLS reverse proxy.
Reverse connect (no inbound ports)
For machines behind NAT or without sshd, run the node with
--connect hub:port: it dials out to a hub (the same binary with
MCP_NODE_HUB_LISTEN), and clients reach it at /n/<name>/mcp on the hub.
See docs/reverse-connect.md.
stdio mode
Clients that spawn their servers as subprocesses can run the node directly
with --stdio (or MCP_NODE_STDIO=1). Messages are newline-delimited
JSON-RPC: one request per line on stdin, one response per line on stdout.
There is no listener and no token, since whoever holds the pipes started the
process. Logs go to stderr only; the node exits 0 when stdin closes and kills
any sessions still running.
{
"mcpServers": {
"mcp-node": {
"command": "mcp-node",
"args": ["--stdio"]
}
}
}--stdio can't be combined with --connect or MCP_NODE_HUB_LISTEN.
Tools
Tool | Does |
| hostname, OS, load, memory, uptime |
| run argv to completion, no shell |
| run a script via |
| start a long-running process as a session |
| read session output from byte offsets |
| block until the session exits or |
| write base64 bytes to stdin; |
| kill the session's process tree |
| free the session; idempotent |
| list live sessions |
| read UTF-8 text with offset/limit |
| write base64 content, create parent dirs, return SHA-256 |
| list entries with type, size, mtime |
Full schemas via tools/list. exec and exec_shell timeouts default to 120 s, clamped to 1–1800 s.
Clients
Every client below uses the same URL and token. Authorization: Bearer <token> works everywhere; X-Node-Token: <token> is accepted for clients that can't set Authorization.
Cursor (.cursor/mcp.json):
{
"mcpServers": {
"mcp-node": {
"url": "http://127.0.0.1:8341/mcp",
"headers": { "Authorization": "Bearer <token>" }
}
}
}VS Code (.vscode/mcp.json):
{
"servers": {
"mcp-node": {
"type": "http",
"url": "http://127.0.0.1:8341/mcp",
"headers": { "Authorization": "Bearer <token>" }
}
}
}Codex (~/.codex/config.toml):
[mcp_servers.mcp-node]
url = "http://127.0.0.1:8341/mcp"
http_headers = { "Authorization" = "Bearer <token>" }Claude Desktop and other JSON clients: same as Cursor, plus "type": "http".
Clients that only speak stdio: run the node itself with --stdio (see stdio mode), or bridge to a remote node over HTTP with npx mcp-remote http://127.0.0.1:8341/mcp --header "Authorization: Bearer <token>"
Security
Token required. The server refuses to start with a missing or empty token file. Comparison is constant-time.
Loopback by default. It binds
127.0.0.1. To expose it, setMCP_NODE_HOSTand add thehost:portclients use toMCP_NODE_ALLOWED_HOSTS.Host and Origin checked before the body is parsed: unknown Host gets 421, unknown Origin gets 403.
No built-in TLS. Put it behind a reverse proxy, tunnel or VPN, for example
caddy reverse-proxy --from node.example.com --to 127.0.0.1:8341.The token is a shell as the user the node runs as. Run it under an account scoped to what the agent should touch.
Configuration
Environment variables only.
Variable | Default |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| unset; |
| unset; |
mcp-node --version prints the version, --help the flags; any other
argument than these, --connect and --stdio is an error.
Reverse-connect variables (MCP_NODE_CONNECT*, MCP_NODE_HUB_*) are listed in
docs/reverse-connect.md.
Troubleshooting
Symptom | Fix |
401 | Wrong or missing token. If you send both headers, they must match. |
421 | Add the |
403 | Browser |
415 | Send |
Connection refused | Check |
Exits at startup | Create a non-empty token file. |
Limitations
Plain HTTP only, no PTY (interactive TUIs won't work), no audit log yet. All three are on the roadmap. The API is 0.1.x and may change.
Build from source
Requires Zig 0.16.
git clone https://github.com/alexchen-sys/mcp-node-zig
cd mcp-node-zig
zig build test
zig build -Doptimize=ReleaseSafe # → zig-out/bin/mcp-nodeOn Linux the binary is static with no libc. macOS links only libSystem; Windows only kernel32/ntdll.
License
MIT
Available Tools
13 toolsexecRun Command (no shell)ADestructive
Run argv to exit, no shell: metacharacters are data. Shell syntax: exec_shell. Long-running or interactive: exec_start. timeout in seconds, default 120, clamped to [1, 1800]. Invalid UTF-8 in stdout/stderr becomes U+FFFD.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Child working directory; omit to inherit. | |
| argv | Yes | Program and arguments, e.g. ["df", "-h", "/"]. Passed verbatim. | |
| timeout | No | Kill after this many seconds (1..1800, default 120). |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No | |
| stderr | No | |
| stdout | No | |
| timeout | No | |
| exit_code | No | |
| duration_ms | No | |
| duration_us | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint/openWorldHint/idempotentHint=false, so the safety profile is covered. The description goes further with non-obvious behavior: no shell interpretation (metacharacters are data), default timeout 120 clamped to [1,1800], and lossy UTF-8 decoding to U+FFFD. It does not describe output shape, but an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Telegraphic but zero-waste: the core contract is front-loaded in the first clause, then alternatives, then timeout semantics, then an edge case. Every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description covers invocation semantics, alternatives, timeout behavior, and encoding edge cases. Only blocking/streaming behavior is left implicit, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds real meaning beyond the schema by stating that argv is passed verbatim with metacharacters as data and reiterating the timeout default/clamp, which clarifies the semantic contract of the required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific operation (run argv to exit, no shell) and its defining property (metacharacters are data), which instantly separates it from exec_shell. The sibling routing names make the distinction explicit rather than inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not and alternatives: shell syntax goes to exec_shell, long-running or interactive goes to exec_start. Both branches an agent could misroute are closed off.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_closeClose SessionADestructiveIdempotent
Kill if running, join threads, free the session. Idempotent: repeat calls return ok, already_closed=true. Close every session you start; sessions persist across requests.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Positive id from exec_start. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No | |
| already_closed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint and idempotentHint, so the bar is lower, yet the description adds real value: it spells out the idempotent repeat-call contract (ok, already_closed=true) and the cross-request persistence trait of sessions. It does not mention permissions or error conditions beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and followed by the constraints. Every clause carries information an agent needs; no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lifecycle tool with an output schema and full annotations, the description covers the key behavioral facts (idempotency contract, session persistence). The only gap is routing relative to exec_kill, which is minor given the output schema handles return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter and 100% schema description coverage ('Positive id from exec_start'), the schema already carries the semantics. The description adds nothing about session_id, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete multi-step operation ('kill if running, join threads, free the session') on a specific resource, which is far more informative than the title alone. It does not explicitly differentiate itself from the sibling exec_kill, which appears to overlap on the 'kill' behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Close every session you start' gives a clear standing directive for when to call it, and 'sessions persist across requests' explains why closure matters. It stops short of naming exec_kill or exec_list as alternatives or stating when NOT to close.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_killKill SessionADestructive
Kill the session's process tree (POSIX: SIGKILL to the process group; Windows: job object). Safe on a dead session. Output stays pollable; free with exec_close.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Positive id from exec_start. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations' destructiveHint, it discloses platform-specific kill semantics (process group SIGKILL vs. Windows job object), that calling on an already-dead session is harmless, and that output remains pollable afterwards. These are behavioral facts the structured fields do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then safety, then cleanup routing. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained; the description still notes output stays pollable and points to exec_close for cleanup. For a one-parameter destructive tool this covers everything an agent needs before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single session_id parameter is documented as a positive id from exec_start. The description adds no further parameter detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (kill) and resource (the session's process tree), with the exact mechanism per platform (POSIX SIGKILL to the process group; Windows job object). It also implicitly separates itself from the sibling exec_close by assigning cleanup of resources to that tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context ("Safe on a dead session") and names the companion tool for freeing resources (exec_close), which routes the agent reasonably well. It stops short of an explicit when-to-use-this-vs-alternatives statement, e.g. kill vs. close vs. wait ordering.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_listList SessionsARead-onlyIdempotent
List live sessions: session_id, pid, argv, done, exit_code, timestamps. No arguments. Recovers ids after a lost reply.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No | |
| sessions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds the recovery use case but its field enumeration largely duplicates the existing output schema, and it says nothing about session lifetime, ordering, or limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse clauses, front-loaded with the verb and resource, with no wasted prose. The only minor redundancy is re-listing fields that the output schema already defines.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the return values need not be explained, annotations carry the safety profile, and there are no parameters, so the description is nearly complete. The remaining gap is minimal: no mention of empty-result or error behavior for a session list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters and schema coverage is 100%, so the required baseline is 4. 'No arguments' correctly confirms the empty-input contract and there is nothing further for the description to disambiguate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List live sessions') and enumerates the returned fields (session_id, pid, argv, done, exit_code, timestamps), which clearly separates it from the action-oriented siblings like exec_poll, exec_wait, and exec_kill. It stops short of naming an alternative sibling, so it is clear but not explicitly differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Recovers ids after a lost reply' gives a concrete when-to-use scenario, which is the key selection cue for a listing tool. It does not state exclusions or name named alternatives, so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_pollPoll SessionARead-onlyIdempotent
Read new session output from byte offsets: stdout/stderr deltas, done, exit_code, truncation flags. Pass back the offsets from the previous reply to resume. Same payload as exec_wait.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Positive id from exec_start. | |
| stderr_offset | No | Stderr byte offset to resume from, as returned by the previous poll (default 0). | |
| stdout_offset | No | Stdout byte offset to resume from, as returned by the previous poll (default 0). |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| done | No | |
| error | No | |
| stderr | No | |
| stdout | No | |
| exit_code | No | |
| duration_ms | No | |
| duration_us | No | |
| stderr_offset | No | |
| stdout_offset | No | |
| truncated_stderr | No | |
| truncated_stdout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safe read-only, idempotent profile, and the description is consistent with them (offset-based resume supports idempotency). It adds behavior the annotations cannot convey: that output is delivered as deltas from byte offsets, that truncation flags exist, and that offsets must be carried forward between calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with zero filler; the read/delta semantics come first and the resumption instruction second, which matches how an agent needs the information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and the description plus schema cover the polling loop adequately. The only gap is the missing routing decision against exec_wait, which is more a usage-guidelines issue than a completeness one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: each parameter is documented, including offset defaults and origin from previous polls. The description reinforces the resume pattern but adds no syntax, format, or edge-case detail (e.g. behavior when offsets are stale or beyond the current buffer) beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read new session output from byte offsets') and enumerates the payload contents (stdout/stderr deltas, done, exit_code, truncation flags), which an agent can distinguish from write/kill siblings. The relationship to exec_wait is only hinted at via 'Same payload as exec_wait', so it stops short of a clear sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Pass back the offsets from the previous reply to resume' gives real procedural guidance for repeated polling, but the description never states when to choose this over exec_wait (the obvious alternative in the sibling list) or when polling is unnecessary. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_shellRun Shell ScriptADestructive
Run one script through one shell (POSIX: bash/sh/fish/zsh -c; Windows: cmd /c, powershell -c). Only for shell syntax: pipes, redirects, &&, globs; otherwise exec. timeout in seconds, default 120, clamped to [1, 1800]. Invalid UTF-8 in stdout/stderr becomes U+FFFD.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Child working directory; omit to inherit. | |
| shell | No | POSIX: bash (default), sh, fish, zsh. Windows: cmd (default), powershell. | |
| script | Yes | Script passed to the shell, e.g. "du -sh /var/log | sort -h". | |
| timeout | No | Kill after this many seconds (1..1800, default 120). |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No | |
| stderr | No | |
| stdout | No | |
| timeout | No | |
| exit_code | No | |
| duration_ms | No | |
| duration_us | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered. The description adds real behavior beyond that: the timeout is defaulted at 120s and clamped to [1, 1800], and invalid UTF-8 in stdout/stderr is silently converted to U+FFFD — a data-mangling detail an agent could not learn elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the verb and routing rule before the mechanics. No filler; the telegraphic fragments (timeout semantics, encoding note) each carry distinct, actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations carry the destructive/open-world safety signal. Combined with shell selection, timeout bounds, cwd inheritance (via schema), and encoding behavior, nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including shell defaults and the timeout range. The description largely restates those values; its only marginal addition is showing how the script string is handed to the shell (-c / /c), which is the baseline 3 case when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Run one script through one shell') and immediately pins down the mechanism per platform (POSIX -c vs Windows /c / -c). It also names the sibling it is not for the non-shell case ('otherwise exec'), so an agent can route between exec_shell and exec without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage to shell syntax (pipes, redirects, &&, globs) and names the alternative tool 'exec' for everything else — a genuine when/when-not pair. It does not, however, address the exec_start/exec_poll/exec_wait background family, so the agent must still infer whether long-running or interactive work belongs here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_startStart SessionADestructive
Start a process as a session; returns session_id at once. Use when it outlives the call, streams output or reads stdin (exec_write). Then exec_poll, exec_wait, exec_kill, exec_close. Sessions persist across requests; finished ones are reaped after the session TTL.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Child working directory; omit to inherit. | |
| argv | Yes | Program and arguments, e.g. ["python3", "-i"]. Passed verbatim. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| pid | No | |
| error | No | |
| session_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly=false, destructive=true, openWorld=true, idempotent=false, and the description adds behavior those annotations don't carry: asynchronous return of session_id, persistence across requests, and TTL-based reaping of finished sessions. It doesn't explain the destructive implications of spawning a process or any resource limits, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return value, then conditions, then lifecycle siblings, then persistence semantics. It is dense and telegraphic but essentially waste-free; the sibling enumeration sentence is the only part that reads as a list rather than prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape needn't be restated, and the description still covers the async contract, session lifetime, and next-step tools. The one gap is the meaning of the destructive annotation for a process-spawning tool, but otherwise an agent has what it needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two params, so the schema already documents 'cwd' (inheritance behavior) and 'argv' (verbatim passthrough with example). The description adds only indirect context via the stdin reference, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a process as a session') plus the immediate return ('returns session_id at once'). Together with the usage condition ('when it outlives the call'), an agent can distinguish this from the synchronous sibling 'exec' and from 'exec_shell' without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit selection criteria: use when the process outlives the call, streams output, or reads stdin, with the stdin counterpart named (exec_write). It also routes the follow-up lifecycle through exec_poll, exec_wait, exec_kill, exec_close, so the agent knows both when to enter and how to proceed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_waitWait For SessionARead-onlyIdempotent
Block until the session exits or timeout elapses (default 30s, max 300s). Returns the exec_poll payload with deltas from the given offsets.
| Name | Required | Description | Default |
|---|---|---|---|
| timeout | No | Max seconds to wait (1..300, default 30). | |
| session_id | Yes | Positive id from exec_start. | |
| stderr_offset | No | Stderr byte offset to resume from (default 0). | |
| stdout_offset | No | Stdout byte offset to resume from (default 0). |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| done | No | |
| error | No | |
| stderr | No | |
| stdout | No | |
| exit_code | No | |
| duration_ms | No | |
| duration_us | No | |
| stderr_offset | No | |
| stdout_offset | No | |
| truncated_stderr | No | |
| truncated_stdout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, and non-destructive traits, so the description adds real value by disclosing blocking semantics, timeout bounds, and that the response is an exec_poll payload containing deltas relative to the supplied offsets. It does not describe behavior on session-not-found or timeout expiry, but the core blocking behavior is well conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the blocking contract and timeout bounds front-loaded and the return shape following. Nothing redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not detail return values, and it correctly focuses on blocking behavior and timing. The only omission is guidance on which sibling to choose for non-blocking reads, a minor gap for a well-annotated read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, including defaults and ranges. The description adds only a mild framing benefit by explaining that offsets produce deltas rather than absolute output, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: blocks until the session exits or the timeout elapses, and frames the return as the exec_poll payload with offset deltas. This clearly separates it from siblings like exec_poll, though the contrast is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The blocking-versus-timeout framing implies when to use it (wait for completion), and referencing exec_poll's payload hints at the relationship to the polling sibling. However, it never says when to prefer exec_poll, exec_kill, or exec_close instead, leaving the exclusion to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_writeWrite To SessionA
Write bytes to session stdin; eof=true closes stdin afterwards. For interactive programs started by exec_start.
| Name | Required | Description | Default |
|---|---|---|---|
| eof | No | Close stdin after writing (default false). | |
| data_b64 | Yes | Standard base64 of bytes for stdin; may be empty. | |
| session_id | Yes | Positive id from exec_start. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| eof | No | |
| bytes | No | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the mutation profile is covered. The description adds the eof side effect (closing stdin) and the interactive-program scope, but says nothing about blocking behavior when the stdin buffer is full or what happens if the session has already exited.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the action and immediately followed by the scoping condition. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers purpose, target session type, and the eof side effect. It leaves a few operational gaps (buffer-full/blocking behavior, session-state requirements) that an agent might want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters are documented there, including eof semantics and data_b64 format. The description restates the eof behavior but adds no encoding, size-limit, or session-state detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write bytes to session stdin') with a precise target ('session stdin'), and immediately ties the tool to sessions created by exec_start, distinguishing it from the sibling write_file which writes to files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'For interactive programs started by exec_start' gives clear context for when this tool applies, but it does not name an alternative or state when-not to use it (e.g., use exec_shell for one-shot commands).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dirList DirectoryARead-onlyIdempotent
List directory entries sorted by name: name, type (d dir, l symlink, f file), size in bytes, mtime in Unix seconds. Max 2000 per call; truncated=true if more.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Directory; "~" expands to home; default: working directory. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| path | No | |
| count | No | |
| error | No | |
| items | No | |
| has_more | No | |
| truncated | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description goes further by disclosing the 2000-entry cap and the truncation=true signal. That is a genuinely useful operational trait not present in annotations, though permission or path-error behavior is not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and packs output format, field list, and the truncation cap with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description helpfully summarizes them plus the truncation flag. It is nearly complete for a one-parameter read; only error handling for invalid paths is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter already documents the '~' expansion and default working directory, so the schema does the heavy lifting. The description adds no additional parameter meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List directory entries') plus the sort order, and describes the exact fields returned (name, type codes, size, mtime). This clearly separates it from siblings like read_file and the exec_* family without needing to read the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb and resource, but there is no explicit guidance on when to prefer this over read_file or exec_list, and no mention of error conditions such as a missing path. The agent can infer intent but is given no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileRead FileARead-onlyIdempotent
Read a text file as UTF-8 (invalid bytes replaced): size, offset, content, has_more. offset/limit count characters, not bytes. Default limit 200000; files over 64 MiB are rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path; "~" expands to home. | |
| limit | No | Max characters (default 200000). | |
| offset | No | Start character (default 0). |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| path | No | |
| size | No | |
| error | No | |
| offset | No | |
| content | No | |
| has_more | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive. Description adds encoding handling, size limit (64 MiB rejection), and offset/limit semantics—useful beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single efficient sentence, front-loaded with core operation and key constraints. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given annotations and output schema, description covers necessary behavioral details like encoding and character-based offset/limit. No major gaps for calling correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented. Description adds that offset/limit count characters not bytes, and the default limit, but mostly repeats schema info.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (text file) with the encoding and behavior (UTF-8, invalid bytes replaced). Clearly distinct from siblings like write_file and list_dir.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or alternatives mentioned. Implied read usage, but no guidance on choosing over exec or other file operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sys_infoSystem InfoARead-onlyIdempotent
Host summary: hostname, OS, load, memory, root disk (disk_root: total/used/free bytes), uptime. No arguments, read-only; use as a connectivity check.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| os | Yes | |
| mem | Yes | |
| node | Yes | |
| machine | Yes | |
| hostname | Yes | |
| disk_root | Yes | |
| uptime_raw | Yes | |
| loadavg_raw | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds useful context beyond that: it is argument-free and the natural probe for verifying host reachability, and it surfaces the disk unit (bytes) that callers would otherwise have to infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with the resource up front, the returned fields in the middle, and the primary use case last. Every clause earns its place; nothing is padded or repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only diagnostic tool with an output schema already present, the description supplies everything an agent needs: what it returns, that it takes no input, and the primary reason to call it. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The description does not invent parameter guidance, and 'no arguments' correctly signals that invocation needs no inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and enumerates exactly what the summary contains (hostname, OS, load, memory, root disk with units, uptime). It is trivially distinguishable from siblings like exec or read_file, which perform actions or fetch file contents rather than reporting host state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case — 'use as a connectivity check' — and notes no arguments are needed, so an agent knows when to reach for it. It does not name an alternative tool or state when not to use it, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_fileWrite FileADestructive
Write base64 content to a file, overwriting; creates parent dirs by default. Returns size and sha256 of the written bytes.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | POSIX permission bits for the new file, 0..0o7777 (default 0o644; ignored on Windows). | |
| path | Yes | File path; "~" expands to home. | |
| mkdirs | No | Create missing parent directories (default true). | |
| content_b64 | Yes | Standard base64 file content. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| path | No | |
| size | No | |
| error | No | |
| sha256 | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered structurally. The description earns credit by elaborating the actual destructive mechanic (overwrites existing files) and the mkdirs-by-default behavior, plus what the response contains, though it omits permission requirements and symlink handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with zero filler. The destructive write semantics are front-loaded in the first clause, and the return-value note is deferred to the end where it belongs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the tool needn't document return values in depth, and the description still notes size and sha256. Annotations cover the safety profile. The only real gap is the missing guidance on when to prefer exec_write over write_file.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (mode, path, mkdirs, content_b64) are already documented in the schema. The description reinforces base64 content and the parent-dir default but adds no format, encoding, or edge-case detail beyond what the schema provides, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb with a specific resource ("Write base64 content to a file") and discloses two consequential behaviors: overwriting and automatic parent-dir creation. It is clearly distinguishable from read_file/list_dir, but it does not explicitly differentiate itself from the sibling exec_write, which an agent could reasonably confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives. With both write_file and exec_write in the sibling list, the absence of any routing guidance or preconditions (e.g., requires an existing/local filesystem context) leaves selection to guesswork.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
exec - First observed
exec_close - First observed
exec_kill - First observed
exec_list - First observed
exec_poll - First observed
exec_shell - First observed
exec_start - First observed
exec_wait - First observed
exec_write - First observed
list_dir - First observed
read_file - First observed
sys_info - First observed
write_file
TDQS
Scored across 13 tools
The nine exec_* tools form a coherent session lifecycle (start/poll/wait/write/kill/close/list) with distinct roles, and the file tools (read_file/write_file/list_dir) are clearly separate. The main overlap is the trio exec vs exec_shell vs exec_start and the fact that exec_poll and exec_wait return identical payloads, which could cause occasional misselection.
All names are lowercase snake_case following a predictable verb_noun pattern. The exec_* family is systematically prefixed, and file tools (read_file/write_file/list_dir) and sys_info-adjacent tools follow consistent conventions throughout.
13 tools is well within the ideal range and each tool earns its place. The session lifecycle is broken into focused operations rather than being crammed into one overloaded tool, and the file/host operations round out the set without bloat.
The surface covers process execution (one-shot, shell, and session-based), full file read/write/list, and host info, which is solid lifecycle coverage. Minor gaps remain (no file delete/stat/move/copy), but these are workaroundable via exec/exec_shell.
Maintenance
Related MCP Connectors
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Open-source Zapier/n8n alternative as an MCP server: agents build, run and debug your workflows.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Zero-install remote MCP server for proof-of-existence file attestation.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceA secure and pluggable MCP server to run terminal commands on your local machine or cloud server — remotely, safely, and with LLMs or agentic clients.-
- AlicenseAqualityBmaintenanceMCP server for SSH and local terminal access. Supports interactive commands, long-running processes, and TUI apps like tmux/zellij63MIT
- FlicenseAqualityBmaintenanceCross-platform MCP server for policy-controlled command execution on Linux and Windows, with no SSH dependency.3-
- AlicenseNot gradedqualityAmaintenanceMCP server and CLI for host and container operations, enabling Docker and Compose control, SSH, host inspection, logs, ZFS, and safe file transfer. It exposes flux and scout MCP tools with parity from the original TypeScript server.2AGPL 3.0