Skip to main content
Glama

SSH MCP Server v2

NPM Version Downloads CI OpenSSF Scorecard OpenSSF Best Practices codecov License GitHub issues

SSH MCP Server is a security-first Model Context Protocol server that gives LLM agents controlled SSH access to remote hosts — with command classification, policy-based authorization, human-in-the-loop approval, and full audit logging.

The risk this server exists to manage. Giving an LLM shell access on a remote host puts private data, untrusted input and network egress in one place — Simon Willison's "lethal trifecta". Prompt injection has no general fix, so ssh-mcp assumes any command may be attacker-influenced: it classifies before executing, authorizes against a role × host-group matrix, gates destructive work behind approval, and records the decision either way. That narrows the blast radius; it does not remove the risk. Two things stay yours: never point it at a root account, and never set auto approval on a production profile. SECURITY.md has the full threat model.


Quick Start

1. Install

npm install -g ssh-mcp

2. Configure

Without a config the server still starts, so a client or directory can complete the MCP handshake and read tools/list — but every tool call is refused until you configure it, with a message naming the path below. Nothing runs on a host until this step is done.

Create the config file at the path for your platform:

Platform

Path

Linux

~/.config/ssh-mcp/config.toml (or $XDG_CONFIG_HOME/ssh-mcp/config.toml)

macOS

~/Library/Application Support/ssh-mcp/config.toml

Windows

%APPDATA%\ssh-mcp\config.toml

[defaults]
defaultProfile = "dev"
approvalMode = "ask-destructive"

[[profiles]]
name = "dev"
host = "192.168.1.100"
port = 22
user = "deploy"           # NOT root!
auth = "key"
keyRef = "~/.ssh/id_ed25519"
role = "admin"
approvalPolicy = "auto"    # dev is permissive
chmod 700 ~/.config/ssh-mcp && chmod 600 ~/.config/ssh-mcp/config.toml

The config decides which hosts, roles and policy rules this server honours, so it checks that nobody but you can read it — and treats the two platforms differently, because the question has a much clearer answer on one of them.

Linux and macOS: enforced. The mode check above, on the file and the directory — which is why chmod 700 is in that command, since mkdir -p under the default umask leaves the directory 0755. The server refuses to start otherwise. "Only the owner" is unambiguous here and chmod is a one-line fix.

Windows: split by what the ACL actually allows. There are no mode bits, so the ACL is read instead — and read exposure and write exposure are not treated alike, because Windows is much clearer about one of them than the other.

The ACL lets another account…

Default

only read the config

reported, and the server starts

change the config

refused

nothing (no ACL at all)

refused — that is full control for everyone

…and if the ACL could not be read

refused, except when icacls is absent or the check timed out

A config under %APPDATA% inherits access for you, SYSTEM and Administrators and needs nothing done to it. One created elsewhere does not: a file under C:\ inherits *read* for every local account and *modify* for every authenticated one. The message names the two icacls commands that fix it either way.

Read exposure is reported rather than refused because that is where Windows is genuinely muddier than POSIX, and refusing over it blocked a config at the documented location (#138). Write exposure is refused because it is not muddy at all: another account being able to rewrite the file that decides which hosts, roles and approval policy this server honours is an authorization bypass, not a disclosure.

Two flags move the whole thing: --strictConfigAcl refuses everything the check objects to, read-only grants included; --allowUncheckedConfigAcl reports everything and refuses nothing. Neither combination leaves you without an exit, which is the lesson of #138.

Exit statuses

Status

Meaning

0

Clean shutdown

1

A defect in the server — printed with a stack trace; please report it

2

How it was invoked or configured — printed as a message, no stack

A supervisor that treats any non-zero status as a failure needs no change. One that matched on 1 to detect a startup problem should match on 2 as well.

Starting with nothing configured is not an exit-2 condition, as of the release that added introspection without a config: the server starts so it can be described, and refuses each tool call instead. A supervisor that used a non-zero exit to catch an unconfigured deployment should watch for starting unconfigured on stderr, or read configured from GET /health when running the HTTP transport.

3. Set credentials via environment variables

export SSH_MCP_PASSWORD="your-password"        # if using auth=password
# OR use SSH agent (recommended):
export SSH_AUTH_SOCK="$SSH_AUTH_SOCK"           # already set if agent running

4. Connect from your MCP client

Claude Code:

claude mcp add --transport stdio ssh-mcp -- ssh-mcp

Claude Desktop / Cursor / Windsurf:

{
  "mcpServers": {
    "ssh-mcp": {
      "command": "ssh-mcp",
      "env": {
        "SSH_MCP_PASSWORD": "your-password"
      }
    }
  }
}

Never pass passwords as CLI arguments — they're visible via ps aux. Use env vars, config files, SSH agent, or OS keychain.


Related MCP server: SSH MCP Server

Tools (14)

Tool

Purpose

readOnly

destructive

list-connections

Discover available hosts and connection status

✅

—

list-sessions

List active sessions per host

✅

—

open-session

Create a named interactive (stateful) or background session

—

—

close-session

Close a session. A background session's command is signalled (INT/TERM/KILL) before its channel is dropped

—

✅

read-session-output

Read output from background sessions (e.g., tail -f)

✅

—

read-command

Execute allowlisted read-only commands (ls, cat, grep, ...)

✅

—

run-command

Execute arbitrary commands (destructive/privileged need approval, unless approvalPolicy = "auto")

—

—

privileged-command

Execute with sudo (needs approval, unless approvalPolicy = "auto")

—

✅

sftp-upload

Upload a file via SFTP (replaces the destination unconditionally)

—

✅

sftp-download

Download a file via SFTP

✅

—

sftp-list

List a remote directory, bounded in entries and bytes

✅

—

sftp-upload-file

Stream a local file to the remote host, never through model context

—

✅

sftp-download-file

Stream a remote file to local disk, never through model context

—

✅

signal-process

Send INT/TERM/KILL to a remote PID

—

✅

Streaming file transfer

sftp-upload/sftp-download move file contents through the model's context: the text is an argument on the way out and a response on the way back. That is what you want for a config snippet and exactly what you do not want for a 200 MB tarball or anything binary.

The two also differ on what they do to an existing destination, and the approved string now says which is which. sftp-upload replaces unconditionally and spells --overwrite every time, because that is what it always does; sftp-upload-file refuses unless you pass overwrite: true, and only then carries the flag. Since sftp-upload takes its content as an argument rather than naming a local file, its string also carries --bytes=<n> --sha256=<32 hex> — without that, two uploads to the same path are the same string, which means one approval covers both and an auditor cannot tell which set of bytes landed. The digest is 128 bits rather than a short prefix because that claim has to hold against a caller who picks both payloads, not only against an accidental repeat.

The path comes last in that string on purpose. [policy].denylist patterns are matched against the whole approved string, so a rule anchored on the path — authorized_keys$, the natural way to write "nothing may write here" — keeps working. A rule anchored on the whole string (^sftp:upload /root/.*$) does not: it was coupled to a format that is ours to change, and this release changes it. Anchor on the path segment instead.

sftp-upload-file and sftp-download-file stream between the remote host and local disk instead. Neither the bytes nor a base64 encoding of them ever reaches the model — the response is a byte count and two paths.

The two transfer tools are off until you configure defaults.transferRoot, and refuse with an explanation until then. (sftp-list has no local side and needs none of this.) That directory is the whole of their local reach:

[defaults]
transferRoot = "/srv/ssh-mcp-transfers"
transferMaxBytes = 268435456        # 256MB per transfer
transferTimeoutMs = 300000          # 5min with no bytes moving

Checked on every call, and refused rather than degraded if any of it fails: the directory must be 0700 and owned by the account running this server, with no group- or world-writable parent, and must not overlap the ssh-mcp installation, the config directory, the audit log directory, or ~/.ssh. Caller paths are confined to it, symlinks are refused rather than followed, and a download is staged as a sibling .part file and published atomically, so a failed transfer never leaves a half-written file at the destination.

transferTimeoutMs is an idle budget, not a total one: it bounds one metadata round-trip, or one stretch of the copy with no bytes moving, and is re-armed on progress. A slow but live transfer of a large file survives it; a stalled channel still fails within one window. That is why the byte cap above does not have to be divided by it — a total budget would have made a 256MB cap mean "only if the link sustains 900 KB/s".

Not available on Windows, where the transfer root cannot yet be verified private; the two transfer tools refuse there rather than writing into a directory other accounts may be able to read.

sftp-upload-file takes an optional mode (1–511; setuid, setgid and the sticky bit are refused). Omitted, a new remote file is published 0600; an overwritten one inherits the replaced file's permission bits — its permission bits only, not its setuid or setgid.

What a transfer is authorized for includes these arguments: the string the policy engine classifies, the approval prompt a human reads, and the audit record all name --overwrite and --mode when they are given. So approving one upload to a path does not approve a different one to the same path, and an approval grant (approvalGrantTtlMs) cannot be replayed with a different mode.

Interactive Sessions

Sessions maintain state (CWD, environment variables) between commands:

Agent: open-session(name="deploy", type="interactive")
Agent: run-command(session="deploy", command="cd /opt/myapp")
Agent: run-command(session="deploy", command="git pull")    # runs in /opt/myapp
Agent: run-command(session="deploy", command="npm ci")      # CWD persists
Agent: close-session(name="deploy")

Background Sessions

Long-running processes (logs, builds):

Agent: open-session(name="logs", type="background", command="tail -f /var/log/syslog")
Agent: read-session-output(name="logs", lines=20)   # poll
Agent: close-session(name="logs")

Remote host support

Tested against Linux (Debian/bash, Alpine/busybox ash), Dropbear, and Windows OpenSSH on Windows 11.

Linux / BSD / macOS

Windows OpenSSH

read-command, run-command, privileged-command, signal-process

✅

✅

sftp-upload, sftp-download

✅

✅

sftp-list

✅

✅

sftp-upload-file, sftp-download-file

✅

❌ (local side unverifiable)

Background sessions

✅

✅

Interactive sessions

✅

❌

Interactive sessions require a POSIX shell (sh, bash, ash, zsh). They work by bracketing each command with printf markers and reading $? and $PWD from a trailer — none of which exist in cmd.exe, the default shell for Windows OpenSSH. Opening one against such a host fails immediately with an explicit error rather than timing out; everything else works normally.

Setting PowerShell as the OpenSSH DefaultShell does not help: the protocol is POSIX-specific, not merely non-cmd.


Configuration

Profile options

[defaults]
defaultProfile = "dev"
sessionMaxPerConnection = 5
sessionIdleTimeoutMs = 600000       # 10min
sessionBackgroundMaxMs = 3600000    # 1hr
commandTimeoutMs = 60000
commandMaxChars = 5000              # 0 = unlimited, the config spelling of --maxChars=none
commandMaxOutputBytes = 1048576     # 1MB
connectionIdleReapMs = 900000       # 15min
commandQuotaPerDay = 0              # 0 = unlimited; circuit breaker for runaway agents
approvalGrantTtlMs = 0              # 0 = always prompt; see "Approval Grants"
approvalMode = "ask-destructive"    # auto | ask-destructive | ask-all | deny
# transferRoot = "/srv/ssh-mcp-transfers"   # enables the streaming file tools; see above
transferMaxBytes = 268435456        # 256MB per streaming transfer
transferTimeoutMs = 300000          # 5min with no bytes moving (idle, not total)

[[profiles]]
name = "prod-web-1"
host = "10.0.1.50"
port = 22
user = "deploy"
auth = "agent"                      # agent | key | password | keychain
keyRef = "~/.ssh/id_ed25519"        # for auth=key
keychainEntry = "ssh-mcp/prod"      # for auth=keychain (requires @napi-rs/keyring)
via = "bastion"                     # ProxyJump — route through bastion profile
group = "prod"                      # Policy tier: prod | staging | dev, or your own (see [policy])
workdir = "/var/www"
trustedHostKey = "SHA256:..."       # Pin host key (optional)
tty = false
role = "operator"                   # viewer | operator | admin
readOnly = false
approvalPolicy = "ask-all"
cert = false                        # SSH CA cert auth — auto-detects keyRef-cert.pub
announceAgent = true                # Send AI_AGENT=ssh-mcp to this host (see below); clear it for a host you do not control
sessionMaxPerConnection = 3         # per-profile override
sessionIdleTimeoutMs = 300000       # stricter for prod
commandQuotaPerDay = 200            # per-profile override
maxChars = 2000                     # per-profile override; stricter for prod
transferMaxBytes = 16777216         # per-profile override; transferRoot is not per-profile
transferTimeoutMs = 60000           # per-profile override

# Optional. Merged over the built-in role matrix; see "Policy Engine" below.
# roleBindings is keyed by role and then by tier, so the block below changes
# operator on prod and leaves operator's other tiers, and viewer and admin,
# on their defaults.
[policy]
denylist = ["^terraform\\s+destroy"]

[policy.roleBindings.operator]
prod = ["read-only", "safe", "destructive"]

Unknown sections and keys are a startup error, not a warning, so a typo cannot leave you running defaults you thought you had overridden. That extends to role and tier names: every one you write under [policy.roleBindings] has to be reachable by some profile, and every profile's role and tier has to resolve to real bindings. Both directions are checked at startup.

Host-side identification

Every command this server runs announces itself to the host as the environment variable AI_AGENT=ssh-mcp, so an operator can tell an agent's session from a person's from anything on the host that reads the session environment — a /etc/profile snippet or a ForceCommand wrapper, for instance. It covers all three channels a command can run on: one-shot commands, background sessions and interactive sessions. SFTP transfers do not carry it, because the SSH library this server uses offers no way to send an environment request on a subsystem channel.

The variable only lands in the session if the host opts in with AcceptEnv AI_AGENT in sshd_config. Without that line sshd ignores it and the session environment is unchanged — measured against Dropbear, which has no AcceptEnv mechanism at all, and against Windows OpenSSH on a default configuration: the command runs, the exit code is unchanged, the variable is simply absent.

The request is nevertheless sent to every host, opted in or not. AcceptEnv is the server's policy about what it stores; it is not a gate on what this client transmits. A host that has not opted in — including one that is hostile — can still see on the wire that an agent rather than a person is driving the session. That is why announceAgent = false exists on the profile: set it for any host you do not control and nothing is sent to it. It is on by default, because the operators who benefit are the ones running the hosts you do control.

No version is sent, deliberately. A version would narrow a hostile host to one specific build; the tool's name alone does not. (The SSH transport already discloses the library version in its identification string, which is one library's version across every program that uses it.)

Treat the variable as a courtesy label, not an attestation. Any SSH client can set it and any agent can omit it, so it is useful for attributing sessions you already trust and must not be used as an authorization or intrusion-detection input. Authoritative attribution belongs to the key or principal that authenticated.

ProxyJump (Bastion)

Reach internal hosts behind a bastion/jump server. The via field specifies a profile name to tunnel through:

[[profiles]]
name = "bastion"
host = "bastion.example.com"
user = "deploy"
auth = "agent"

[[profiles]]
name = "internal-db"
host = "10.0.1.50"                 # private IP — not directly reachable
user = "dbadmin"
auth = "key"
keyRef = "~/.ssh/db_key"
via = "bastion"                     # tunnel through bastion

No agent forwarding — only a TCP tunnel via forwardOut. The bastion stays connected and reusable for multiple internal hosts.

SSH CA Certificates

For enterprise setups with a central SSH Certificate Authority:

[[profiles]]
name = "prod-db"
host = "db.internal"
user = "admin"
auth = "key"
keyRef = "~/.ssh/id_ed25519"
cert = true                         # enable CA cert auth

The certificate file is auto-detected using OpenSSH convention (keyRef + -cert.pub, e.g. ~/.ssh/id_ed25519-cert.pub). You can override the path with SSH_MCP_<NAME>_CERT env var. The cert is concatenated with the private key per ssh2 convention.

Credential Resolution Order

  1. SSH agent (SSH_AUTH_SOCK) — no key material in process memory

  2. OS keychain (macOS Keychain / Windows Credential Manager / Linux Secret Service) — requires auth = "keychain" and @napi-rs/keyring

  3. Environment variables — SSH_MCP_PASSWORD, SSH_MCP_KEY, SSH_MCP_SUDO_PASSWORD, or profile-specific SSH_MCP_<NAME>_PASSWORD

  4. Key file — keyRef path or SSH_MCP_KEY env var

Never CLI arguments. v2 removes --password, --sudoPassword, --suPassword entirely.


Policy Engine

Roles

Role

Dev

Staging

Prod

viewer

read-only

read-only

read-only

operator

read-only, safe, destructive

read-only, safe, destructive

read-only, safe

admin

all

all

read-only, safe, destructive

Which column applies comes from the profile's group. Set it explicitly — without it the tier is guessed from the profile name (prod/staging/dev, local, test, sandbox), and an unrecognised name resolves to prod, the strictest tier. A production host named web-01 is therefore treated as production rather than silently getting dev permissions.

Note what this means for sudo: admin has no privileged on prod, so privileged-command is refused there by design — including on a quick-start profile, which has no name to infer from and therefore lands on prod. If the host is not production, say so:

npx ssh-mcp --host=10.0.0.5 --user=deploy --group=dev
[[profiles]]
name = "build-box"
group = "dev"

Configuring the matrix

The table above is the default, not a limit. An optional [policy] section is merged over it at startup, so granting sudo on a host you have honestly labelled prod is a reviewable line in a config file rather than a relabelling:

[policy.roleBindings.admin]
prod = ["read-only", "safe", "destructive", "privileged"]

The merge is at role and tier depth. That block changes admin on prod and nothing else: admin on staging and dev keep their defaults, and viewer and operator are untouched. Roles and tiers the defaults have never heard of are added rather than rejected, which is what makes a custom group resolve to real bindings instead of falling back to the strictest tier:

[[profiles]]
name = "build-box"
role = "admin"
group = "tier-1"

[policy.roleBindings.admin]
"tier-1" = ["read-only", "safe", "destructive"]

Extra deny patterns live in the same section, and are applied on top of the never-allowed list rather than replacing it:

[policy]
denylist = ["^terraform\\s+destroy"]

Because role and tier names are free strings, nothing in the merge itself can tell a new custom role from a misspelling of an existing one. A cross-check at startup does, and these all fail there rather than at the point of use:

  • a command class outside read-only | safe | destructive | privileged, so a priviledged typo cannot parse into a grant of nothing and then read as a policy decision when a command is refused;

  • any unrecognised section or key anywhere in the config, so a block the parser does not understand is an error rather than a clean startup with none of the behaviour you configured;

  • a role or tier under [policy.roleBindings] that no profile uses, so [policy.roleBindings.operater] cannot merge in as a fourth role while the profiles you meant to restrict keep running on defaults;

  • a profile whose role has no bindings, or whose tier has none under that role, so a host cannot end up on read-only for a reason nobody wrote down.

The last one covers the tier you did not set as well as the one you did. A profile with no group still resolves to one by name, and that inferred tier has to exist under the profile's role like any other.

A tier with no bindings for a role grants read-only, and never another tier's classes. There is no fallback between tiers: while the matrix was compiled in, falling back to prod meant falling back to a role's strictest cell, but a [policy] block can write that cell now.

An OPA sidecar is not an alternative route to the same grant. OPA is consulted only for commands the local policy already allows, so it can refuse more but never widen. Widening happens here or not at all.

Command Classification

Every command is classified before execution:

  • read-only: Allowlisted commands (ls, cat, grep, df, stat, systemctl status, ...), and only when every argument is provably data under a grammar declared for that binary — see below

  • safe: Non-destructive mutations (npm install, git pull, ...)

  • destructive: mutations that need approval (rm -rf /tmp/build, ...)

  • privileged: sudo, su, doas, pkexec

read-only is granted by argument grammar, not by binary name alone. An allowlisted binary's argument list is checked word by word against a grammar declared for it (src/policy/reader-grammar.ts); an option the grammar does not list, an abbreviation of one it does, a short-option cluster with an unlisted letter, an argument carrying an unquoted shell glob (*, ?, [), or an operand shaped like something the binary would write to, all fall the whole command to safe instead. This is what a readOnly profile is confined to, what viewer holds on prod and staging (on dev the viewer holds safe too), and what read-command requires — all of these refuse safe except the viewer on dev, so a command that used to run under the old, name-only rule can now be refused. The refusal names the rejected word (for example, `journalctl` is read-only only with the options and operands its grammar lists; `--foo` is not accepted there.), so the caller can see which spelling of the same command still works. See SECURITY.md for the residual risks this leaves open.

A separate forbidden list is never allowed, whatever the role or approval policy: rm -rf /, mkfs, dd of=/dev/, shutdown, curl|sh, fork bombs, writes to /etc/cron, /etc/systemd or authorized_keys, iptables -F, and recursive chmod 777 / / chown /. Add your own patterns via the policy denylist; an invalid pattern fails at startup rather than degrading silently.

Approval Modes

  • auto — no prompts (dev only!)

  • ask-destructive — prompt for destructive/privileged (default). Narrower than it sounds: outside the never-allowed list, destructive is one rm -rf /path pattern, find with a write/exec flag, an unresolvable command word, a program handed to an interpreter this server cannot read (python3 -c, perl -e, node -e, a program arriving on a pipe — but not awk, whose program is not read), and sftp-upload/interactive open-session — elevation classifies privileged, which also prompts. Ordinary writes, service control and signals do not. See SECURITY.md before relying on this in production.

  • ask-all — prompt for every command

  • deny — reject destructive/privileged commands outright (no prompt)

Approval Grants (just-in-time)

approvalGrantTtlMs lets one explicit approval cover repeats of the exact same command on the same profile for a bounded time (e.g. 300000 for five minutes). It exists because approving rm -rf /tmp/build every few seconds during an iterative task trains you to click through prompts — which is worse for safety than a grant you chose deliberately.

A grant is bound to the exact command text, the profile and the command class: approving rm -rf /tmp/build does not cover rm -rf /tmp/build-prod, the same command on another host, or the same command escalated to sudo. Runs covered by a grant appear in the audit log with approver: "jit-grant", so they stay distinguishable from a fresh human answer.

Off by default (0 = always prompt). Auto-approval weakens the gate that makes destructive commands safe, so turning it on should be a decision.

Answering the prompt

Approval goes through the MCP elicitation request, so what you see is your client's dialog. Accepting it approves the command — there is no second field to fill in.

You have 10 minutes to answer. Past that the request expires and the command is refused rather than left pending, and the refusal says so; the prompt may still be open in your client, in which case run the command again once you are ready. If your client does not support elicitation at all, every destructive and privileged command is refused with APPROVAL_UNAVAILABLE naming that cause — approval fails closed by design.

Command Quota

commandQuotaPerDay bounds how many commands a profile may run in a rolling 24-hour window (0 = unlimited). The approval gate stops destructive commands and the HTTP rate limiter caps request rate, but neither bounds total work — a prompt-injected agent looping over allowed commands stays under both. The quota is the circuit breaker for that case.

Counted after policy allows a command and before it runs, so a denied command does not spend budget. The window slides rather than resetting at midnight, which would let an agent spend a full quota just before the reset and another immediately after.

External Policy Engine (OPA)

For organizations that standardize on Open Policy Agent / Rego:

ssh-mcp --opaUrl=http://localhost:8181

When --opaUrl is set, commands the built-in engine allows are additionally evaluated by OPA. OPA can only narrow. A command the built-in engine has already denied returns that denial without OPA being consulted at all, so a sidecar answering allow cannot grant a class the role bindings withhold. To widen, edit [policy].

An outage falls back to the local decision and logs one warning per minute. That is the default because OPA is an additional deny layer and stopping all work would be the worse failure — but an operator who deployed OPA as the authorization gate loses that gate during the outage, and the only signal is a stderr line MCP clients usually discard. --opaFailClosed makes the gate being down mean no; the refusal carries ruleId: "opa-unavailable" so the audit record says the gate was down rather than implying a policy refused the command.

The request shape follows the AuthZEN Access Evaluation contract:

{
  "input": {
    "subject": { "role": "operator", "profile": "prod-web-1" },
    "action": { "tool": "run-command", "commandClass": "destructive" },
    "resource": { "command": "rm -rf /tmp/cache", "binary": "rm", "host": "10.0.1.50" },
    "context": { "readOnly": false }
  }
}

OPA responds with { "result": true/false }. If OPA denies (result: false), the command is blocked even if the built-in engine allows it. If OPA is unreachable, the built-in engine's decision stands by default (fail-open, to avoid locking out access); --opaFailClosed refuses instead. A 200 that carries no boolean result counts as unreachable — that is what OPA answers for an undefined document, so a misnamed package or an unactivated bundle is an outage rather than consent.

Example Rego policy (ssh-mcp.rego):

package ssh.mcp

default allow := false

# Admins pass the OPA gate on dev hosts. The built-in policy still applies on
# top: this widens nothing that the role bindings withhold.
allow if {
  input.subject.role == "admin"
  startswith(input.subject.profile, "dev")
}

# Deny all destructive commands on prod
deny if {
  input.action.commandClass == "destructive"
  startswith(input.subject.profile, "prod")
}

Security

Threat Model

See SECURITY.md for the full threat model, vulnerability reporting policy, and deployment checklist.

Supply chain

Releases carry signed attestations, published through Sigstore and recorded in its public transparency log. They live in two different stores, which is what decides how each is verified:

Attestation

Predicate

Stored by

Since

Build provenance — SLSA Build Level 2

slsa.dev/provenance/v1

npm

every release

SBOM — CycloneDX and SPDX

cyclonedx.org/bom, spdx.dev/Document

GitHub

releases after v2.4.0

Provenance comes from npm trusted publishing: the release workflow authenticates with a short-lived OIDC token and no stored credential, so there is no long-lived npm token to leak.

npm audit signatures        # provenance, against an installed tree

npm pack ssh-mcp            # the SBOM attestation is bound to the tarball, so fetch it
gh attestation verify ssh-mcp-*.tgz --repo tufantunc/ssh-mcp --predicate-type https://cyclonedx.org/bom

Both flags on the last command are load-bearing. gh attestation verify defaults to the SLSA predicate, so without --predicate-type it filters the SBOM out and reports nothing found — and the provenance it would look for instead is in npm's store, not the GitHub store --repo queries. Use https://spdx.dev/Document for the SPDX one.

Both SBOMs are also attached to each GitHub release, for reading rather than verifying.

Level 2, not 3. Provenance is signed by the generic GitHub-hosted runner — builder.id is https://github.com/actions/runner/github-hosted — which the build itself can influence; Build L3 requires an isolated builder it cannot. Reaching L3 is not currently compatible with trusted publishing: npm turns on its own provenance whenever that setting is left at its default, and then ignores any externally generated one. So L3 today would mean returning to a long-lived npm token — trading the property described above for a level number.

Safe Defaults

  • Non-root user in all examples

  • TOFU host key verification (accept on first connect, verify after — within one process; see SECURITY.md)

  • RFC 9142 algorithm allow-list (no SHA-1, no CBC, no ssh-rsa)

  • exec()-only for elevation (privileged-command takes no session, so sudo never runs in a persistent shell — fixes PTY leak. Interactive sessions exist for ordinary commands; this bullet is about elevation, not about every path being an exec channel)

  • Sudo via stdin (not argv — fixes process list leak)

  • Sanitizer strips CR/LF/NUL from all metadata

  • 3-layer redaction (field → regex → entropy) on audit logs

  • No CLI-arg secrets (use env vars, keychain, or config)

  • Agent announcement is disclosed, not hidden — every command tells the host AI_AGENT=ssh-mcp, including hosts that never opted in to storing it; clear announceAgent on a profile pointing at a host you do not control (see Host-side identification)

Hardening Checklist

  • Create dedicated low-privilege service account on target hosts

  • Use command-specific sudoers instead of NOPASSWD: ALL

  • Enable ask-all approval for production profiles

  • Restrict network egress on target hosts

  • Use readOnly = true for monitoring profiles

  • Review audit logs regularly

  • Run chmod 700 <config dir> && chmod 600 config.toml (Windows: the ACL under %APPDATA% is already restricted)


Transports

stdio (default)

For local MCP clients (Claude Code, Cursor, Windsurf). No network exposure.

ssh-mcp                          # reads config from XDG path
ssh-mcp --config=/path/to.toml   # custom config path

HTTP (optional)

For remote/web clients behind a reverse proxy with TLS:

ssh-mcp --transport=http --httpPort=3000 --bearerToken=secret
ssh-mcp --transport=http --httpPort=3000 --bearerToken=secret --rateLimit=60

Flag

Default

Description

--bearerToken

required

Bearer token for authentication (all routes except GET /health)

--httpPort

3000

HTTP listen port

--httpHost

127.0.0.1

Bind address

--rateLimit

0 (off)

Max requests per minute (0 = unlimited)

--authFailureLimit

10

Failed bearer-auth attempts allowed per client per minute (0 = off)

--trustProxy

false

Read the client address from X-Forwarded-For, but only when the peer is the proxy — bare means a loopback peer

--trustedProxies

—

Comma-separated peer addresses allowed to send X-Forwarded-For. Empty means loopback only

Endpoints: POST / (MCP Streamable HTTP), GET /status, GET /health

GET /health answers {"healthy": true, "configured": <bool>}. It stays 200 either way — healthy is liveness — while configured is false when no profile is set, which is the case of a config bind mount that silently did not attach: the server binds the port and refuses every tool call. GET /status carries the profile list itself and stays behind the bearer token.

When rate limit is exceeded, the server returns HTTP 429 with Retry-After header and a JSON-RPC error body so MCP clients can handle it gracefully.

Failed authentication is throttled separately, and on by default. --rateLimit never saw a wrong bearer token, because the auth check answers before the limiter is reached — so guessing ran at network speed. --authFailureLimit gives each client its own small budget, spent only on a 401; a correct token never consumes from it, so a working client never throttles itself. Once an address has spent its budget every request from it waits, including one with the right token — that is deliberate, since answering the guess would otherwise tell the caller which token was right. Clients are told apart by socket address. Behind a reverse proxy that means every client shares one budget, so set --trustProxy when the proxy is yours — the server prints a warning the first time it sees X-Forwarded-For without it. --trustProxy takes the rightmost X-Forwarded-For entry, which is the hop the proxy itself appended; everything to its left came from the client, so reading the leftmost would let a client choose its own budget or spend a victim's. That only holds if a proxy actually appended the entry, so the header is read only when the peer is the proxy — bare --trustProxy means a loopback peer, which is the deployment above; name a proxy elsewhere with --trustedProxies. When the header cannot be read as an address, or the peer is not trusted, the server falls back to the socket address and says so once, so a flag that is not taking effect is not silent. One trusted hop is assumed. A malformed --authFailureLimit is refused at startup rather than silently disabling the check; only 0 turns it off.

Always terminate TLS at a reverse proxy (Caddy/nginx). The server listens on 127.0.0.1 only.


Docker

# Build
docker build -t ssh-mcp .

# Run (config file + env vars for credentials)
docker run -i \
  -v ./config.toml:/home/appuser/.config/ssh-mcp/config.toml:ro \
  -e SSH_MCP_PASSWORD=secret \
  ssh-mcp

Or with docker-compose:

docker-compose --profile app up

The Docker image runs as non-root UID 65532, with a minimal node:22-slim base.


CLI Flags (v2)

Secrets are never passed as CLI arguments.

Flag

Default

Description

--config

platform config dir (see Configure)

Path to TOML config file

--host

—

Quick start: SSH host (creates single-profile config)

--user

—

Quick start: SSH username

--port

22

Quick start: SSH port

--key

—

Quick start: Path to private key

--workdir

—

Quick start: Working directory for commands and sessions

--group

prod

Quick start: Policy tier — prod, staging or dev

--timeout

60000

Command timeout in ms

--maxChars

5000

Max command length (none or 0 disables the limit; in a config file the same setting is commandMaxChars = 0)

--sessionMax

5

Max concurrent sessions per connection

--sessionTtl

600000

Session idle timeout in ms

--transport

stdio

stdio or http

--httpPort

3000

HTTP transport port

--httpHost

127.0.0.1

HTTP bind address

--bearerToken

—

Bearer token for HTTP transport auth (required for --transport=http)

--rateLimit

0

HTTP requests per minute on the MCP route (0 = unlimited)

--authFailureLimit

10

Failed bearer-auth attempts allowed per client per minute (0 = off)

--trustProxy

false

Read the client address from X-Forwarded-For, but only when the peer is the proxy — bare means a loopback peer

--trustedProxies

—

Comma-separated peer addresses allowed to send X-Forwarded-For. Empty means loopback only

--allowedHosts

bind address + localhost

Comma-separated Host headers accepted by the DNS-rebinding guard

--hostKeyMode

tofu

tofu | strict | insecure. strict accepts only hosts pinned with trustedHostKey. See SECURITY.md

--insecureHostKey

false

Disable host key verification for hosts with no trustedHostKey — a pin still refuses (test only!)

--allowUncheckedConfigAcl

false

Windows: report every ACL finding and refuse none

--strictConfigAcl

false

Windows: refuse on every ACL finding, including a read-only over-grant

--disableApproval

false

Skip the approval gate (quick start profile only)

--opaUrl

—

OPA sidecar URL for external policy

--opaFailClosed

false

Refuse every command while OPA is unreachable, instead of falling back to local policy

--opaTimeoutMs

10000

How long to wait for the OPA sidecar. Lower makes the fail-open cheaper to reach; higher makes an outage slower to notice

--commandQuota

0 (off)

Max commands per rolling 24h per profile

--approvalGrantTtl

0 (off)

Auto-approve an identical command for this many ms after approval

--auditEntropyScan

false

Enable entropy-based secret scanning in audit

--auditTamperEvident

false

Enable hash-chained tamper-evident audit log

--otelEndpoint

—

OTLP/HTTP endpoint for OpenTelemetry traces

--otelServiceName

ssh-mcp

Service name reported on trace spans

--dumpToolHashes

—

Print SHA-256 hashes of the tool descriptions and exit


Migrating from v1

v2 is a breaking release. Passing a removed flag now fails at startup with the replacement, rather than failing later as a confusing auth error.

Tools

v1

v2

Notes

exec

read-command

Allowlisted read-only commands. Prefer this for reads.

exec

run-command

Arbitrary commands. Destructive and privileged ones go through the approval gate, unless approvalPolicy = "auto".

sudo-exec

privileged-command

Requires approval unless approvalPolicy = "auto". Password is piped via stdin.

description parameter

—

Removed. It was an injection vector (#44) and never reached the host.

Command results now carry status. In v1 a failed command rejected with Error (code N). In v2 a non-zero exit comes back as an error result including the exit code and stderr — so an empty response no longer means "it worked".

Flags

v1 flag

Replacement

--password

SSH_MCP_PASSWORD env var (or SSH_MCP_<PROFILE>_PASSWORD)

--suPassword

SSH_MCP_SUDO_PASSWORD env var

--sudoPassword

SSH_MCP_SUDO_PASSWORD env var

--disableSudo

Use a role/policy that disallows the privileged class

Credentials moved off the command line because CLI arguments are world-readable via /proc/<pid>/cmdline on Linux (CWE-214). Credentials now resolve through an SSH agent → OS keychain → env var → key file cascade.

Example

// v1
{ "command": "npx", "args": ["ssh-mcp", "--host=1.2.3.4", "--user=root", "--password=hunter2"] }

// v2 — credentials via env
{
  "command": "npx",
  "args": ["ssh-mcp", "--host=1.2.3.4", "--user=root"],
  "env": { "SSH_MCP_PASSWORD": "hunter2" }
}

For more than one host, move to a TOML config file (see Configuration) and pass --config <path>; profiles carry per-host roles and approval policy.

Host key verification

v1 did not verify host keys. v2 defaults to trust-on-first-use and records the key in memory, for the life of the process; a later mismatch in that same process fails the connection. Nothing is written to disk and ~/.ssh/known_hosts is not consulted, so a restart accepts afresh — see SECURITY.md. Pin explicitly with trustedHostKey in a profile, which is the only control here that survives a restart — and which no host key mode overrides, so --insecureHostKey is an opt-out only for hosts you have not pinned (test environments only).


Testing

# Start test SSH servers. --build matters: compose reuses an existing image, and a
# fixture change (such as the alpine target's AcceptEnv line) then never lands,
# which shows up as a product failure in the integration suite.
docker compose --profile test up -d --build

# Run all tests
npm test

# Run only unit tests
npm test -- test/unit/

# Run with coverage
npm run coverage

MCP Inspector

npm run inspect

Contributing

See CONTRIBUTING.md. Please follow the security checklist in all PRs.

Support

If you find SSH MCP Server helpful, consider starring the repository or sponsoring!

Listed on

SSH MCP Server on Glama

Also on the official MCP registry as io.github.tufantunc/ssh-mcp.

Available Tools

14 tools
close-sessionA
Destructive

Close a named session. A background session's command is signalled on the host (INT, then TERM, then KILL) before its channel is dropped; an interactive session's shell is ended. The response says so if the command could not be signalled or had not stopped in time.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSession name to close
profileNoProfile name

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses the full shutdown sequence for background sessions (INT, then TERM, then KILL), that the channel is dropped, that interactive shells are ended, and that the response indicates failures to signal or stop in time. This is substantial behavioral context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core action is front-loaded, followed by essential behavioral details and response semantics. Every sentence contributes information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no output schema, the description covers the shutdown behavior, the difference between session types, and how failures are reported. Given the annotations and schema, nothing critical is missing for correct invocation and interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description adds the context of 'named session' but does not add new meaning to the parameters themselves. Baseline 3 is appropriate when the schema fully carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Close a named session,' which precisely states what the tool does. It further distinguishes behavior for background versus interactive sessions, separating it from list-only or read-only sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool's use case is implied: call it when you need to close a named session. However, it does not explicitly state when to use this over alternatives like signal-process or run-command, nor does it provide exclusion criteria. The context is clear but the guidance is not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-connectionsA
Read-only

List all configured SSH profiles and their connection status. Use this to discover available hosts before running commands.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safe-read nature is covered. The description adds that the tool lists profiles and connection status, which clarifies what data is returned but does not elaborate on side effects, permissions, or what 'connection status' means. This is modest additional context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first states the action and scope, the second explains when to use it. The key information is front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with no output schema, the description is complete: it names what is listed, the status attribute, and the intended use case. An agent can correctly decide to invoke this tool and know what to expect without additional guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts zero parameters and schema description coverage is 100% (vacuously). Per the baseline for zero-parameter tools, no further parameter documentation is required. The description' s functional overview is sufficient for an agent to understand there a re no inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource ('all configured SSH profiles') and an additional attribute ('their connection status'). It distinguishes from sibling 'list-sessions' by focusing on profiles rather than active sessions, and adds the intent of discovering available hosts before running commands.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage context: 'Use this to discover available hosts before running commands.' It clearly implies this is a preliminary, read-only discovery step. However, it does not name alternatives like 'list-sessions' or state when not to use this tool, which would have made the guidance stronger.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-sessionsB
Read-only

List active sessions for a given SSH profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoProfile name (uses default if omitted)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already declares this as a read-only operation, and the description's 'List' wording is consistent. The description adds the scope qualifier 'for a given SSH profile' but does not explain what 'active' means or how the tool behaves with an invalid profile. This is minimal added value beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence with the verb front-loaded and no redundant words. It communicates the essential action and scope immediately and efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter, read-only listing tool, the description covers the main selection and invocation needs: what it lists and for which scope. However, it omits guidance on the distinction from list-connections or the shape of the result, which would help an agent use it correctly. It is nearly complete but not fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'profile' is fully described in the schema ('Profile name (uses default if omitted)'). The description's phrase 'for a given SSH profile' restates the parameter without adding new meaning. Since schema coverage is 100%, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('active sessions') with a scope ('for a given SSH profile'), so the purpose is clear. It does not explicitly contrast with the sibling list-connections, so an agent must infer the distinction between sessions and connections. The verb+resource is specific but sibling differentiation is implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like list-connections or open-session. The description only restates what the tool does; it does not provide context or exclusions. An agent has no explicit direction for selecting this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open-sessionA

Open a named session on a remote host. Use type="interactive" for stateful shell (CWD/env persists between commands) or type="background" for long-running processes.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSession name (alphanumeric, dash, underscore, max 64 chars)
typeNoSession typeinteractive
commandNoCommand for background sessions
profileNoProfile name (uses default if omitted)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavior context: interactive sessions preserve CWD/env between commands, and background sessions are for long-running processes. However, with no annotations, it does not disclose lifecycle behavior such as whether sessions persist until explicitly closed, what happens on name collisions, or cleanup requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the core open-session purpose is front-loaded before type guidance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core semantics are covered, but with no output schema and no annotations, the agent is left guessing about return values, session lifecycle, and how this relates to close-session or run-command. The background command requirement is also not fully clarified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema by explaining when to use interactive vs. background and what stateful behavior means. It does not enrich name or profile parameters, but the type clarification is valuable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Open'), a specific resource ('named session'), and a target ('remote host'). It also distinguishes session types, which differentiates this from siblings like list-sessions, read-session-output, and close-session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear choice guidance for the type parameter: interactive for stateful shell, background for long-running processes. It does not explicitly contrast with run-command or close-session, so it stops short of full alternative-based routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

privileged-commandA
Destructive

Execute a command with sudo elevation. Goes through the approval gate; approvalPolicy on the profile decides whether that is a prompt, an automatic allow, or a refusal. The sudo password is piped via stdin (never visible in process list). Single-line only: a command containing a line break is refused, so upload a multi-line script with sftp-upload and run it by path.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesCommand to execute with sudo
profileNoProfile name

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint: true. The description goes well beyond this by disclosing the approval gate mechanics (prompt/allow/refusal based on profile approvalPolicy), the stdin password handling (never visible in process list), and the single-line-only refusal. This adds rich behavioral context beyond what annotations provide, and it does not contradict the destructive hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences that are front-loaded with the core purpose and immediately provide the most critical operational constraints (approval gate, password handling, single-line rule). Every sentence earns its place; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, elevated execution tool, the description covers all critical operational aspects: approval behavior, password security, input constraints, and the alternative path for multi-line scripts. The absence of an output schema is fine since command execution tools typically return stdout/stderr, and that is implied. Nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining that the profile parameter determines the approval policy behavior, and it clarifies the command's single-line constraint. It doesn't elaborate on the exact format of command, but the schema already covers that. The added profile context justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Execute a command'), a resource ('with sudo elevation'), and the approval gate context. It distinguishes from siblings by making clear this is the privileged variant (run-command is non-privileged) and explicitly references sftp-upload for multi-line scripts, so an agent can select it correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use (when sudo elevation is needed) and provides a clear alternative: 'upload a multi-line script with sftp-upload and run it by path' for anything containing a line break. It also implies that non-privileged commands belong in run-command, which is evident from naming and context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read-commandA
Read-only

Execute a READ-ONLY command from an allowlist (ls, cat, grep, find, stat, df, etc.). This tool does NOT modify the system. Prefer this tool for all read operations. Single-line only: a command containing a line break is refused, so upload a multi-line script with sftp-upload and run it by path.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesRead-only shell command (must be in the allowlist)
profileNoProfile name

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral constraints beyond the readOnlyHint annotation: it clarifies the tool does not modify the system, enforces a single-line-only restriction, and explains that line breaks cause refusal. It also provides a workaround, giving the agent actionable operational knowledge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, immediately stating the tool's core purpose and read-only nature. Every sentence adds value, covering scope, preference, and a key limitation with a clear alternative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for this tool's complexity: it covers the allowlist, read-only behavior, single-line limitation, and how to handle multi-line scripts. Combined with the readOnlyHint annotation and fully documented parameters, an agent has sufficient information to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents both parameters with 100% coverage: command is described as a read-only shell command in the allowlist, and profile is described as a profile name. The description mainly reinforces the command parameter's constraints but adds little beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool executes a READ-ONLY command from an allowlist, naming common examples like ls, cat, grep, and find. It also distinguishes itself from siblings such as run-command and privileged-command by emphasizing the read-only, allowlisted nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to prefer this tool for all read operations, establishing when it should be used. It also gives a concrete alternative for multi-line scripts by directing users to upload with sftp-upload and run by path, which helps an agent avoid rejected commands.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read-session-outputA
Read-only

Read recent output from a background session (e.g., tail -f logs).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesBackground session name
linesNoNumber of recent lines to read
profileNoProfile name

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a read-only operation, and the description does not contradict it. It adds context about targeting background-session output but does not disclose return format, ordering, or potential edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with a helpful example and no filler. Every part contributes to understanding the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with fully documented parameters and a readOnlyHint, the description is adequate for basic invocation. It could mention prerequisites like an active session or output format, but those are not critical for a tool this straightforward.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents 'name', 'lines', and 'profile'. The description adds no parameter-level meaning beyond the tail-like example, keeping this at the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'read' and the resource ('recent output from a background session'), which identifies the tool's purpose. It does not explicitly distinguish itself from the sibling read-command, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example 'tail -f logs' and the phrase 'background session' imply when this tool should be used, but there is no explicit guidance about alternatives or exclusions. It provides usable context without clear routing among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run-commandA
Destructive

Execute an arbitrary shell command on the remote server. May modify the system. Commands classified destructive or privileged go through the approval gate; approvalPolicy on the profile decides whether that is a prompt, an automatic allow, or a refusal. Single-line only: a command containing a line break is refused, so upload a multi-line script with sftp-upload and run it by path.

ParametersJSON Schema
NameRequiredDescriptionDefault
ttyNoAllocate a pseudo-terminal
commandYesShell command to execute
profileNoProfile name
sessionNoRun in an existing interactive session (stateful)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide destructiveHint=true. The description adds substantial value: 'May modify the system' reinforces destructive nature, but more importantly it discloses the approval gate behavior (prompt, auto-allow, refusal based on approvalPolicy) and the strict single-line refusal. These are non-obvious behaviors not present in the schema or annotations. It does not describe return values or error handling, but given the tool's nature and no output schema, this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and front-loaded: the core purpose appears first, followed by essential behavioral constraints and an alternative for a common edge case. No redundant phrases; each sentence contributes a distinct piece of information. It is concise yet comprehensive for the space it covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a command execution tool with 4 parameters, no output schema, and only a destructiveHint annotation, the description covers the key operational constraints (approval gate, single-line restriction) and gives an alternative for multi-line scripts. It lacks explicit differentiation from the sibling 'privileged-command' and does not describe session behavior beyond the schema, but overall it provides enough for an agent to call it correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning to the 'command' parameter by specifying the single-line restriction, and to 'profile' by explaining that approvalPolicy on the profile decides the approval gate. It does not elaborate on 'tty' or 'session', but the schema already describes them clearly. Thus it adds some value beyond the schema, warranting a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Execute an arbitrary shell command on the remote server.' This is unambiguous and distinct from siblings like read-command (which likely reads output) and list-sessions. It also implies scope (remote server) and the generality (arbitrary command), setting it apart from more specialized tools like sftp-upload.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: single-line only, with an explicit alternative for multi-line scripts ('upload a multi-line script with sftp-upload and run it by path'). It also explains the approval gate for destructive/privileged commands, which helps decide when this tool is appropriate. However, it does not explicitly contrast with the sibling 'privileged-command' – it only mentions that privileged commands go through approval, leaving the relationship ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sftp-downloadB
Read-only

Download a file from the remote server via SFTP.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoProfile name
remotePathYesRemote file path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters: the approval prompt and the audit record quote this path back.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds minimal behavioral context beyond that, such as the fact that it downloads from a remote server. It does not disclose details like whether a session must be open, how the file is transferred, or what happens on failure, but the annotation covers the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the action and resource. It is efficient and easy to parse, though it could add a bit more context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple download tool with full schema coverage and a readOnlyHint annotation, the description is mostly adequate. However, it lacks guidance on prerequisites (e.g., an open session) and does not clarify the relationship to sibling tools like sftp-download-file, which could lead to incorrect selection in a larger toolset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds no additional meaning beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Download') and resource ('a file from the remote server via SFTP'), which clearly identifies the tool's core action. It does not explicitly distinguish itself from sibling tools like sftp-download-file, but the action and resource are clear enough for basic selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (downloading via SFTP) but provides no explicit guidance on when to choose this tool over alternatives such as sftp-download-file or sftp-list. There is no mention of prerequisites like an open session or connection, which would help an agent decide when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sftp-download-fileA
Destructive

Download a remote file to local disk over SFTP, streaming it without passing the contents through model context. The destination must be inside the transferRoot directory the operator configured; without that setting this tool refuses. Use this for binary or large files; use sftp-download when you need to read the contents.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoProfile name
localPathYesLocal destination, inside defaults.transferRoot
overwriteNoReplace an existing local file (default false)
remotePathYesRemote file to download. No leading or trailing whitespace, and no control, bidirectional or zero-width characters.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits beyond annotations: it streams without passing contents through model context and refuses if transferRoot is not configured. Annotations only include destructiveHint: true, so the description adds valuable context. It doesn't cover overwrite behavior, but that is in the schema; overall it provides solid behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences. The first front-loads the core action and a key trait (streaming, no model context), while the second covers usage guidance and a critical constraint. No waste; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a download tool with full schema coverage and no output schema, the description covers the purpose, usage guidance, and key behavioral constraints. It doesn't mention error cases like missing files or authentication, but those are typical and not required. The description is sufficiently complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description does not add extra semantics about parameters beyond what the schema already documents. It mentions the streaming behavior but does not explain parameter-specific details or constraints beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Download a remote file to local disk over SFTP' with the specific streaming behavior. It also distinguishes itself from the sibling tool sftp-download by specifying when to use each ('Use this for binary or large files; use sftp-download when you need to read the contents'), making the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is provided: the tool is for binary/large files, and sftp-download is the alternative for reading contents. It also states a key prerequisite (the destination must be inside transferRoot) and that the tool refuses otherwise, which tells the agent when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sftp-listA
Read-only

List a remote directory over SFTP, with a bounded number of entries and a bounded response size. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoProfile name
maxEntriesNoMax entries to return (default 200, hard cap 1000)
remotePathYesRemote directory to list. No leading or trailing whitespace, and no control, bidirectional or zero-width characters.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already indicates no mutation, and the description reinforces this with 'Read-only.' It adds value beyond the annotation by disclosing bounded response size and bounded entry count, which are useful behavioral constraints an agent should know before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One crisp sentence that front-loads the core operation, then adds the two most important behavioral constraints. No wasted words or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list operation with one required parameter, full schema coverage, and a read-only annotation, the description is nearly complete. It lacks only an explicit statement of what the response contains (e.g., entry names/types), but the tool's purpose makes that reasonably inferable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters including maxEntries defaults and caps. The description's 'bounded number of entries' loosely echoes maxEntries, and 'bounded response size' adds a general constraint not tied to a specific parameter, but it does not meaningfully deepen parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List a remote directory over SFTP') and adds distinguishing constraints: bounded number of entries and bounded response size. This is clearly distinct from sibling tools like sftp-upload/download or list-connections, even though it does not name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the tool is appropriate: listing a remote directory. However, it offers no exclusions or guidance about when to prefer an alternative sibling tool, which would be helpful given the large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sftp-uploadA
Destructive

Upload a file to the remote server via SFTP (secure file transfer, not shell-based). Replaces an existing file at that path unconditionally — there is no overwrite flag and no way to require a new destination.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesFile content to upload
profileNoProfile name
remotePathYesRemote file path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters: the approval prompt and the audit record quote this path back.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses that the tool overwrites an existing file unconditionally, that there is no overwrite flag, and that there is no way to force a new destination. This gives the agent concrete insight into side effects. It does not cover edge cases like whether missing parent directories are created, but the main destructive trait is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two terse sentences, with the primary action first and the critical overwrite caveat immediately after. No filler or restatement of the name/schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter upload with complete schema and a destructiveHint annotation, the description covers the core action and danger. It is incomplete regarding the sibling sftp-upload-file relationship and the behavior when the target path does not already exist, which an agent may need to know before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents content, profile, and remotePath; the baseline is therefore 3. The description adds no new parameter-specific meaning beyond reinforcing that there is no overwrite flag, which is already evident from the schema. No deduction is warranted, but neither is any elevation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action (upload a file) and resource (remote server via SFTP), and adds the key caveat of unconditional replacement. It does not differentiate from the similarly named sibling sftp-upload-file, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'secure file transfer, not shell-based' gives implied context for choosing SFTP over command tools, and the overwrite warning implies avoiding it when a new destination is required. However, it never explicitly names an alternative (e.g., sftp-upload-file) or states when-not-to-use, so guidance remains implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sftp-upload-fileA
Destructive

Upload a local file to the remote host over SFTP, streaming it without passing the contents through model context. The local file must be inside the transferRoot directory the operator configured; without that setting this tool refuses. Use this for binary or large files; use sftp-upload for short text you already have.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoRemote file mode, 1 to 511 (0o777); setuid, setgid and the sticky bit are refused. e.g. 420 for 0644. Omit for 0600 on a new file, or the replaced file’s permission bits.
profileNoProfile name
localPathYesLocal file to upload, inside defaults.transferRoot
overwriteNoReplace an existing remote file (default false)
remotePathYesRemote destination path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, so the description does not need to re-state that this is potentially destructive. It adds helpful behavioral context: streaming behavior, the transferRoot prerequisite, and refusal behavior. However, it does not describe what happens on overwrite or failure beyond what the parameter schema already implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler. The core action and key behavior are front-loaded, followed by a prerequisite and explicit sibling routing. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with a fully documented schema and the destructive annotation, gives an agent enough to select and invoke the tool correctly. It lacks explicit return-value or failure-behavior details, but for an upload tool the streaming, constraint, and usage guidance cover the critical decisions. A brief note on overwrite consequences would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the structured parameter docs already carry the semantic load. The description's transferRoot mention is useful but also appears in the localPath schema. It does not add meaningful parameter information beyond what is already available.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Upload a local file to the remote host over SFTP') and adds a defining behavior: streaming without passing contents through model context. It also explicitly distinguishes this tool from sftp-upload, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use guidance ('Use this for binary or large files') and names the alternative ('use sftp-upload for short text you already have'). It also states a hard prerequisite: the local file must be inside transferRoot or the tool refuses.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signal-processA
Destructive

Send a signal (INT, TERM, KILL) to a remote process by PID.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidYesProcess ID to signal (positive integer)
signalNoSignal to sendTERM
profileNoProfile name

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The destructiveHint annotation already signals that this is a destructive operation. The description adds that the target is a remote process and enumerates the available signals, but it does not disclose consequences such as process termination, irreversibility, or that KILL is more forceful than TERM. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It states the action, signal options, target, and identifier in minimal words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter mutation with a destructive annotation, this is mostly adequate. However, it lacks operational details: what the profile is for, whether an active session is required, and what a successful call returns (no output schema). These are not fatal for a simple tool, but they are clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds little beyond the schema: it mentions PID and the signal enum, both already defined. The 'profile' parameter remains unexplained in the description, though the schema labels it 'Profile name'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('send'), a precise object ('signal'), and a clear target ('remote process by PID'). It also lists the supported signal values (INT, TERM, KILL), making the operation unambiguous and naturally distinct from sibling session/command/SFTP tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement about when to choose this tool, what prerequisites are required (e.g., an open session/connection or profile), or when not to use it. The purpose implies the use case, but no guidance is provided and sibling tools like run-command are not ruled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv2.11.0
    • Changedsftp-download1 field changed
      • changedInput schema / properties / remotePath / description
        Previous value: -"Remote file path to download"New value: +"Remote file path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters: the approval prompt and the audit record quote this path back."
    • Changedsftp-download-file1 field changed
      • changedInput schema / properties / remotePath / description
        Previous value: -"Remote file to download"New value: +"Remote file to download. No leading or trailing whitespace, and no control, bidirectional or zero-width characters."
    • Changedsftp-list1 field changed
      • changedInput schema / properties / remotePath / description
        Previous value: -"Remote directory to list"New value: +"Remote directory to list. No leading or trailing whitespace, and no control, bidirectional or zero-width characters."
    • Changedsftp-upload1 field changed
      • changedInput schema / properties / remotePath / description
        Previous value: -"Remote file path"New value: +"Remote file path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters: the approval prompt and the audit record quote this path back."
    • Changedsftp-upload-file1 field changed
      • changedInput schema / properties / remotePath / description
        Previous value: -"Remote destination path"New value: +"Remote destination path. No leading or trailing whitespace, and no control, bidirectional or zero-width characters."
  2. 3 tool updatesv2.9.0
    • Addedsftp-download-file
    • Addedsftp-list
    • Addedsftp-upload-file
  3. 11 tool updatesv2.5.0
    • First observedclose-session
    • First observedlist-connections
    • First observedlist-sessions
    • First observedopen-session
    • First observedprivileged-command
    • First observedread-command
    • First observedread-session-output
    • First observedrun-command
    • First observedsftp-download
    • First observedsftp-upload
    • First observedsignal-process

TDQS

A3.9/5.0

Scored across 14 tools

Disambiguation4/5

Tools are mostly distinct, but sftp-upload and sftp-upload-file (and similarly download) have overlapping purposes; however, their descriptions clearly differentiate by use case (text content vs. streaming files). run-command and privileged-command are also similar but clearly separated by privilege level.

Naming Consistency4/5

All tools use lowercase hyphenated names, but there are two patterns: verb-first (list-sessions, open-session) and sftp-prefixed (sftp-upload, sftp-download). This is consistent within groups, though not a single uniform verb_noun pattern, making it slightly inconsistent.

Tool Count5/5

14 tools cover a broad range of SSH operations—session management, command execution, file transfer, process signaling—without being excessive. Each tool serves a clear purpose, and the count feels well-scoped for the domain.

Completeness3/5

The tool set covers most common SSH workflows, but there is a notable gap: open-session creates an interactive/stateful session, yet there is no tool to send commands into that session. This limits the usefulness of interactive sessions and represents a missing lifecycle operation.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A server that enables remote command execution over SSH through the Model Context Protocol (MCP), supporting both password and private key authentication.
    1
    17 npm
    2
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    A local Model Context Protocol server that allows LLMs to securely execute shell commands on remote Linux and Windows systems via SSH connections.
    6
    19 npm
    2
    -
  • A
    license
    A
    quality
    D
    maintenance
    A Model Context Protocol server for secure local system operations, enabling shell command execution and file management via a standardized interface.
    14
    1
    Apache 2.0