ssh-hop
Allows managing Ubuntu servers over SSH/SFTP, including running commands, checking connectivity, and transferring files.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ssh-hoprun 'df -h' on ubuntu-01"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ssh-hop
给 AI 用的 SSH / SFTP 代理。 本机 stdio MCP 服务,AI 只认别名,凭据永远留在本机进程里。
局域网自用、单人使用、无端口、无容器、无守护进程。每次工具调用复用(或新建)一条 SSH 连接, 执行完就返回。附带一个功能等价的命令行工具。
AI 客户端 ──stdio──► ssh-hop ──SSH/SFTP──► hosts.json 里配置的机器
(只看得到别名) (凭据唯一的持有者)支持 Cherry Studio / Claude Code / Codex / Cursor 等任何 MCP 客户端,以及 OMP。
目录
Related MCP server: ssh-mcp-server
为什么需要它
直接让 AI 连 SSH 有几个现实问题:
问题 | ssh-hop 的做法 |
把 IP、账号、密码写进提示词 | 写进本机 |
AI 可能连错机器 | 别名精确匹配,写错直接报错并列出可用别名,绝不模糊匹配 |
每次都要重新握手,很慢 | 连接池复用,空闲 30 秒内不再握手 |
不知道 AI 到底执行了什么 | 每次远程操作追加一行 JSON 审计日志,自动脱敏 |
有的机器不想让它改 | 每台机器单独配置命令正则白名单/黑名单 |
安装
需要 Python 3.10+。推荐用 uv(会把依赖装进独立环境,不污染系统 Python)。
方式一:从源码装(推荐,可改代码)
git clone https://github.com/onlineY/ssh-mcp.git
cd ssh-mcp
uv tool install .会生成两个命令:ssh-hop(命令行)和 ssh-hop-mcp(MCP 服务)。
方式二:不装,直接跑
git clone https://github.com/onlineY/ssh-mcp.git
cd ssh-mcp
uv venv && uv pip install -e .然后用 .venv/Scripts/python.exe -m ssh_hop(Windows)或 .venv/bin/python -m ssh_hop(Linux/macOS)当 MCP 命令。
方式三:从 Release 装
到 Releases 下载 .whl 文件:
uv tool install ./ssh_hop-0.1.0-py3-none-any.whl安装后会得到什么
项目 | 路径 |
命令入口 |
|
程序本体 |
|
验证:
ssh-hop --version快速开始
# 1. 生成配置文件
ssh-hop init
# 2. 编辑它,填入真实的主机、账号、密码(或密钥路径)
# Windows: C:\Users\<你>\.ssh-hop\hosts.json
# Linux: ~/.ssh-hop/hosts.json
# 3. 校验格式(不连网)
ssh-hop check
# 4. 校验 + 真实连接测试
ssh-hop check --connect
# 5. 试一条命令
ssh-hop run ubuntu-01 'uname -a'check --connect 成功的样子:
hosts file OK: C:\Users\me\.ssh-hop\hosts.json (2 host(s))
ubuntu-01 ok 1197ms Linux app 6.8.0-1062-azure ... x86_64 GNU/Linux
nas ok 120ms Linux nas 5.10.0 ... aarch64 GNU/Linux配置 hosts.json
默认位置:~/.ssh-hop/hosts.json(Windows 是 C:\Users\<你>\.ssh-hop\hosts.json)。
查找顺序:--hosts-file 参数 → 环境变量 SSH_HOP_HOSTS → 当前目录 ./hosts.json → ~/.ssh-hop/hosts.json。
最小配置
{
"hosts": {
"ubuntu-01": {
"host": "192.168.1.10",
"user": "deploy",
"password": "你的密码"
}
}
}完整示例
{
"defaults": {
"port": 22,
"timeout": 60,
"connectTimeout": 10,
"netTimeout": 20,
"idleReuseSec": 30,
"knownHosts": "auto-add",
"allowCommands": ["*"],
"denyCommands": []
},
"hosts": {
"ubuntu-01": {
"desc": "Ubuntu 应用服务器,docker compose 在 /opt/app",
"host": "192.168.1.10",
"user": "deploy",
"password": "你的密码",
"defaultCwd": "/opt/app"
},
"nas": {
"desc": "NAS,用密钥登录",
"host": "192.168.1.20",
"user": "admin",
"keyFile": "C:/Users/me/.ssh/id_ed25519",
"passphrase": null
},
"router": {
"desc": "路由器,只准看不准改",
"host": "192.168.1.1",
"user": "root",
"password": "你的密码",
"allowCommands": [
"^(logread|ubus|uci|ip|iw|df|free|ps|cat|ls|tail|grep|uname)\\b"
],
"denyCommands": ["(sysupgrade|firstboot|mtd)"],
"allowUpload": false
}
}
}字段说明
defaults 里的每一项都可以被单台主机覆盖。
字段 | 默认值 | 说明 |
| — | 必填。IP 或域名 |
| — | 必填。登录用户名 |
| — | 密码认证。和 |
| — | 私钥路径(不是 |
| — | 私钥有密码时填 |
|
| 给 AI 看的说明。想隐藏就不要写 IP/账号 |
|
| SSH 端口 |
| — | 每条命令前自动 |
|
| 单条命令超时秒数(上限 1800) |
|
| TCP / 认证超时秒数 |
|
| 整次握手的硬上限,防止对端卡住导致永久挂起 |
|
| 空闲多久内复用连接; |
|
| 命令白名单正则, |
|
| 命令黑名单正则,优先级高于白名单 |
|
| 是否允许传文件 |
|
| 远端可写/可读目录;空 = 任意路径 |
|
| 本地可读写目录;空 = 任意路径 |
|
|
|
|
| 执行 |
|
| 每条命令前注入的环境变量 |
|
| 自由标签,会返回给 AI |
两种认证方式
// 密码
{ "user": "deploy", "password": "hunter2" }
// 密钥(私钥路径,不是 .pub)
{ "user": "deploy", "keyFile": "C:/Users/me/.ssh/id_ed25519" }
// 密钥 + 私钥密码
{ "user": "deploy", "keyFile": "/home/me/.ssh/id_rsa", "passphrase": "xxxx" }踩坑提醒
Windows 路径用正斜杠
/或双反斜杠\\;单个\在 JSON 里是转义符,会解析失败。程序拒绝
REPLACE_ME/CHANGE_ME/your-password这类占位符,防止你忘了改就能连上。
hosts.json是明文凭据,不要提交到 git(仓库的.gitignore已经排除)。
命令权限(正则)
默认 ["*"],什么都不限制。需要收窄时用正则:
写法 | 含义 |
| 全部允许(默认) |
| 只允许以 docker 开头的命令 |
| 精确到一条 |
| 任意位置含 docker 就放行( |
规则细节:
用
re.search匹配整条命令,开启DOTALL(.能跨行,所以裸*也能匹配多行的 heredoc)。想钉住整条命令就用
^/$锚定。每个列表里第一条命中的规则生效;
denyCommands先判,永远压过allowCommands。没命中任何白名单规则就拒绝,报错信息会列出你配置的规则,AI 一步就能自查。
白名单写成空列表会回退成
["*"],避免把自己锁死。正则写错会在加载
hosts.json时报错,不会等到执行才炸。
不改配置也能先查
不想真跑,只想问"这条命令会不会被放行":
ssh-hop classify router 'reboot'host: router
command: reboot
verdict: refused
reason: refused on 'router': matches denyCommands rule '(sysupgrade|firstboot|mtd)'...
allow: '^(logread|ubus|uci|ip|iw|df|free|ps|cat|ls|tail|grep|uname)\\b'
deny: '(sysupgrade|firstboot|mtd)'MCP 里对应 ssh_run_policy 工具,同样不连网、不执行。
客户端接入
通用做法:MCP 客户端按 命令 + 参数 + 环境变量 拉起一个子进程,用 stdio 说 JSON-RPC。
下面是各客户端的配置位置和写法(把路径换成你自己的)。
Cherry Studio
设置 → MCP 服务器 → 添加。类型选 stdio,命令填 ssh-hop-mcp.exe 的完整路径。
Cherry Studio 把 MCP 配置存在 SQLite 里(
%APPDATA%\CherryStudio\Data\cherrystudio.sqlite的mcp_server表),没有 JSON 文件可编辑,所以只能在界面里加。
Claude Code
claude mcp add ssh-hop --scope user \
-e SSH_HOP_HOSTS="C:/Users/你/.ssh-hop/hosts.json" \
-- "C:/Users/你/.local/bin/ssh-hop-mcp.exe"验证:claude mcp list 应显示 ✓ Connected。
Codex
编辑 ~/.codex/config.toml:
[mcp_servers.ssh-hop]
command = "C:/Users/你/.local/bin/ssh-hop-mcp.exe"
env = { SSH_HOP_HOSTS = "C:/Users/你/.ssh-hop/hosts.json" }
startup_timeout_sec = 30
tool_timeout_sec = 300验证:codex mcp list。
Cursor / Claude Desktop / 其他
{
"mcpServers": {
"ssh-hop": {
"command": "C:/Users/你/.local/bin/ssh-hop-mcp.exe",
"env": { "SSH_HOP_HOSTS": "C:/Users/你/.ssh-hop/hosts.json" }
}
}
}Cursor:
~/.cursor/mcp.jsonClaude Desktop:
%APPDATA%\Claude\claude_desktop_config.json
OMP
编辑 ~/.omp/agent/mcp.json:
{
"mcpServers": {
"ssh-hop": {
"type": "stdio",
"command": "C:/Users/你/.local/bin/ssh-hop-mcp.exe",
"env": { "SSH_HOP_HOSTS": "C:/Users/你/.ssh-hop/hosts.json" }
}
}
}路径要点:
--hosts-file之外的路径请用绝对路径 + 正斜杠。 环境变量SSH_HOP_HOSTS指定配置文件,SSH_HOP_HOME指定审计日志和 host key 的存放目录。 即使一个环境变量都不传,也能自动找到~/.ssh-hop/hosts.json。
改完配置要重启客户端,它才会重新读取工具列表。
AI 可用的 8 个工具
工具 | 作用 |
| 列出所有别名、说明、命令规则、传输策略。AI 应先调它 |
| 测试连通性,返回 |
| 执行 shell 命令,返回 stdout / stderr / 退出码 / 耗时。 |
| 只判断命令会不会被放行、命中哪条规则,不执行 |
| 同一命令并行跑多台,逐台返回结果 |
| 上传本地文件/目录到远端 |
| 下载远端文件/目录到本地 |
| 列远端目录,或 stat 单个文件 |
失败时返回结构化结果,而不是抛异常堆栈,方便 AI 自行纠正:
{ "ok": false, "error": "refused on 'router': matches denyCommands rule ...", "kind": "refused" }kind 取值:refused(权限拒绝) / ssh(连接或远端错误) / config(配置问题) / unknown-host / usage / internal。
命令行用法
不开 MCP 客户端也能用,功能等价:
# 列主机和它们的策略
ssh-hop ls
# 修命令能不能跑(不连网)
ssh-hop classify router 'reboot'
# 执行命令
ssh-hop run ubuntu-01 'docker compose ps'
# 后台执行长任务,返回 pid 和日志路径
ssh-hop run ubuntu-01 'nohup ./deploy.sh > /tmp/deploy.log 2>&1 & echo started' --background
# 传文件
ssh-hop put ubuntu-01 ./dist/app.tar.gz /opt/app/app.tar.gz
ssh-hop get ubuntu-01 /var/log/syslog ./syslog --json
# 列远端目录
ssh-hop rls ubuntu-01 /opt/app
# 校验配置
ssh-hop check --connect
# 生成配置模板
ssh-hop init退出码:0 成功 / 其他为命令退出码 / 1 连接或配置错误 / 2 被策略拒绝。
安全边界
凭据不外泄
AI 的工具返回里没有 IP、端口、用户名、密码、密钥路径。
别名精确匹配,写错就报错并列出可用别名,不会连错机器。
审计
每次远程操作追加一行 JSON 到
~/.ssh-hop/audit.jsonl,password=/ token 自动脱敏。用
SSH_HOP_HOME可改存放位置。
传输
远端 host key 记入
~/.ssh-hop/known_hosts(首次遇到新 key 时创建),同时尊重~/.ssh/known_hosts。auto-add(默认)会接受没见过的 key 并记下来——新装的路由器必须这样;但记下来之后密钥变了就会硬失败并给出处理指引。strict模式只接受已知 key。
这不是沙箱
命令策略只是一道正则闸门——规则只决定这条命令发不发出去,发出去之后就是你账号的完整权限。
要真正的隔离,请用受限账号、sudo 规则或容器。
常见问题
Q:为什么都是 npm 的 MCP,Python 怎么让 AI 启动?
MCP 本质就是"用 stdio 跑一个子进程说 JSON-RPC",客户端不关心是 node 还是 python。
npm 生态多是因为它由 Anthropic 的 TS SDK 起步、npx -y 能免安装运行。
Python 的区别只有一个:依赖要先装好(npx 会自动下载,python 不会)。
所以别用裸 python,而是用 uv tool install 把依赖固化成一个可执行文件
(~/.local/bin/ssh-hop-mcp.exe),客户端直接拉起它,不需要你手动激活任何虚拟环境。
Q:会每条命令都重新建连接吗?
不会。一台主机保持一条连接,空闲 30 秒内复用。实测:
场景 | 耗时 |
冷启动(新建 SSH) | 1522 ms |
复用中 | 287–520 ms |
空闲超窗后重连 | 1318 ms |
想更激进地复用,把 idleReuseSec 调大即可(比如 600 表示 10 分钟)。
Q:连不上,怎么排查?
ssh-hop probe <别名>会返回 uname / 延迟 / SSH 服务端版本。如果报 netTimeout ... exceeded,说明握手卡住了
(对端没发 SSH banner),可以调大 netTimeout。
Q:为什么 uuidgen 这种命令被拒绝了?
说明那台主机配了 allowCommands 白名单,而 uuidgen 不在里面。用 ssh-hop classify <别名> '<命令>'
看是哪条规则拦的,或者把 "*" 加进该主机的 allowCommands。
Q:连接偶尔断,命令会执行两次吗?
不会。只有当命令还没送到远端(连接在网络层就死了)才会自动重连重试一次;
一旦命令已经发出,绝不重试,避免重复执行。重试过会在返回里带一条 warnings 提示。
Q:改了 hosts.json 要重启客户端吗?
ssh-hop 命令行立即生效;MCP 客户端需要重启(它会缓存工具列表)。若客户端传了环境变量,
改完配置重启即可。
开发
git clone https://github.com/onlineY/ssh-mcp.git
cd ssh-mcp
uv venv
uv pip install -e ".[dev]"
# 跑测试
.venv/Scripts/python.exe -m pytest # Windows
.venv/bin/python -m pytest # Linux/macOS
# 静态检查
uv pip install pyflakes
.venv/Scripts/python.exe -m pyflakes src/ssh_hop/*.py tests/*.py项目结构
src/ssh_hop/
config.py hosts.json 加载与校验、按主机的策略
guard.py 命令正则策略、路径收敛、输出截断
client.py 连接池、执行、SFTP、审计(凭据只存在这里)
server.py MCP 服务,8 个工具
cli.py 命令行
tests/
fake_ssh.py 进程内 SSH 服务器(真实 paramiko 传输层 + SFTP 子系统)测试用 tests/fake_ssh.py —— 一个跑在临时目录里的真实 SSH 服务器
(真的 SSH 传输层、真的 SFTP 子系统、一个小型 shell 解释器),
所以 SSH/SFTP/MCP 全链路都能端到端验证,不需要真机。
发布新版本
改 pyproject.toml 里的 version,然后:
git tag v0.1.1
git push origin v0.1.1GitHub Actions 会自动跑测试、构建 wheel、创建 Release 并附上 .whl 文件。
许可证
Available Tools
8 toolssftp_downloadA
Download a remote file or directory to the local machine over SFTP. Use to pull logs, config or artifacts for offline analysis. remotePath and localPath are both plain paths; if the host sets downloadRoots/localRoots, each must sit inside its list.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| localPath | Yes | ||
| overwrite | No | ||
| remotePath | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose a meaningful constraint: paths must sit inside downloadRoots/localRoots when the host defines them. However, it is silent on the overwrite default of true (which can clobber local files), on directory recursion behavior, and on auth/permission needs, leaving real behavioral gaps for a tool that writes to the local filesystem.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences: action first, then use case, then the constraint. Every sentence carries weight, though the run-on feel of the final sentence slightly blurs the structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. Against 0% schema coverage and no annotations, the description covers purpose, use case, and path constraints but omits the overwrite clobber risk and directory-handling behavior an agent would want before invoking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds genuine meaning for remotePath and localPath (plain paths, path-root containment) that the schema lacks entirely, but says nothing about host or the consequential overwrite flag, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Download), resource (remote file or directory), destination (local machine), and transport (SFTP). The direction is unmistakable and cleanly separates it from the sftp_upload sibling, which pulls/pushes in the opposite direction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use to pull logs, config or artifacts for offline analysis' gives a concrete usage scenario and intent. It stops short of explicitly naming when NOT to use it or pointing at alternatives like sftp_list, so it is clear context without routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sftp_listA
List a directory on a remote host, or stat a single remote file. Use to check whether a deployment landed, inspect permissions, or find log files before downloading them.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | ||
| host | Yes | ||
| remotePath | No | / |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It usefully discloses the two operating modes (directory listing vs single-file stat), but says nothing about authentication expectations, error behavior on missing paths, or that the operation is non-mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the core behavior is front-loaded, followed immediately by practical use cases. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary. However, with zero annotation coverage and zero schema description coverage, the undocumented 'all' parameter and absent behavioral cues (read-only nature, error handling) leave real gaps for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It hints at path semantics via 'directory' vs 'single remote file', but adds nothing about remotePath's default of '/', and the 'all' parameter (presumably hidden-file inclusion) is unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States concrete verbs (list, stat) and resources (remote directory, single remote file), so the dual-mode behavior is immediately clear. It doesn't name which siblings to prefer over it, though the reference to finding files 'before downloading them' implicitly points at sftp_download.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete when-to-use scenarios: verifying a deployment landed, inspecting permissions, and locating log files prior to download. No explicit when-not or named alternatives, but the context is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sftp_uploadA
Upload a local file or directory to a remote host over SFTP. Use to ship code, config or binaries before running deploy commands. remotePath must be absolute; localPath is any readable local path. If the host sets uploadRoots, the remote path must sit inside one.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| mode | No | ||
| localPath | Yes | ||
| overwrite | No | ||
| remotePath | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a genuine behavioral constraint (the uploadRoots containment rule) plus remotePath absoluteness. However it is silent on the default overwrite=true semantics, i.e. that existing remote files are replaced by default, and says nothing about permissions/mode behavior or authentication needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with what the tool does and why you would reach for it, and no filler. The path-constraint sentences are dense but each carries a distinct rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. But for an unannotated mutation tool with 5 parameters, the description omits the destructive-by-default overwrite behavior and the meaning of mode, leaving an agent without enough to call it safely in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must compensate. It meaningfully explains remotePath (must be absolute, must sit under uploadRoots) and localPath (any readable local path), but leaves mode and overwrite completely undefined — overwrite being the most consequential omission since it defaults to true.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (upload) and resource (local file or directory to a remote host over SFTP), which is unambiguous on its own. It does not explicitly contrast itself with the mirror-image sibling sftp_download, so an agent must infer the direction distinction rather than being told.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use to ship code, config or binaries before running deploy commands" gives a concrete usage context and even implies ordering relative to the ssh_run deploy siblings. It stops short of naming when-not-to-use or explicitly pointing at sftp_download for the reverse direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ssh_list_hostsA
List the remote hosts reachable through this server. Call this first whenever you need to touch a remote machine, to learn valid aliases and each host's permissions. Only aliases and descriptions are returned - never addresses or credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses what is and is not returned (aliases and descriptions, never addresses or credentials), which is security-relevant transparency. It stops short of covering auth requirements, failure modes, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose and followed by usage and return-scope. Every sentence carries distinct information; nothing is redundant padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail need not be explained here, and there are no parameters to document. The description supplies the discovery/ordering context an agent needs to place this tool correctly in the ssh workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4 per the rubric. The only parameter-like semantics introduced is that returned values are valid aliases usable with other tools, which is helpful framing but does not change the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the remote hosts reachable through this server') and immediately scopes it against siblings by clarifying that only aliases and descriptions are returned, not addresses or credentials. An agent can distinguish this discovery tool from ssh_probe, ssh_run, and sftp_list without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit ordering guidance ('Call this first whenever you need to touch a remote machine') and states the goal of that call (learning valid aliases and permissions). It does not name specific alternatives or state when-not to use it, so it falls just short of the 5-level criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ssh_probeA
Test connectivity to a host and return its identity (uname, uptime, id) plus round-trip latency. Use to diagnose 'cannot connect' problems before running real commands.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the runtime behavior — it probes the host and returns identity fields plus latency, and it is positioned as a non-mutating check to run before real commands. It does not mention auth requirements, timeout behavior, or whether the probe leaves any trace on the host, so it is good but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and followed by the usage trigger, with no filler. The parenthetical return list is slightly dense but earns its place by previewing the diagnostic payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no further explanation, and a one-parameter diagnostic tool has limited surface to cover. The description addresses purpose, usage trigger, and runtime behavior; the only real gap is host-parameter format guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter 'host' has no documentation anywhere. The description never clarifies the expected format (hostname, IP, user@host, or ssh-config alias), which is exactly the ambiguity that matters for a connectivity tool with a low-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation (test connectivity), its target (a host), and exactly what it returns (uname, uptime, id, round-trip latency). It also implicitly separates itself from the sibling ssh_run by framing itself as a pre-flight diagnostic, so an agent can pick it out of the ssh_* family without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: 'Use to diagnose cannot connect problems before running real commands,' which routes the agent to use this before ssh_run/ssh_run_many. It lacks an explicit when-not-to-use statement or naming of the sibling tools, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ssh_runA
Execute a shell command on a remote host and return stdout, stderr and the exit code. Use for inspecting state (logs, processes, ports, containers, service status) and for changing it (deploy, restart, edit). Prefer one command per call; chain with && when steps must be ordered. Set background=true for long-running or daemonizing commands so the call returns at once. Each host has its own allowCommands/denyCommands rules; a refusal names the rule that blocked it.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| host | Yes | ||
| command | Yes | ||
| timeout | No | ||
| background | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and largely succeeds: it discloses the return shape, the background/daemonizing behavior, and the policy-refusal model ('a refusal names the rule that blocked it'). It does not explain timeout behavior, cwd semantics, or whether commands are run as a privileged user, leaving a few operational unknowns for a mutation-capable tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with purpose then immediately branching into read vs. write usage, sequencing advice, the background flag, and the policy caveat. No filler and every sentence carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be spelled out, and the description still adds policy-refusal and background context. However, for a tool that can mutate remote state, the absence of any note on timeout defaults, cwd resolution, or privilege level leaves small but real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it only partially does: it explains background=true and ties host to per-host allowCommands/denyCommands rules. cwd, timeout, and the format/escaping of command are undocumented in both schema and description, so roughly half the parameters remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Execute a shell command on a remote host') and even states the return payload (stdout, stderr, exit code). It implicitly separates itself from ssh_run_many by emphasizing 'one command per call', and the policy/allowCommands note ties it to the ssh_run policy siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use is given for both read ('inspecting state: logs, processes, ports, containers, service status') and write ('deploy, restart, edit') scenarios, plus sequencing guidance ('chain with &&') and a concrete rule for background=true. The only gap is that it never names ssh_run_many as the multi-host alternative, but it does constrain its own scope clearly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ssh_run_manyA
Run the same command across several hosts in parallel and return one result per host. Use to compare versions, check a service fleet-wide, restart several boxes, or survey disk usage. Each host applies its own allowCommands/denyCommands, so one host may be refused while others run; per-host status is reported rather than failing the whole call.
| Name | Required | Description | Default |
|---|---|---|---|
| hosts | Yes | ||
| command | Yes | ||
| timeout | No | ||
| stop_on_error | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses that each host applies its own allowCommands/denyCommands policy, that one host may be refused while others run, and that per-host status is reported rather than failing the whole call. Missing are auth requirements, parallelism/concurrency limits, and the interaction between timeout and stop_on_error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then usage examples, then the key behavioral caveat. No filler and every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description still adds the useful 'one result per host' framing. For a 4-parameter fan-out tool with no annotations, the omission of timeout semantics and any auth/parallelism notes keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for four undocumented parameters. It clarifies that 'hosts' is a multi-host fan-out and 'command' is identical across hosts, and it implicitly characterizes the default stop_on_error=false behavior ('per-host status is reported rather than failing the whole call'). However, 'timeout' is entirely unaddressed and stop_on_error is only implied, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run the same command across several hosts in parallel') plus the output shape ('one result per host'), which cleanly distinguishes it from the single-host ssh_run sibling. An agent can pick this over ssh_run without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives four concrete usage scenarios (compare versions, fleet-wide service check, restart several boxes, survey disk usage) that map well to when this tool is appropriate. It stops short of naming alternatives or stating when not to use it (e.g. single host → ssh_run), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ssh_run_policyA
Check whether a command would be allowed on a host, WITHOUT running it. Returns the rule that matched. Use before a risky or unfamiliar command, or when a previous command was refused and you want to know why.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| command | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden — and it does disclose the critical trait that nothing is executed (no side effects), plus that the matched rule is returned. It omits auth/permission prerequisites and whether the check is bound to the active session/credentials, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler. The most important fact (no execution) is front-loaded in the first sentence, and the additional return/usage notes follow in priority order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the key dry-run semantics plus when to call it. Missing only permission/session prerequisites, which would matter for a policy-evaluation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required params, so the description must compensate — but 'host' and 'command' are only mentioned obliquely ('a command on a host') with no format, naming convention, or syntax guidance. An agent has to guess how hosts are addressed and whether multi-command strings are allowed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check/would be allowed) and resource (a command on a host), and the all-caps 'WITHOUT running it' immediately separates it from what a naive reader would assume ssh_run does. An agent can distinguish this dry-run/authorization check from ssh_run without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives two concrete triggering situations: before a risky/unfamiliar command, or to diagnose a prior refusal. It does not explicitly name ssh_run as the alternative being gated, so routing is implied by the description rather than stated as a rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
sftp_download - First observed
sftp_list - First observed
sftp_upload - First observed
ssh_list_hosts - First observed
ssh_probe - First observed
ssh_run - First observed
ssh_run_many - First observed
ssh_run_policy
TDQS
Scored across 8 tools
Each tool targets a distinct action: listing hosts, probing connectivity, checking policy, running commands (single or multiple), and SFTP upload/download/list. Boundaries are clear from descriptions; no two tools appear to do the same thing. An agent can confidently select the right tool for a given task.
Tool names follow a predictable pattern: ssh_ prefix for SSH operations and sftp_ prefix for file transfer operations, with consistent verb_noun or verb phrasing (list_hosts, probe, run_policy, run, run_many, upload, download, list). The subdomain prefixes aid grouping without breaking consistency. No mixed conventions or vague verbs.
Eight tools are well-scoped for an SSH/SFTP hop server, covering host discovery, connectivity testing, policy validation, command execution, and file transfer. Each tool earns its place, and the count is neither bloated nor thin. The set is focused on the core workflows of remote access and transfer.
The surface covers the essential lifecycle: list hosts, probe, policy check, run, run_many, and SFTP upload/download/list. Minor gaps exist (e.g., no explicit sftp_delete or sftp_mkdir), but agents can work around these by using ssh_run for filesystem operations if permitted. No critical dead ends for the stated purpose.
Maintenance
Related MCP Connectors
Scoped, audited SSH exec, sessions, and SFTP on your saved servers without exposing credentials
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
- emisarOAuthdev.emisar
Let AI operate servers without SSH. Choose actions, approve risky changes, and audit every step.
Remote shell and detached long-running jobs on your own machines — no SSH, open ports or VPN.
Related MCP Servers
- AlicenseBqualityAmaintenanceProvides policy-driven, auditable SSH access to server fleets for AI assistants with zero-trust security controls, command whitelisting, and comprehensive audit logging to safely manage infrastructure.1328Apache 2.0
- AlicenseAqualityCmaintenanceEnables AI assistants to securely execute remote SSH commands, perform file transfers, and monitor system status through a standardized interface. It features robust security controls including command whitelisting, blacklisting, and credential isolation to prevent unauthorized operations.1017 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to securely execute SSH commands on remote servers with connection pooling, session isolation, and a web audit panel.4MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to securely execute commands on remote hosts via SSH and SFTP, with persistent shells, file transfers, screenshots, and an audit log.4MIT