tianshu-mcp
tianshu-mcp is an MCP orchestration layer that dispatches development tasks to external AI agents, then independently verifies the results against a git baseline with fail-closed acceptance and an auto-rework loop.
Dispatch work —
run_tasklaunches an AI agent (Codex, ZCode, TraeWork, Kimi Code, Qoder CN, Open Design, MiniMax Code, or CLI profiles) on a project, returning ataskIdimmediately; supportsmodel,reasoningLevel,mode,planDoc,designSystem,autoVerify,autoFixRounds,dryRun,acceptanceOverride,idempotencyKey, andallowCreateProject.Track async tasks —
query_task(status, progress, log tail, recent events),list_tasks(filter by project/status),wait_taskandwait_any(block until a task or the first of up to 20 tasks reaches a stop point: terminal state orneeds_user).Verify objectively —
verify_taskruns configured command checks (typecheck/lint/test/build) plus programmatic code analysis relative to the pre-work git baseline; fail-closed on zero-test-case green runs, no net changes, or cancellation.Read evidence —
get_task_reportreturns the full acceptance report (.md) for a given round, defaulting to the latest.Close the loop on failure —
rework_taskre-queues a failed/needs_attentiontask withfeedbackand optional structuredrepairHint;continue_taskresumes aneeds_usersession (agent question, login, permission, confirmation) per-agent semantics.Stop tasks —
cancel_taskkills the process tree for CLI agents or clicks stop via CDP for GUI agents, and doubles as a manual confirmation entry for terminal GUI tasks.Visual acceptance —
prepare_visual_baselinegenerates candidate screenshot/import baselines with a digest summary, andapprove_visual_baselinecommits them only after explicit user approval.Inspect the environment —
get_profilesshows agent adapters and executable discovery results.Safety model — tools are tiered as
read(no approval),write(approval required), orexecute(verify_task, no approval); no credentials are stored or forwarded, no auto commit/stash/rollback, commands run withshell:false, and path gates reject system/root directories.
Enables verification of agent changes against a Git baseline, including command checks, code analysis, diffstat, and fail-closed detection of missing net changes.
Provides tools for dispatching and controlling OpenAI Codex desktop agent tasks via CDP, including task execution, monitoring, cancellation, and handling login/confirmation states.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@tianshu-mcpHave Codex implement the login page in D:/my-app, verify and rework up to 2 rounds"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
面向 AI-Agent 的编排层与客观验收仪
tianshu-mcp 是被 天枢(Tianshu) 当作标准 MCP server 接入的编排层。天枢是总指挥与用户交互面,本 server 承担三件事:调度(队列 / 并发闸 / 状态机 / 取消)、执行面(把任务书送达外部 AI-Agent)、客观验收仪(相对 git 基线做命令检查、代码分析与可选视觉比对)。
它要回答的核心问题是:Agent 说「做完了」,谁来证明真的做完了。 为此「完成」必须有运行时证据,验收不合格自动生成修复计划并返修,轮次耗尽则交由天枢裁决。
天枢 Tianshu(TUI × GUI) ← 总指挥 / 交互面 / 裁决
↓ MCP over stdio(stdout 仅承载 JSON-RPC)
tianshu-mcp ← 调度 · 执行面 · 验收仪
↓
Codex · TraeWork · ZCode · Kimi Code · Qoder CN · Open Design · MiniMax Code
↓(GUI 经 CDP 驱动桌面 UI;CLI 走子进程)
目标项目工作区 ← git 仓库 + 测试 + .tianshu-mcp/8 个 MCP 工具(v0.9.0 起由 13 个按域合并)——
run_task / query_task / manage_task / verify_task / query_info / wait_task,外加视觉验收的prepare_visual_baseline / approve_visual_baseline。合并映射:cancel_task+continue_task+rework_task→manage_task;list_tasks+get_task_report+get_profiles→query_info;wait_any→wait_task(增强)。异步契约,长任务不卡
tools/call——run_task秒回taskId,用wait_task阻塞等到停点(终态或needs_user)、或用query_task轮询;进度只落盘、不推送,调用方看到的始终是「最后一次落盘的事实」。等待原语(issue #28;v0.9.0 合并) ——
wait_task(taskId)或wait_task(taskIds)一次调用即等到任务到达停点(终态或needs_user),专为回合驱动调用方设计:run_task后在本回合内直接等结果,无需自行轮询;纯只读、超时/中断对任务本体零影响。客观验收,fail-closed —— 自动命令检查 + 程序化代码分析,全部相对动工前的 git 基线,绝不自动 commit / stash / 回滚;「测试退出码 0 但零用例」「git 项目零净变更」都判失败,杜绝假绿。
失败返修闭环 —— 自动返修(
autoFixRounds)+ 手动rework_task;失败原因被解析为可直接执行的动作随计划喂回 agent,轮次耗尽转needs_attention等天枢裁决。六个 GUI 执行面(CDP) —— 各 agent 使用隔离的 CDP 流程驱动桌面 UI,并在关键节点上报细粒度事件,
query_task因此能区分「agent 正在干活」与「卡在弹窗等人工介入」。可扩展 —— 新 agent = 一个 profile(数据)+(如需)一个 adapter 文件,零改编排核心。
本 server 是标准 MCPstdio server:stdout 只承载 MCP JSON-RPC 消息,所有级别诊断日志(DEBUG/INFO/WARN/ERROR)写入 stderr 并同源追加到 <数据目录>/logs/server.log。因此 stderr 里出现 INFO / WARN 不代表服务器出错。数据目录默认 ~/.tianshu-mcp,可用环境变量 TIANSHU_MCP_HOME 覆盖。
目录
Related MCP server: olondunge
项目图解
五幅概念插画,概括本 server 的形态与机制(插画为概念示意,非逐字对应的结构图)。
为什么需要编排层与验收仪
问题:Agent 说「做完了」,谁来证明
把开发任务交给 AI-Agent 之后,真正的困难不在「它能不能干活」,而在怎么确认它真的干完了、干对了:
自述不可信 —— agent 的「已完成」是自然语言结论,不是证据。没有独立验收时,半成品与真交付长得一模一样。
环境是黑盒 —— 桌面 agent 的请求在传输层加密(如 TraeWork 的 TTNet 层 TDE),无法在客户端外构造,唯一可行路径是驱动 UI 取结果。
宿主工具面很窄 —— 天枢的 MCP 工具只回文本(
content[].text被拼成字符串、isError透传),且按次同步调用,长任务必须自己异步化,也不能依赖服务端推送。失败后没人接手 —— 验收不通过时,如果没有人把「哪一行错了、该改什么」喂回去,agent 只会重复同一次错误。
本项目的形态不是自由设计的结果,而是这几条实测硬约束逼出来的:
# | 实测约束 | 架构后果 |
C1 | 天枢的 MCP 工具只回文本 | 所有结果统一为「人类可读文本 + |
C2 | 天枢按次同步调用 | 长任务异步化: |
C3 | 桌面 agent 请求在传输层加密,无法在客户端外构造 | 只能 CDP 驱动桌面 UI,从 DOM 提取结果 |
C4 | Codex 桌面端是 MSIX 商店包,无法直接 | 必须经 COM 激活并注入专属 |
解法:把「谁来干活」与「怎么算干得好」拆开
调度层负责纪律 —— 每项目串行队列 + 全局并发闸(默认 2)、显式状态机、超时与取消语义。
执行面负责投递 —— 一份
AgentAdapter契约:GUI agent 走 CDP 驱动,CLI agent 走子进程;新增 agent 通常只是一个 profile。验收仪负责证据 —— 相对动工前 git 基线做命令检查、代码分析、可选视觉比对,并以 fail-closed 拦住「假绿」;报告分人读
.md与机读.json。返修闭环负责收敛 —— 失败轮次把原因解析成可执行动作喂回同一 agent,轮次耗尽转人工裁决。
两条边界是硬性的:agent 的「完成」不是验收结论(只有
verdict.passed才算);环境 / 认证类错误不进验收与返修(hardFailure直接终态失败,避免把基础设施问题当成代码问题烧掉返修轮次)。
核心特性
异步派单与等待 ——
run_task秒回taskId;wait_task阻塞等到任务到达停点(终态或needs_user),wait_any等一组任务的先到者;需要进度细节时用query_task看状态 / 进度 / 日志尾 / 最近细粒度事件(eventLimit,1..50,默认 10)。详见 等待原语。客观验收引擎 —— 自动命令检查(typecheck/lint/test/build,缺则跳过 + 技术栈推导)+ 程序化代码分析(变更清单 / diffstat / TODO·debugger·密钥形态等可疑标记),全部相对 git 基线;命令默认有界并行(
verifyConcurrency,默认 2,范围 1–4,1即完全串行)。三项 fail-closed 保护 —— 测试退出码为 0 但零用例判失败;git 项目默认要求相对基线产生变更(纯分析任务可在
.tianshu-mcp/acceptance.json设"requireChanges": false显式关闭);本轮被取消即passed=false。验收配置三级继承(issue #20)——
<数据目录>/acceptance.default.json(全局兜底)→<项目>/.tianshu-mcp/acceptance.json(项目覆盖)→acceptanceOverride参数(任务级临时覆盖,不落盘)。用tianshu-mcp config acceptance <projectPath> [--task <id>]查看最终生效配置。详见 验收配置规范。结构化修复指令(issue #19)—— 失败轮次把原因解析为可直接执行的动作(
文件:行 / 问题 / 做什么),随返修计划与返修消息一起喂给 agent;提取不到时显式回退到完整报告(不静默留空)。rework_task另可选repairHint。详见 结构化修复指令。dryRun 干跑模式(issue #21)——
run_task(dryRun=true)让 agent 只分析规划、输出将要修改的文件清单与方案、不动源码;验收只做静态分析(引用文件是否存在、拟改位置是否存在、明显逻辑冲突),跳过 typecheck/test/build;方案有问题 →needs_attention(人工裁决),不进入自动返修、不消耗验收轮次。详见 dryRun 干跑模式。幂等重试(issue #15)——
run_task/verify_task接受可选idempotencyKey:同一 key 在 TTL(默认 24h)内重试不重复派单(恒返回原taskId)或不重跑验收;同键异参 fail-closed 报错。映射落盘于<数据目录>/idempotency.json,跨 server 重启仍生效。详见 v0.5.10 发布说明。无项目派发(ZCode 专用,issue #12)——
run_task的projectPath可省略,任务在 ZCode 的default工作区运行,不登记 / 导入项目、不采集 Git 基线、不执行项目验收(结果以verificationNotApplicable: "no_project"结构化标注)。配套allowCreateProject: false可在目标目录未登记时于任何导入副作用之前停止派发。详见 ZCode CDP 适配器。细粒度事件流(issue #18)—— 适配器在关键节点上报语义事件(
task_dispatched/confirmation_dialog_detected/awaiting_user_authorization/file_modification_started/rework_triggered),query_task经eventLimit回传最近 N 条。事件上报是可选能力:未实现的适配器行为不变。详见 事件流。终态通知 webhook(issue #22)—— 可选
notifications.webhook(全局config.json):任务完成 / 失败 / 进入needs_attention时向指定 URL 异步 POST 一条 JSON(含taskId/event/status/ 时间戳 / 报告路径),可选 HMAC-SHA256 签名。默认关闭,发送失败只记日志、绝不影响状态机。详见 任务终态通知。视觉验收(可选模块,v0.5.0 起) —— 页面截图对比、静态图片规格校验、基准两阶段批准与规则冻结;缺基准不得判通过,自动返修禁止调用批准入口。另有可选 AI 内容校验(v0.5.4,默认关闭):判定完全委托给你自备的本地命令,MCP 不读取 / 不存储 / 不转发任何密钥、不内置模型客户端,默认仅告警。详见 视觉验收。
技能自检安装(issue #16)—— 启动时把包内
skills/tianshu-mcp/幂等同步到~/.rivet/skills/tianshu-mcp/;仅在可证未被改动时自动升级,检出本地修改或来源不明一律保留 + 告警。详见 运行时契约。独立交付面:日志台 GUI ——
mcp-gui/提供本地只读的桌面应用,把四类日志与任务产物统一到一个界面(Tauri 2.x + Vue 3,独立版本与 tag,不随 MCP 主包发布);详见 日志台 GUI。不碰密钥 —— 各 agent 使用自己的登录态,本 server 不保存 / 转发任何 API key(详见 SECURITY.md)。
想理解内部结构 —— 见 ARCHITECTURE.md(分层模型、模块边界、状态机、验收流水线、扩展点与已知缺口)。
快速开始
前置条件
项 | 要求 |
Node.js | ≥ 20(CI 覆盖 20 / 22 / 24) |
包管理器 | npm(仓库含 |
操作系统 | Windows / macOS / Linux(CI 三平台矩阵验证) |
Git | 可选;验收的基线分析在 git 仓库内更完整 |
数据目录默认 ~/.tianshu-mcp,可用环境变量 TIANSHU_MCP_HOME 覆盖;首次启动自动创建。
从源码构建
git clone https://github.com/lanlan0811/tianshu-mcp.git
cd tianshu-mcp
npm ci
npm run build # sync-version + tsc → dist/
npm test # 1383 passed / 12 skipped(1395 项,115 个测试文件)安装 npm 包
npx -y tianshu-mcp # 免安装直接拉起
# 或
npm install -g tianshu-mcp在天枢里添加(推荐)
天枢「设置 → MCP 服务器 → 添加」,按下面填写即可(传输方式选 stdio(本地进程)):
字段 | npm 分发(推荐) | 本地开发 |
服务器 ID |
|
|
传输方式 |
|
|
命令 |
|
|
参数(空格分隔) |
|
|
服务器 ID 即工具前缀:填
tianshu-mcp后工具名为mcp__tianshu-mcp__run_task等 8 个(v0.9.0 起)。参数按空格分隔填写,不要加引号;本地开发模式请把
<仓库绝对路径>换成真实绝对路径。界面未提供环境变量输入框;如需自定义数据目录,改用下面的
config.json方式设置TIANSHU_MCP_HOME。添加后连接成功即完成;新开会话即可看到 8 个工具(v0.9.0 起)。
或改 config.json(可配环境变量)
{
"mcp": {
"servers": {
"tianshu-mcp": {
"command": "node",
"args": ["<仓库绝对路径>/dist/index.js"],
"env": { "TIANSHU_MCP_HOME": "<仓库绝对路径>/.tianshu-mcp" }
}
}
}
}新开会话后,工具面出现 mcp__tianshu-mcp__run_task 等 8 个工具(v0.9.0 起)。一次典型闭环:
run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
model="GPT-5.6 Sol", reasoningLevel="高", autoVerify=true, autoFixRounds=5)
→ taskId → wait_task(taskId) 阻塞等到停点 → succeeded / failed / needs_attention → get_task_report 读报告
(回合驱动调用方:wait_task 一次调用即等到停点;超时返回后再次调用本工具继续等待,或用 query_task 看进度细节)给天枢的提示语(推荐用法)
「在项目
D:\xxx用 codex 实现『任务』。先跑run_task(autoVerify:true, autoFixRounds:2),完成后用wait_task等到停点再看结果;若报告显示needs_attention,把get_task_report的失败项摘要作为feedback调rework_task再验一轮;全部通过后向我汇报changedFiles与diffstat。」
「在项目
D:\xxx用 traework、mode=Code实现『任务』;它会先切到 Code 模式再绑定项目,然后发任务、自动验收,失败自动生成修复计划并返修。」
工具面
8 个工具(v0.9.0 起由 13 个按域合并),按能力分为三族:read(读 / 查询,无副作用)、write(有副作用,全部需审批)、execute(执行项目侧命令但不改源码,当前仅 verify_task,仍免审批)。
工具 | 能力 / 审批 | 作用 |
| write + 审批 | 派活(可带自动验收 / 自动返修),异步返回 |
| read | 轮询状态 / 进度 / 日志尾 / 最近细粒度事件(可选 |
| write + 审批 | 任务生命周期管理, |
| execute(不改源码,免审批) | 对任务 / 项目路径做一次验收。会跑项目配置命令、可能产生构建产物,故 MCP |
| read | 统一信息查询, |
| read | 阻塞等待任务到达停点(终态或 |
| write + 审批 | 截图或导入参考图,生成待审阅候选和摘要 |
| write + 审批 | 用户审阅后校验摘要并写入基准与审批记录 |
返回统一为「人类可读文本 +
---tianshu-mcp-meta---JSON 块」,便于宿主正则抽取。路径安全闸门(v0.4.0 起):
projectPath在提交时校验——必须绝对路径、目录必须存在、符号链接经 realpath 归一;主目录本身与系统 / 根级目录直接拒绝,防止 worker 写权限覆盖整棵系统子树;git 仓库有未提交变更时回执附带共处警示。
支持的 Agent
driver: "gui" 由显式 adapter 驱动桌面 UI(各自使用隔离的 CDP 流程);driver: "spawn" 走外部 CLI 子进程。
agentId | driver / adapter | status | 说明 |
|
| ready(macOS 为 | Codex 桌面端 GUI(Windows:MSIX COM 激活 + CDP;macOS:spawn .app + CDP);支持 |
|
| ready(Windows 真机闭环;macOS 未验证) | CDP GUI adapter;支持无项目派发、 |
|
| ready | CDP 驱动 TRAE SOLO CN 桌面 UI;支持 |
|
| ready(macOS 为 | Kimi Code 桌面端(Electron);双渲染进程(主窗口 + |
|
| Windows 真机闭环通过;macOS research | 仅 Qoder CN;必须提供已有 |
|
| ready(macOS 为 | Open Design 桌面端 GUI;选择器取自产品自身 Web 前端的 |
|
| ready(macOS 为 | MiniMax Code 桌面端(Electron);双渲染进程(主窗口 + |
|
| 仅测试 |
|
mode支持Work/Code/Design(仅 TraeWork),不传时从任务书文本识别。Kimi Code 的reasoningLevel按界面实际渲染的档位集合校验(官方模型低/高/max,非官方模型仅on/off)。MiniMax Code 的reasoningLevel/contextWindow同样按界面实际候选校验(如M3.1-Flash-Preview为default/low/medium/high/xhigh/max与512K/1M,而M3无档位组、deepseek-v4.1-flash无窗口组),越权或读不到即 fail-closed。新增 agent 通常只需加一个 profile,详见 docs/agent-profiles.md 与 CONTRIBUTING.md。
macOS 无头路径:codex-cli(用户 profile)
内置 codex 走桌面端 GUI 驱动;若不想依赖 GUI 自动化,codex CLI 无头模式在 macOS 全程可用——无需改 server 代码,在数据目录加一个 driver=spawn 的用户 profile 即可:
{
"profiles": {
"codex-cli": {
"displayName": "Codex CLI (OpenAI 无头)",
"type": "cli",
"driver": "spawn",
"status": "ready",
"command": null,
"argsTemplate": ["exec", "<prompt:arg>", "--skip-git-repo-check", "--sandbox", "workspace-write"],
"promptMode": "arg",
"cwd": "task",
"env": {},
"timeoutMs": 1800000,
"killTree": "taskkill",
"authNote": "复用 ~/.codex 登录态;勿与 --approve-for-me 同用(实测互斥)",
"executableDiscovery": {
"dirs": ["/opt/homebrew/bin", "/usr/local/bin"],
"fileNames": ["codex"],
"fallbackCommand": "codex"
}
}
}
}前置:
npm i -g @openai/codex(⚠️ 请保持最新,≤0.130.0 签名证书已被吊销,macOS Gatekeeper 会直接Killed: 9)并已codex login。用法与内置 agent 一致:
run_task(projectPath=/path/to/项目, agentId=codex-cli, task="任务书", autoVerify=true, autoFixRounds=2)。model参数对 spawn agent 不生效——CLI 使用~/.codex/config.toml的默认模型;要锁模型可在argsTemplate追加"-m", "<模型名>"。
权限与安全边界
能力三族
能力 | 含义 | 审批 | 工具 |
| 只读 / 查询,无副作用 | 免审批 |
|
| 有副作用 | 需审批 |
|
| 执行项目侧命令,不改源码 | 免审批 |
|
readOnlyHint由capability === "read"推导,因此verify_task的该注解为 false;它不是审批信号——审批与否由_meta.requireApproval单独承载。
硬性红线
绝不按进程树盲杀 GUI 实例 —— 只终止本模块创建、且命令行核对通过的 PID。
默认复用用户实例 —— 绝不新起第二个;受管实例也不触碰用户手动打开的实例。
computer-use 白名单 —— 仅允许 TraeWork 文件夹选择对话框(窗口标题 + 宿主进程双校验)。
凭证零管理 —— 不读取 / 解密 / 转发任何 agent 凭证;GUI adapter 只驱动 UI。
命令不拼 shell —— 验收命令是结构化 argv,
shell:false。不自动 commit / stash / 回滚 —— 动工前采集 git 基线,报告相对基线计算。
路径不硬编码 —— 机器路径 / 用户名 / 端口走 profile 或占位符。
stdout 只承载 JSON-RPC —— 所有诊断日志走 stderr(并同源追加到
logs/server.log)。技能内容只来自包自身 —— 待安装技能经
import.meta.url相对包定位,不从process.cwd()发现内容。
运行时契约
stdio 与日志
本 server 严格遵守 MCP stdio 传输契约:stdout 只承载 JSON-RPC 消息,任何诊断日志都写入 stderr 并同源追加到 <数据目录>/logs/server.log(UTF-8,ISO 时间戳,含级别标签)。排查连接问题时以 server.log 为准;不要因为 stderr 有输出就判定 server 异常。只有启动失败(tianshu-mcp 启动失败:)才是致命错误,并会以非 0 退出码结束。
数据目录
<数据目录>/ 默认 ~/.tianshu-mcp(可用 TIANSHU_MCP_HOME 覆盖)
├── config.json server 配置(并发、超时、技能开关、通知)
├── agent-profiles.json 用户自定义 / 覆盖的 agent profile
├── projects.json 项目登记表(含每项目验收配置)
├── idempotency.json 幂等键映射(TTL + 容量裁剪)
├── logs/server.log 全级别诊断日志(与 stderr 同源)
└── tasks/<taskId>/ 单任务隔离目录(事件流 / 报告 / 日志 / 视觉证据)技能自检安装
启动时把包内 skills/tianshu-mcp/ 幂等同步到 ~/.rivet/skills/tianshu-mcp/,让宿主在新会话里读到编排技能。三个要点:
技能内容只来自包自身 —— 源目录由
import.meta.url相对定位(dev 直跑与 dist 运行都指向包内skills/),不从当前工作目录发现内容。找不到源时跳过安装并告警。不一致时不静默覆盖 —— 安装目录内维护清单
<目标>/.tianshu-mcp-install.json(版本 + 内容 hash),据此仅在可证未被改动时自动升级;检出你改过文件或来源不明 → 默认保留你的版本并告警。覆盖是原子的 —— 先装到
.incoming-*,再备份旧目录为.bak-<时间戳>,最后换入;失败回滚,不留半成品。
// <数据目录>/config.json
{
"skills": {
"autoInstall": true, // true(默认)| "prompt" | false
"backupKeep": 3 // 覆盖后保留的历史备份个数;0 = 不清理
}
}放行与关闭(命令行参数或等价环境变量;--no-skill-install / autoInstall:false 的否决权最高):
--approve-skill-update(或TIANSHU_MCP_APPROVE_SKILL_UPDATE=1):本次启动允许「需变更」的技能目录由包内版本覆盖(先备份)。对已确证含用户本地修改的目录不生效。--no-skill-install(或TIANSHU_MCP_NO_SKILL_INSTALL=1):本次启动不做任何技能安装与检查。
里程碑
阶段 | 版本 | 交付概要 |
编排骨架 | 0.1.x | 8 工具、状态机 / 队列 / 并发闸 / 取消(kill tree)、验收引擎、自动返修;TraeWork CDP 驱动接入与模式切换 |
ZCode GUI | 0.2.0 | ZCode 统一闭环(开发 → 受控失败 → 同会话返修 → |
Codex 桌面端 | 0.3.x | Codex MSIX COM 激活 + CDP(破坏性: |
macOS 与闸门 | 0.4.x | macOS 双驱动(spawn .app + CDP)、 |
视觉与幂等 | 0.5.x | 视觉验收(0.5.0)+ 可选 AI 内容校验(0.5.4)、ZCode 无项目派发、Kimi Code / Qoder CN 适配、幂等键(0.5.10) |
加固与可观测 | 0.6.x | 技能自装加固(0.6.0)、GUI 选择器漂移修复(0.6.2)、细粒度事件流、结构化修复指令、dryRun、验收配置三级继承、终态通知 |
Open Design | 0.7.x | Open Design 桌面端适配(0.7.1)、ZCode 3.14.x 绑定契约修复(0.7.4) |
MiniMax Code | 0.7.8 | 第七个 GUI agent 接入(0.7.8);真机取证修正三处结构假设(二级子菜单 / 集合随模型变化 / 项目创建两步),新增 |
恢复语义修正 | 0.8.0 | TraeWork 移除跨模式项目绑定兜底(#35:兜底结构性不可达且静默改写目标模式);ZCode 恢复轮保留原会话权限(#30:发送前无条件覆盖默认值已移除) |
完成判定加固 | 0.8.1 | 四个 driver(ZCode / Kimi Code / MiniMax Code / Open Design)补上「曾观测到运行信号」门(#31:选择器漂移时不再把进行中的任务误判成功); |
Codex 模型回读 | 0.8.2 | 模型触发器回读改读结构(#34:真机按钮的 |
TraeWork 绑定链路 | 0.8.3 | footer 点击改「副作用驱动三级阶梯」(#38: |
工具面瘦身 | 0.8.4 |
|
日志台 GUI |
|
|
完整逐版记录见 CHANGELOG.md,交接状态与排障手册见 HANDOFF.md,工程质量口径见 ARCHITECTURE.md。
日志台 GUI
mcp-gui/ 是本仓库的第二个交付面(issue #25):一个本地只读的桌面应用(Tauri 2.x + Vue 3 + Vite + TypeScript),把 MCP 落盘的日志与任务产物统一到一个界面里查看。它与 MCP server 的关系只有一条——共享同一批落盘事实,不产生第二个事实来源:
不依赖 server 在运行 —— 纯读文件系统,数据目录按与 server 完全相同的规则解析(
TIANSHU_MCP_HOME→~/.tianshu-mcp),并可在多个数据目录之间切换 / 追加 / 移除。只读消费方 —— 全程不改动任何业务数据(唯一写入是应用自身偏好,落在系统应用配置目录),也不替代面向机器的
query_task/get_task_report。四类日志与产物 —— 全局运行日志、任务事件流、原始执行日志、验收报告;视觉离线 HTML 在 sandbox iframe 中渲染(禁用脚本、阻断外部资源)。
数据源 | 路径(相对数据目录) | 界面位置 |
全局运行日志 |
| 工作区 · 运行日志 |
任务事件流 |
| 工作区 · 事件流 |
原始执行日志 |
| 工作区 · Agent 日志 / 验收日志 |
验收报告 |
| 工作区 · 验收报告 |
主要能力:
大日志与实时跟随 —— 首屏只读尾部 64 KiB 窗口、向前按块加载并显示「已加载 N / 共 M」;文件被追加时增量刷新,上翻自动暂停跟随,可一键「跳到最新」。
洞察(效能 / 归因 / 趋势) —— 只读聚合:按 Agent 与按项目的效能看板(任务数 / 成功率 / 平均轮次 / 一次通过率 / 平均验收耗时 / 报告缺失)、四类失败归因 TOP 列表(
errorType/ 失败检查项 / 阻塞问题 / 代码信号)、按天 / 按周的任务量与成功率、返修率趋势(纯内联 SVG,不引图表库)。口径显式标注(UTC 日期、周一为周始、「一次通过」= 成功且仅 1 轮、只统计每个任务的最新一轮报告),只读统计、不提供删除 / 清理。结构化筛选 / 多任务对比 / 命令面板 —— 概览页筛选新增错误类型 / 干跑 / 返修 / 视觉验收四项(口径与任务快照字段一一对应,前后端同口径);洞察页新增「任务对比」子分区,勾选 2–4 个 任务并排看状态 / 轮次 / 验收耗时 / 改动行数 / 最新报告判定等指标(报告按需读取并缓存,缺失一律显示
—,不编造);Ctrl/Cmd + K打开命令面板(子序列模糊匹配,跳页面 / 切数据目录 / 直接打开任务),Ctrl/Cmd + R刷新,全部为只读交互。基线与复盘 / 阶段甘特 / 磁盘占用 / 深链 —— 工作区新增**「基线」分区(动工前
baseline.json摘要并与最新报告改动对照,缺失如实提示);事件流新增「阶段」视图**(状态跃迁甘特,按需读一次全量,最后一段标「进行中」不编造时长);洞察页新增**「磁盘占用」**(总量 / logs 占比 / 体积 TOP 20 / 只提示不删除的「可清理」相对判据);支持tianshu://task/<任务ID>深链(冷启动 + 热启动、单实例唤出已有窗口,Rust 侧处理故不给 webview 多余权限)。报告与多轮对比 ——
.md渲染、.json结构化卡片、视觉.html沙箱预览;dry-run-report-*与report-*分开展示(静态分析 vs 真实命令验收,结论口径不同),多轮报告可并排对比。界面与主题 —— 中英双语、跟随系统 / 浅色 / 深色三选一;自研「黑曜石终端」设计系统,零 UI 库、零外链、零字体文件,图标一律内联 SVG。
系统托盘与关闭行为 —— 常驻托盘(「显示日志台 / 退出日志台」,文案随界面语言即时切换),默认 关闭窗口 = 缩小到托盘,可在设置面板改为「关闭应用」。
更新日志窗口与双源自动更新 —— 启动静默检查更新,命中即弹「更新日志」(下载并安装 / 忽略此版本 / 稍后),正文即该版本的双语发行说明,与发行页同源同一份(正文缺失即拒绝发版,不产出空正文 / 单行标题);更新源由 Gitee / GitHub 并发实测择优(不依赖系统区域)决定并如实展示,包体经 minisign 验签,验签不通过一律拒绝安装。
解耦与发布边界(改这里之前先读):
边界 | 约定 |
数据 | GUI 只读业务目录;唯一写入是应用自身偏好与用户显式选择的导出 / 更新文件 |
代码 |
|
打包 | 根 |
发版 | GUI 独立版本号与独立 tag( |
构建 | 本机不执行 Rust 侧构建与检查( |
双份 schema 的防漂移:事件分类在 Rust 侧与前端各有一份镜像,真源始终是
src/tasks/task.ts与src/agents/agent-events.ts;mcp-gui/scripts/check-schema-parity.mjs在 CI 中做三方集合比对,任一不一致即 fail。使用与开发说明见 日志台文档,真机记录见 issue-25 记录。
文档导航
使用与集成
文档 | 说明 |
架构说明:分层模型与模块边界、状态机、验收流水线、驱动层契约、扩展点与已知缺口 | |
核心原理分析:四条硬约束如何逼出当前架构、核心机制逐条拆解与自洽性总结 | |
天枢 config.json 两种接入模式、UI / API 操作、冒烟步骤、FAQ | |
agent profile 字段说明 + 真实机器样例 | |
各 Agent 能力调研矩阵 | |
npm 发布步骤与凭证说明 | |
教天枢编排本 MCP 的技能(含使用示例) |
Agent 适配(CDP 驱动)
文档 | 说明 |
Codex 桌面端:MSIX COM 激活、CDP 接管、选择器、运行检测、验收返修 | |
TraeWork:原理、配置、模式切换、选择器、安全红线、踩坑记录 | |
ZCode:安装探测、精确项目 / 模型、完全访问、暂停继续、无项目派发与双平台状态 | |
Kimi Code:双渲染进程、工作区完整路径绑定与原生导入、模型三级选择与思考档位 | |
Qoder CN:安装发现与实例复用、工作区原生导入、 | |
Open Design:数据目录推导、sidecar 根进程判定、选择器取证表与 12 步执行链、传输层双路径、失败码表 | |
docs/minimax-cdp.md | MiniMax Code:双渲染进程、模型二级子菜单(推理等级 / 上下文窗口)与逐模型候选、完整路径项目绑定、 |
验收与可观测
文档 | 说明 |
项目级与三级继承的验收配置规范 | |
结构化修复指令:来源、回退语义与已知限制 | |
dryRun 干跑模式:只读约束、零改动门禁、方案文档 | |
细粒度事件流:词表、落盘与读取侧有界窗口 | |
docs/notifications.md | 任务终态通知:webhook 契约、去重与签名 |
docs/wait-task.md | 等待原语: |
视觉验收入门与完整配置(含可选 AI 内容校验) | |
视觉验收验证进度与平台证据 | |
上述验证的原始机器可读记录 | |
日志台 GUI( |
真机验收记录
文档 | 说明 |
M2 真实 codex 冒烟与 rework 闭环记录 | |
docs/zcode-windows-smoke.md · docs/zcode-issue-8-10-validation.md · docs/zcode-issue-12-windows-evidence.md | ZCode Windows 真机验收记录 |
Codex Windows 真机验收记录(含失败 → 自动生成计划 → 返修通过) | |
docs/host-integration-record.md · docs/issue-1-host-reconnect-record.md | 天枢宿主真实接入与重连验收 |
docs/dod7-release-record.md · docs/dod8-session-record.md · docs/s7-session-recheck.md | npm 发布、真实会话实测与二次整改复测 |
docs/issue-16-skill-install-hardening-record.md · docs/issue-17-small-fixes-record.md · docs/issue-23-selector-drift-record.md | 技能自装加固、小项扫尾、选择器漂移记录 |
docs/issue-18-21-real-machine-record.md · docs/issue-19-22-real-machine-record.md | 事件流 / 修复指令 / dryRun / 验收继承 / 通知的真机记录 |
docs/issue-25-gui-real-machine-record.md · docs/gui-0.1.0-release-record.md | 日志台 GUI 真机验收与正式版发布记录 |
面向开发者
Node.js ≥ 20 · TypeScript 5.7 · Vitest · tsup-free(tsc 直出 dist/)+ tsx 开发。
npm ci
npm run build # sync-version + tsc → dist/
npm test # 全量用例
npm run typecheck # 类型检查(tsc --noEmit)
npm run lint # ESLint(--max-warnings 0)
npm run check:stdio # 严格 stdio 冒烟(真实进程字节流校验)新增 CLI agent —— 通常只需在
<数据目录>/agent-profiles.json加一个driver: "spawn"的 profile,零改代码。新增 GUI agent —— 新写一个 adapter 目录(
adapter.ts/discovery.ts/cdp.ts/selectors.ts/project.ts/liveness.ts/run.ts),并在agents/registry.ts与agents/builtin.ts注册。新增 MCP 工具 ——
src/mcp/tools.ts增元数据 +src/mcp/handlers.ts增实现 +src/config/schema.ts增入参 schema。调 UI 选择器 —— profile
gui.selectors覆盖(客户端升级导致选择器漂移时,先用npm run probe:*诊断)。版本号三处必须同步:
package.json、package-lock.json、src/version.generated.ts(后者由scripts/sync-version.mjs在 build 前生成,勿手改)。
安全
路径边界强制 ——
projectPath经 realpath 归一;主目录与系统 / 根级目录子树拒绝;glob/grep/diff 拒绝..穿越。不自动改动仓库历史 —— 动工前采集 git 基线,报告相对基线计算;MCP 从不自动 commit / stash / checkout。
凭证零管理 —— 不读取 / 解密 / 转发任何 agent 凭证;AI 内容校验同样不引入凭证管理——判定命令自己管密钥。
命令不拼 shell —— 验收命令是结构化 argv,
shell:false,无 shell 注入面。桌面自动化边界 —— 默认复用用户实例、computer-use 白名单、归属核对后才终止进程。
安全漏洞请按 SECURITY.md 私密报告,不要开公开 Issue。
社区与支持
使用问题 / 讨论 → GitHub Issues(附
logs/server.log输出可加速定位)主仓库 → https://github.com/lanlan0811/tianshu-mcp(GitHub)
镜像仓库 → https://gitee.com/lan0811/tianshu-mcp(Gitee)
贡献代码 → CONTRIBUTING.md · 安全模型 → SECURITY.md · 行为准则 → CODE_OF_CONDUCT.md
交接状态 / 排障手册 → HANDOFF.md · 版本变更 → CHANGELOG.md · 依赖清单 → DEPENDENCIES.md
贡献者
感谢以下通过 Issue 与 PR 为本项目做出贡献的社区成员(按首次参与顺序排列):
Star History
许可证
本项目以 Apache License 2.0 发布,完整法律文本见 LICENSE。版权归 tianshu-mcp 贡献者所有(Copyright 2026 tianshu-mcp contributors)。简言之:你可以商业使用、修改、分发与私用,并获授贡献者专利许可;分发时须随附 LICENSE 全文并标注修改;本许可不授予商标使用权,对贡献者发起专利诉讼将导致专利授权自动终止;软件按「现状」提供,不附带任何担保。
第三方依赖许可
运行时依赖的许可如下(完整依赖清单——逐项版本、开发依赖、桌面端 Rust 依赖、间接依赖许可证分布与 SBOM 复现命令——见 DEPENDENCIES.md):
依赖 | 许可 | 用途 |
MIT | MCP 协议实现 | |
MIT | 外部输入校验 | |
MIT | 跨平台子进程 | |
Apache-2.0 | 视觉验收驱动无头浏览器 | |
Apache-2.0 | 托管 Chrome/Edge 的安装与版本锁定 | |
ISC | 页面截图像素比对 | |
| Apache-2.0 | 图片解码与规格校验;缺失时视觉模块明确阻塞 |
sharp本体为 Apache-2.0,但其可选平台二进制(@img/sharp-*)声明为 LGPL-3.0-or-later,以未修改的预编译动态库使用——不安装sharp时依赖树中不含任何 LGPL 组件。
开发依赖(TypeScript、ESLint、Prettier、Vitest、Vite、tsx 等)各自遵循其开源许可,且不随 npm 发布产物分发。
与安全边界的关系
本 MCP 不保存、不读取、不转发任何 AI-Agent 的 API key 或登录态(详见 SECURITY.md)。许可条款不改变这一设计边界。
英文文档见 README.en.md 与 ARCHITECTURE.en.md;完整文档地图与状态快照见 HANDOFF.md。
Available Tools
13 toolsapprove_visual_baselineapprove_visual_baselineADestructive
仅在用户明确审阅并授权后批准视觉基准。必须核对候选摘要与批准说明;自动返修禁止调用。宿主必须实施实际审批控制。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | No | ||
| candidateId | Yes | ||
| approvalNote | Yes | ||
| expectedDigest | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, so the mutation risk is known. The description adds critical behavioral context: this is a human-gated approval action, automatic rework must not call it, and the host must enforce real approval. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all substantive and front-loaded with the core approval condition. The second sentence adds a safety constraint and the third adds an implementation requirement. No wasted words, though the Chinese phrasing is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, human-gated approval tool with no output schema, the description covers the critical safety context and usage constraints. However, it does not explain what happens after approval, what the expectedDigest is for, or how the candidate summary should be obtained, leaving some operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, but it only explains the approvalNote verification concept and does not explain candidateId, expectedDigest, or taskId semantics beyond what the schema's names and formats imply. The description adds some context about the approval note but leaves parameter meaning mostly to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: approve a visual baseline only after explicit user review and authorization. It clearly distinguishes this from preparing a baseline or other task operations, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage conditions: only after user review and authorization, must verify candidate summary against approval note, and explicitly forbids automatic rework invocation. It also states the host must implement actual approval control, which is strong when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_taskcancel_taskADestructive
取消运行中任务:CLI agent 终止进程树;GUI agent(codex 等)尽力点击界面停止按钮并等待 GUI 空闲(有界超时),未确认停止时结果中明示。排队中任务直接移除。对已处于终态的 GUI 任务,本调用兼任人工确认入口:人工核实窗口中已无残留运行后调用,可清除 meta 的 guiStopUnconfirmed 待确认标记(不改终态)。
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral detail beyond the destructiveHint annotation: CLI agent terminates the process tree, GUI agent attempts to click stop and waits with a bounded timeout, queued tasks are simply removed, unconfirmed stops are surfaced in results, and terminal-state calls clear the guiStopUnconfirmed flag without changing final state. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: it front-loads the core action, then efficiently covers CLI behavior, GUI behavior, queued tasks, and the terminal-state confirmation role without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with this many behavioral branches, the description is unusually complete: it covers CLI vs GUI handling, queued removal, timeout expectations, unconfirmed-stop reporting, and the confirmation semantics for terminal GUI tasks. The remaining gap is parameter-level guidance, and there is no output schema, but the behavior itself is sufficiently specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it never explains taskId or reason. taskId is somewhat inferable from context, but reason's purpose is entirely undocumented in both the description and the schema, leaving the agent to guess whether and how it affects cancellation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair, “取消运行中任务,” and then disambiguates by execution mode (CLI vs GUI), queue state, and terminal-state confirmation. This clearly separates it from siblings such as run_task, continue_task, query_task, and rework_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly enumerates when to call the tool: for running CLI/GUI tasks, for queued tasks, and as the manual confirmation entry for terminal GUI tasks with guiStopUnconfirmed. It does not explicitly name alternatives or state when not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
continue_taskcontinue_taskA
恢复处于 needs_user 的任务。zcode:agent_question 时 message 发往原会话,关闭旧实例/登录/系统权限场景中 message 仅作已处理确认。codex:user_confirmation 时重新接入观察 GUI 内运行(不发送消息);login_required 时复检环境后重发任务书。qoder:Agent 提问通过专用答题控件回复;多题 message 使用完整问题文字到答案的 JSON 对象。审批或环境处理后仅恢复观察,提交不明时禁止重发。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| message | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only set readOnlyHint=false, destructiveHint=false, etc., which convey a non-mutating, non-destructive profile. The description adds significant behavioral detail: for codex user_confirmation it says '不发送消息' (does not send messages), for login_required it says to recheck environment and resend the task, and for qoder it describes using a dedicated control. It also states that after approval/environment handling only observation is resumed. This goes well beyond the flat annotations, though it leaves some outcomes unspecified (e.g., return values).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured in a scannable list format per agent and condition, with each clause delivering a distinct fact. The opening phrase is a clear verb+resource statement, and the rest is organized into conditional blocks. Although it is longer than a one-sentence blurb, every sentence earns its place by conveying unique scenario-specific behavior. It is not repetitive and front-loads the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—handling multiple agent types and distinct scenarios—the description covers the major behavioral branches (zcode, codex, qoder) and the environmental conditions under which they apply. It also expresses constraints like '审批或环境处理后仅恢复观察' (only resume observation after approval/environment processing). Missing is any explicit statement about the output or success criteria, but since there is no output schema and the safety profile is already captured by annotations, the description covers the essential usage context well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the meaning of the parameters. It explains that 'message' is sent to the original session for zcode agent_question, serves as a processed confirmation in certain scenarios, is not sent for codex user_confirmation, and is used for replies via a control for qoder. It also gives a format hint for multi-question messages (JSON object from full question text to answer). However, 'taskId' is never explicitly defined beyond the implicit notion of identifying a task, and the exact JSON structure is not fully specified. This adds some meaning but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete action and resource: '恢复处于 needs_user 的任务' (resume tasks in the needs_user state), clearly distinguishing it from sibling tools like query_task or run_task. It further refines the purpose by enumerating agent-specific behaviors (zcode, codex, qoder), which unambiguously defines what the tool does and when it applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit context for when to use the tool (tasks in needs_user state) and gives scenario-specific instructions per agent (agent_question, user_confirmation, login_required). While it does not name alternative tools, it clearly states a prohibition: '提交不明时禁止重发' (do not resend when submission is unclear), which helps prevent misuse. This is strong guidance but stops short of explicit when-not-to-use comparisons against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_profilesget_profilesARead-only
查看当前 agent 适配与可执行探测结果(含未安装/调研占位提示)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and destructiveHint=false, so the read-only nature is known. The description adds context that the output may include placeholder prompts for uninstalled/research items, which is useful but not extensive. No deeper behavioral details (e.g., caching or side effects) are disclosed, but that is acceptable for a simple read getter with read-only annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no fluff. It conveys the core purpose and a relevant caveat in minimal words, earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only tool, the description covers the essential return subject ('agent adaptation and executable detection results') and notes the placeholder edge case. There is no output schema, so the description reasonably carries the explanatory burden, and it does so adequately. It could be slightly more explicit about the exact shape or granularity of results, but it is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema coverage is trivially 100% and there is nothing for the description to add. Per the baseline rule for 0-param tools, a score of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('view') and a specific resource ('current agent adaptation and executable detection results'), and adds a clarifying parenthetical about placeholder prompts. This clearly differentiates it from sibling task-management tools, which all deal with running or reporting on tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description only states what the tool does, not when it should be used. It provides no explicit context, prerequisites, or alternatives, and does not reference any sibling tools. An agent would have to infer when to call get_profiles versus get_task_report or other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_reportget_task_reportARead-only
取某轮验收报告全文(report.md)。round 缺省取最新一轮。
| Name | Required | Description | Default |
|---|---|---|---|
| round | No | ||
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds useful behavior beyond annotations: it returns the full text of report.md and defaults round to the latest round when omitted. It doesn't claim any write side effects and doesn't contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler; the core action and the only non-obvious default are both stated. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple read-only, two-parameter fetch with annotations covering safety. It states what is returned (full report.md content) and the default round behavior; no output schema exists, so that statement matters. Missing minor context such as behavior when no report exists or when task has no acceptance report, but not essential for a basic read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It explains round semantics and its default ('缺省取最新一轮') and clarifies the output is the report's full text. taskId is not described in prose, but the required parameter and tool name make its role obvious; schema supplies type constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('取' / get), a specific resource (acceptance report / report.md), and parameter scope (by round), making the tool's main function clear. It doesn't explicitly distinguish itself from sibling tools like query_task or list_tasks, but the resource is unique enough that the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: use when you need the full acceptance report for a given or latest round. The optional-round behavior and default are stated, but no alternatives or when-not-to-use conditions are given, so it falls short of explicit sibling routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_taskslist_tasksARead-only
列出历史任务(可按项目路径 / 状态过滤,limit 默认 50)。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | ||
| projectPath | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the default limit of 50 and the filtering capabilities, which are useful behavioral details. However, it does not disclose pagination behavior, ordering, or what happens when no filters are provided, which would add further transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the core purpose (list historical tasks) and then adds the key filtering and limit details. Every word earns its place, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no output schema, the description covers the main purpose and key parameters. However, it lacks details on return format, ordering, pagination, and how the status filter values are specified. Given the tool's simplicity and the annotations covering safety, this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. The description mentions filtering by project path and status, which maps to the 'projectPath' and 'status' parameters, and mentions the default limit of 50, which maps to 'limit'. However, it does not explain the format of 'status' values or the exact semantics of 'projectPath', leaving some ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('列出' = list) and resource ('历史任务' = historical tasks), and mentions filtering by project path/status and a default limit of 50. It is clear about what the tool does, though it doesn't explicitly distinguish it from sibling tools like query_task or get_task_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: listing historical tasks with optional filters. It does not explicitly state when to use this tool versus alternatives like query_task or get_task_report, nor does it mention any exclusions or prerequisites. The filtering options give some context, but no explicit routing guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_visual_baselineprepare_visual_baselineA
准备视觉基准候选,返回摘要与预览;不采用正式基准。需要用户授权。
| Name | Required | Description | Default |
|---|---|---|---|
| caseIds | No | ||
| imports | No | ||
| projectPath | Yes | ||
| viewportIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide no safety hints (all false), so the description carries the burden. It adds that the tool does not formally adopt a baseline and requires user authorization, which clarifies the side-effect profile beyond the annotations. It also mentions the return of a summary and preview, giving a partial picture of behavior, though it does not detail any state changes or object creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with no fluff; the core action is front-loaded and followed by key behavioral caveats. Every clause adds information. It is appropriately sized for the content it covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no output schema, and no annotation hints, the description leaves significant gaps: no parameter explanations, no usage guidance versus siblings, and no description of the returned summary/preview structure. It is enough to hint at purpose but not enough for an agent to correctly populate the complex inputs. The tool appears non-trivial (imports array of objects) and needs richer context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the burden falls entirely on the description, but it mentions no parameters at all. The text does not explain the roles of projectPath, caseIds, viewportIds, or imports or how they relate to the prepared baseline candidates. An agent must guess parameter semantics from property names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'prepare visual baseline candidates' and its immediate outputs ('returns summary and preview'). It also distinguishes itself from the sibling approve_visual_baseline by noting it does not adopt a formal baseline. This gives an agent a specific verb-object pair and a differentiating constraint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a pre-approval step by stating it does not adopt a formal baseline, and it notes user authorization is required. However, it never explicitly states when to use this tool versus approve_visual_baseline or the other task siblings, nor gives any exclusion criteria. The usage context is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_taskquery_taskARead-only
查询任务状态 / 进度 / 最近日志尾部(默认 agent.log 末 40 行)/ 最近细粒度事件。返回任务 meta 与日志片段。meta.recentEvents 为最近 N 条 agent 事件(eventLimit 缺省 10、上限 50),取值 task_dispatched / confirmation_dialog_detected / awaiting_user_authorization / file_modification_started / rework_triggered —— 长任务下可据此区分「正常执行」与「卡在弹窗等人」。未实现事件上报的适配器该数组为空,其余字段不变。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| tailLines | No | ||
| eventLimit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description reveals important behavioral details: the log tail defaults to the last 40 lines of agent.log, eventLimit defaults to 10 with a max of 50, and adapters that don't report events return an empty array while other fields remain unchanged. It also lists the exact event enum values, giving the agent concrete expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph, but it front-loads the core purpose and then packs defaults, event values, and adapter fallback behavior into later clauses. Some run-on punctuation makes it slightly harder to scan, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explains the return shape (task meta + log snippet), details the recentEvents field and its adapter-dependent empty-array behavior, and provides interpretation guidance for long-running tasks. It doesn't describe other meta fields, log formatting, or error conditions, but it is sufficiently complete for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It documents eventLimit's default and maximum, tailLines via the default 'agent.log 末 40 行', and even explains the values returned in meta.recentEvents. taskId is left implicit, but its meaning is obvious from the tool name and required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb and resource set: '查询任务状态 / 进度 / 最近日志尾部 / 最近细粒度事件' (query task status / progress / recent log tail / recent fine-grained events). It clearly positions the tool as a read-only inspection tool distinct from siblings like run_task or verify_task by also stating it returns task meta and log snippets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete use context: '长任务下可据此区分「正常执行」与「卡在弹窗等人」' (under long tasks, distinguish normal execution from being stuck waiting on a dialog). It does not explicitly name alternatives or say when not to use it, but the context is clear enough for an agent to select it for status/progress/log inspection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rework_taskrework_taskADestructive
手动返修:把终态任务(failed/needs_attention)重新入队续跑,同一 agent/项目与轮次记账。feedback 为追加指示(建议带上一次验收失败摘要)。repairHint 为可选的结构化修复提示(自由字符串,上限 4000 字符)——写「文件:行 / 问题 / 做什么」,会以【结构化修复提示】块置于 feedback 之前,便于 agent 先精确定位再读整段说明;不传则行为不变。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| feedback | No | ||
| repairHint | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses the re-queue/continue behavior, preservation of agent/project/round accounting, feedback append semantics, and repairHint placement before feedback. It even clarifies that omitting repairHint leaves behavior unchanged, which is valuable side-effect transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core action before parameter details. It slightly repeats schema facts such as maxLength=4000 and the free-string type, but there is no wasted or misleading content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no output schema, the description covers the target state, parameters, and behavioral consequences well. It does not describe return values or what happens if the task is not in a terminal state, but the essential invocation information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description compensates well: feedback is defined as an appended instruction, and repairHint is explained as an optional structured hint with format, limit, and placement. taskId's role is clear from context and required status, so no parameter is left ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: manually rework terminal tasks (failed/needs_attention) by re-queueing them to continue running, including same agent/project and round accounting. It clearly defines scope, but does not explicitly name or contrast sibling tools such as continue_task or run_task, so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool: for terminal-state tasks that need manual rework. It does not provide exclusions or name alternatives, so an agent gets a clear trigger but no explicit guidance about when a sibling would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_taskrun_taskAIdempotent
派活:启动外部 AI-Agent 开发任务并可自动验收返修,异步返回 taskId。ZCode 要求 model=供应商/模型,不支持 mode;TraeWork 的 model 可选并支持 Work/Code/Design mode。task/context 内的 ZCode 项目路径引用会在发送前校验。Qoder CN 要求已有 projectPath 和可读 planDoc;modelSource 可选 default/custom,省略模型或等级则沿用当前设置。思考等级通过模型管理保存为全局偏好,权限模式不变;macOS research 禁止派发。可选 idempotencyKey(1..128 字符):同一 key 在 TTL(默认 24h)内重复提交恒返回原 taskId 与当前状态、不新建任务,参数变更则报冲突——重试请复用同一 key。 可选 dryRun=true 进入干跑模式(先审后做):agent 只分析规划、输出将要修改的文件清单与方案、不动源码;验收引擎只做静态分析(引用文件是否存在、拟改位置是否存在、明显逻辑冲突),跳过 typecheck/test/build。dryRun 需提供 projectPath、忽略 autoVerify、不进入自动返修;产物为独立报告(meta.dryRunReportFiles,不消耗验收轮次)与方案文档(meta.dryRunPlanDoc,可直接作为后续正式任务的 planDoc)。默认关闭。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| task | Yes | ||
| model | No | ||
| dryRun | No | ||
| agentId | No | ||
| context | No | ||
| planDoc | No | ||
| autoVerify | No | ||
| modelSource | No | ||
| projectPath | No | ||
| designSystem | No | ||
| autoFixRounds | No | ||
| taskTimeoutMs | No | ||
| idempotencyKey | No | ||
| reasoningLevel | No | ||
| designDirection | No | ||
| acceptanceOverride | No | ||
| allowCreateProject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorld/write/idempotent/non-destructive; the description adds real behavioral detail beyond them: idempotency key TTL of 24h, conflict on changed parameters, dryRun's no-source-change contract and static-only acceptance skipping typecheck/test/build, and a global side effect (reasoning level persisted as a global preference while permission mode is unchanged). It does not describe return payload shape or failure modes, keeping it at 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and return contract, then provider-specific rules, then the optional idempotencyKey and dryRun behavior. It is one dense block with no headings or bullets, which makes the long provider-rule sentence hard to scan, but every clause carries operational information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter tool with a deeply nested acceptanceOverride and no output schema, the description covers the primary flow, provider prerequisites and dry-run/idempotency semantics well. It nevertheless leaves several top-level parameters (agentId, designSystem, designDirection, autoFixRounds, taskTimeoutMs, allowCreateProject) and the entire acceptance-override mechanism unexplained, and no output schema exists to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it explains task/context, model, mode, projectPath, planDoc, modelSource, reasoningLevel, idempotencyKey, dryRun and partially autoVerify. However it says nothing about agentId, designSystem, designDirection, autoFixRounds, taskTimeoutMs, allowCreateProject, or the large nested acceptanceOverride object, so a substantial share of the 18 parameters remains undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"启动外部 AI-Agent 开发任务并可自动验收返修,异步返回 taskId" states a specific verb (start a dev task), the resource (external AI-agent task), the async return contract (taskId), and the post-processing behavior (auto verification/rework). An agent can distinguish this from siblings like cancel_task, verify_task, rework_task and continue_task purely from the text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use conditions per provider family (ZCode requires model=provider/model and no mode; TraeWork model optional with Work/Code/Design; Qoder CN requires projectPath plus readable planDoc), plus when to set dryRun and when to reuse idempotencyKey. It stops short of naming sibling tools as alternatives (e.g. use verify_task vs relying on autoVerify), so it is strong context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_taskverify_taskAIdempotent
对已完成任务或项目路径执行一次验收(不改源码):自动命令检查 + 代码分析(相对 git 基线)。可用 extraChecks 临时加验。需任务/项目二选一。可选 idempotencyKey:同一 key 重试不重跑验收——执行中的同键请求返回进行中提示,已完成的直接返回既有报告与轮次,参数变更则报冲突。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | No | ||
| checksMode | No | ||
| baselineRef | No | ||
| extraChecks | No | ||
| projectPath | No | ||
| idempotencyKey | No | ||
| acceptanceOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral detail beyond the annotations: it guarantees no source changes, explains idempotency semantics in detail (same key returns in-progress status, completed reruns return existing reports, parameter changes cause conflicts), and mentions the relative git baseline. This aligns with idempotentHint=true and destructiveHint=false, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, front-loading the core purpose and safety guarantee before moving to parameter constraints and idempotency details. It is somewhat run-on but every clause carries useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex, with no output schema and a very large nested acceptanceOverride parameter, yet the description provides no information about return values, report structure, or the meaning of verification outcomes. It also does not explain checksMode or how to use acceptanceOverride, leaving significant gaps for an agent deciding how to configure a verification run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains taskId/projectPath mutual exclusivity, extraChecks, baselineRef via '相对 git 基线', and idempotencyKey behavior. However, it omits key parameters like checksMode and the large acceptanceOverride object, leaving important semantic gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs acceptance verification ('执行一次验收') on completed tasks or project paths, using automatic command checks and code analysis relative to a git baseline, and explicitly notes it does not modify source code. This distinguishes it from siblings like run_task, rework_task, and get_task_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies when to use it: for completed tasks or project paths, and it states the mandatory selection between task and project ('需任务/项目二选一'). It also explains how extraChecks can augment verification. However, it does not explicitly name alternatives or exclusions relative to siblings such as query_task or get_task_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_anywait_anyBRead-only
等待一组任务中首个到达停点(终态或 needs_user)的任务;返回该任务快照与全部任务当前状态。taskIds 1..20 个,开始前校验全部存在,缺一即报错。
| Name | Required | Description | Default |
|---|---|---|---|
| taskIds | Yes | ||
| timeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, idempotentHint=false), so the description adds value by disclosing pre-flight behavior: it validates ALL taskIds exist before starting and errors if any is missing, plus what it returns (first snapshot + all statuses). It does not describe blocking/timeout behavior, which limits it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the core wait semantics lead, followed by the return value and the pre-flight validation constraint. Every clause earns its place; only the omission of timeoutMs keeps it from being optimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations governing timing, the description must explain return and blocking behavior. It covers the return payload and validation but omits timeoutMs behavior and what happens on timeout, which is material for a group-wait tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It partly does for taskIds (restates the 1..20 bound and adds the existence-validation/error semantics), but timeoutMs — a meaningful blocking-control parameter — is not mentioned at all, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: waits for the FIRST task in a group to reach a stopping point (terminal state or needs_user), and returns that snapshot plus all statuses. The 'first of a group' framing implicitly distinguishes it from the single-task wait_task sibling, but it never names the alternative, so sibling differentiation is only inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The condition for using it (wait until the first of several tasks stops) is implied by the description, and the 1..20 group constraint is stated, but there is no explicit when-to-use/when-not guidance or naming of the alternative (wait_task) for single-task cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_taskwait_taskARead-only
等待任务到达停点(阻塞只读原语):轮询至终态(succeeded/failed/needs_attention/cancelled/interrupted)或 needs_user,或超时(timeoutMs 缺省 50000ms、上限 600000ms)后返回当前状态快照。适合回合驱动的调用方:run_task 后在本回合内等待结果。超时返回时请再次调用本工具继续等待——本调用不影响任务本体,超时/中断均无害。
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| timeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish a safe read (readOnlyHint=true, destructiveHint=false), and the description adds substantial context beyond them: it is a blocking polling primitive, timeoutMs defaults to 50000ms with a 600000ms cap, and on timeout/interruption the call is harmless and does not affect the task. This tells the agent retrying is safe and the task state is untouched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core behavior, terminal states, and timeout rules front-loaded; every clause carries information. It is heavy in one block with limited visual separation, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description covers the return sufficiently as a 'current state snapshot' and names the terminal states. Combined with the timeout and retry guidance, an agent has what it needs to call it correctly, though it could be clearer about what the snapshot contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for timeoutMs by giving both a default (50000ms) and an upper bound (600000ms) that the schema lacks. taskId semantics (which task to wait on) remain implicit, so it is strong but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (wait for a task to reach a stop point), defines the blocking-read nature, and enumerates the exact terminal states it waits for (succeeded/failed/needs_attention/cancelled/interrupted/needs_user). An agent can distinguish it from siblings like get_task_report, query_task, and wait_any without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says it fits turn-driven callers that call run_task and then wait within the same turn, and instructs to call it again on timeout. It does not, however, address when to prefer it over the sibling wait_any, leaving that choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.7.7- Added
wait_any - Added
wait_task
1 tool update
v0.7.6- Changed
run_task1 field changed- added
Input schema / properties / designDirectionAdded value: +{ + "minLength": 1, + "type": "string" +}
1 tool update
v0.6.7- Changed
run_task1 field changed- added
Input schema / properties / dryRunAdded value: +{ + "type": "boolean" +}
4 tool updates
v0.6.4- Changed
query_task1 field changed- added
Input schema / properties / eventLimitAdded value: +{ + "maximum": 50, + "minimum": 1, + "type": "integer" +}
- Changed
rework_task1 field changed- added
Input schema / properties / repairHintAdded value: +{ + "maxLength": 4000, + "type": "string" +}
- Changed
run_task1 field changed- added
Input schema / properties / acceptanceOverrideAdded value: +{ + "additionalProperties": false, + "properties": { + "checks": { + "items": { + "additionalProperties": false, + "properties": { + "cmd": { + "anyOf": [ + { + "minLength": 1, + "type": "string" + }, + { + "items": { + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" + } + ] + }, + "name": { + "minLength": 1, + "type": "string" + }, + "optional": { + "type": "boolean" + }, + "timeoutMs": { + "exclusiveMinimum": 0, + "type": "integer" + } + }, + "required": [ + "name", + "cmd" + ], + "type": "object" + }, + "type": "array" + }, + "requireChanges": { + "type": "boolean" + }, + "verifyConcurrency": { + "type": "number" + }, + "visual": { + "additionalProperties": false, + "properties": { + "allowedOrigins": { + "default": [], + "items": { + "format": "uri", + "type": "string" + }, + "type": "array" + }, + "baselineRoot": { + "default": "tests/visual/baselines", + "minLength": 1, + "type": "string" + }, + "browser": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "mode": { + "const": "managed", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "chrome", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "edge", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "executable", + "type": "string" + } + }, + "required": [ + "mode", + "executablePath" + ], + "type": "object" + } + ], + "default": { + "mode": "managed" + } + }, + "content": { + "additionalProperties": false, + "default": {}, + "properties": { + "allowRemote": { + "default": false, + "type": "boolean" + }, + "argsTemplate": { + "items": { + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "cache": { + "default": true, + "type": "boolean" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "enabled": { + "default": false, + "type": "boolean" + }, + "env": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/env" + }, + "minConfidence": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "samples": { + "default": 3, + "maximum": 9, + "minimum": 1, + "type": "integer" + }, + "timeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 90000 + } + }, + "type": "object" + }, + "contents": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "allowRemote": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/allowRemote" + }, + "argsTemplate": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/argsTemplate" + }, + "blocking": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/blocking" + }, + "command": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/command" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/cwd" + }, + "env": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/env" + }, + "expect": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/expect" + }, + "files": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "minItems": 1, + "type": "array" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "samples": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/samples" + } + }, + "required": [ + "expect", + "id", + "files" + ], + "type": "object" + }, + "type": "array" + }, + "defaults": { + "additionalProperties": false, + "default": {}, + "properties": { + "capture": { + "default": "viewport", + "enum": [ + "viewport", + "fullPage", + "element" + ], + "type": "string" + }, + "colorScheme": { + "default": "light", + "enum": [ + "light", + "dark", + "no-preference" + ], + "type": "string" + }, + "locale": { + "default": "en-US", + "minLength": 1, + "type": "string" + }, + "maxDiffRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold", + "default": 0.001 + }, + "pixelThreshold": { + "default": 0.1, + "maximum": 1, + "minimum": 0, + "type": "number" + }, + "selector": { + "minLength": 1, + "type": "string" + }, + "timezone": { + "default": "UTC", + "minLength": 1, + "type": "string" + } + }, + "type": "object" + }, + "enabled": { + "default": false, + "type": "boolean" + }, + "images": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "aspectRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "dpi": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "fileSizeBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "files": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "minItems": 1, + "type": "array" + }, + "formats": { + "items": { + "enum": [ + "png", + "jpeg", + "webp" + ], + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "height": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "optional": { + "default": false, + "type": "boolean" + }, + "transparency": { + "enum": [ + "transparent", + "opaque" + ], + "type": "string" + }, + "width": { + "additionalProperties": false, + "properties": { + "exact": { + "exclusiveMinimum": 0, + "type": "number" + }, + "max": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width/properties/exact" + }, + "min": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width/properties/exact" + } + }, + "type": "object" + } + }, + "required": [ + "id", + "files" + ], + "type": "object" + }, + "type": "array" + }, + "limits": { + "additionalProperties": false, + "default": {}, + "properties": { + "artifactBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 524288000 + }, + "concurrency": { + "default": 1, + "exclusiveMinimum": 0, + "maximum": 4, + "type": "integer" + }, + "decodedPixels": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 32000000 + }, + "inputBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 20971520 + }, + "itemTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 60000 + }, + "navigationTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 30000 + }, + "roundTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 300000 + }, + "serviceTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 60000 + }, + "stabilitySamples": { + "default": 3, + "exclusiveMinimum": 0, + "maximum": 3, + "minimum": 2, + "type": "integer" + } + }, + "type": "object" + }, + "pages": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "baseline": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "capture": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/capture" + }, + "content": { + "additionalProperties": false, + "properties": { + "allowRemote": { + "type": "boolean" + }, + "argsTemplate": { + "items": { + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "blocking": { + "default": false, + "type": "boolean" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "env": { + "additionalProperties": { + "minLength": 1, + "type": "string" + }, + "propertyNames": { + "pattern": "^[A-Za-z_][A-Za-z0-9_]*$" + }, + "type": "object" + }, + "expect": { + "maxLength": 4000, + "minLength": 1, + "type": "string" + }, + "samples": { + "maximum": 9, + "minimum": 1, + "type": "integer" + } + }, + "required": [ + "expect" + ], + "type": "object" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "maskSelectors": { + "default": [], + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": "array" + }, + "maxDiffRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "optional": { + "default": false, + "type": "boolean" + }, + "pixel": { + "default": true, + "type": "boolean" + }, + "pixelThreshold": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "readySelector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "route": { + "default": "/", + "type": "string" + }, + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "source": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "type": { + "const": "existing", + "type": "string" + }, + "url": { + "format": "uri", + "type": "string" + } + }, + "required": [ + "type", + "url" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "port": { + "exclusiveMinimum": 0, + "maximum": 65535, + "type": "integer" + }, + "root": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "type": { + "const": "static", + "type": "string" + } + }, + "required": [ + "type", + "root" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "args": { + "default": [], + "items": { + "type": "string" + }, + "type": "array" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot", + "default": "." + }, + "env": { + "additionalProperties": { + "pattern": "^[A-Za-z_][A-Za-z0-9_]*$", + "type": "string" + }, + "default": {}, + "type": "object" + }, + "readyUrl": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/source/anyOf/0/properties/url" + }, + "type": { + "const": "command", + "type": "string" + } + }, + "required": [ + "type", + "command", + "readyUrl" + ], + "type": "object" + } + ] + }, + "steps": { + "default": [], + "items": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "click", + "type": "string" + } + }, + "required": [ + "type", + "selector" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "input", + "type": "string" + }, + "value": { + "type": "string" + } + }, + "required": [ + "type", + "selector", + "value" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "hover", + "type": "string" + } + }, + "required": [ + "type", + "selector" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "scroll", + "type": "string" + }, + "x": { + "type": "number" + }, + "y": { + "type": "number" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "durationMs": { + "exclusiveMinimum": 0, + "maximum": 60000, + "type": "integer" + }, + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "state": { + "enum": [ + "visible", + "hidden" + ], + "type": "string" + }, + "type": { + "const": "wait", + "type": "string" + }, + "url": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/source/anyOf/0/properties/url" + } + }, + "required": [ + "type" + ], + "type": "object" + } + ] + }, + "type": "array" + }, + "storageState": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "viewports": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "minItems": 1, + "type": "array" + } + }, + "required": [ + "id", + "source" + ], + "type": "object" + }, + "type": "array" + }, + "viewports": { + "default": [ + { + "deviceScaleFactor": 1, + "height": 720, + "id": "desktop", + "width": 1280 + }, + { + "deviceScaleFactor": 1, + "height": 844, + "id": "mobile", + "width": 390 + } + ], + "items": { + "additionalProperties": false, + "properties": { + "deviceScaleFactor": { + "default": 1, + "exclusiveMinimum": 0, + "maximum": 4, + "type": "number" + }, + "height": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width" + }, + "id": { + "maxLength": 100, + "pattern": "^[A-Za-z0-9][A-Za-z0-9_-]*$", + "type": "string" + }, + "width": { + "exclusiveMinimum": 0, + "type": "integer" + } + }, + "required": [ + "id", + "width", + "height" + ], + "type": "object" + }, + "minItems": 1, + "type": "array" + } + }, + "type": "object" + } + }, + "type": "object" +}
- Changed
verify_task1 field changed- added
Input schema / properties / acceptanceOverrideAdded value: +{ + "additionalProperties": false, + "properties": { + "checks": { + "items": { + "$ref": "#/properties/extraChecks/items" + }, + "type": "array" + }, + "requireChanges": { + "type": "boolean" + }, + "verifyConcurrency": { + "type": "number" + }, + "visual": { + "additionalProperties": false, + "properties": { + "allowedOrigins": { + "default": [], + "items": { + "format": "uri", + "type": "string" + }, + "type": "array" + }, + "baselineRoot": { + "default": "tests/visual/baselines", + "minLength": 1, + "type": "string" + }, + "browser": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "mode": { + "const": "managed", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "chrome", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "edge", + "type": "string" + } + }, + "required": [ + "mode" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "executablePath": { + "minLength": 1, + "type": "string" + }, + "mode": { + "const": "executable", + "type": "string" + } + }, + "required": [ + "mode", + "executablePath" + ], + "type": "object" + } + ], + "default": { + "mode": "managed" + } + }, + "content": { + "additionalProperties": false, + "default": {}, + "properties": { + "allowRemote": { + "default": false, + "type": "boolean" + }, + "argsTemplate": { + "items": { + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "cache": { + "default": true, + "type": "boolean" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "enabled": { + "default": false, + "type": "boolean" + }, + "env": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/env" + }, + "minConfidence": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "samples": { + "default": 3, + "maximum": 9, + "minimum": 1, + "type": "integer" + }, + "timeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 90000 + } + }, + "type": "object" + }, + "contents": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "allowRemote": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/allowRemote" + }, + "argsTemplate": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/argsTemplate" + }, + "blocking": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/blocking" + }, + "command": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/command" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/cwd" + }, + "env": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/env" + }, + "expect": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/expect" + }, + "files": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "minItems": 1, + "type": "array" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "samples": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/content/properties/samples" + } + }, + "required": [ + "expect", + "id", + "files" + ], + "type": "object" + }, + "type": "array" + }, + "defaults": { + "additionalProperties": false, + "default": {}, + "properties": { + "capture": { + "default": "viewport", + "enum": [ + "viewport", + "fullPage", + "element" + ], + "type": "string" + }, + "colorScheme": { + "default": "light", + "enum": [ + "light", + "dark", + "no-preference" + ], + "type": "string" + }, + "locale": { + "default": "en-US", + "minLength": 1, + "type": "string" + }, + "maxDiffRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold", + "default": 0.001 + }, + "pixelThreshold": { + "default": 0.1, + "maximum": 1, + "minimum": 0, + "type": "number" + }, + "selector": { + "minLength": 1, + "type": "string" + }, + "timezone": { + "default": "UTC", + "minLength": 1, + "type": "string" + } + }, + "type": "object" + }, + "enabled": { + "default": false, + "type": "boolean" + }, + "images": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "aspectRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "dpi": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "fileSizeBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "files": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "minItems": 1, + "type": "array" + }, + "formats": { + "items": { + "enum": [ + "png", + "jpeg", + "webp" + ], + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "height": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "optional": { + "default": false, + "type": "boolean" + }, + "transparency": { + "enum": [ + "transparent", + "opaque" + ], + "type": "string" + }, + "width": { + "additionalProperties": false, + "properties": { + "exact": { + "exclusiveMinimum": 0, + "type": "number" + }, + "max": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width/properties/exact" + }, + "min": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/images/items/properties/width/properties/exact" + } + }, + "type": "object" + } + }, + "required": [ + "id", + "files" + ], + "type": "object" + }, + "type": "array" + }, + "limits": { + "additionalProperties": false, + "default": {}, + "properties": { + "artifactBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 524288000 + }, + "concurrency": { + "default": 1, + "exclusiveMinimum": 0, + "maximum": 4, + "type": "integer" + }, + "decodedPixels": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 32000000 + }, + "inputBytes": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 20971520 + }, + "itemTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 60000 + }, + "navigationTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 30000 + }, + "roundTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 300000 + }, + "serviceTimeoutMs": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width", + "default": 60000 + }, + "stabilitySamples": { + "default": 3, + "exclusiveMinimum": 0, + "maximum": 3, + "minimum": 2, + "type": "integer" + } + }, + "type": "object" + }, + "pages": { + "default": [], + "items": { + "additionalProperties": false, + "properties": { + "baseline": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "capture": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/capture" + }, + "content": { + "additionalProperties": false, + "properties": { + "allowRemote": { + "type": "boolean" + }, + "argsTemplate": { + "items": { + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" + }, + "blocking": { + "default": false, + "type": "boolean" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "env": { + "additionalProperties": { + "minLength": 1, + "type": "string" + }, + "propertyNames": { + "pattern": "^[A-Za-z_][A-Za-z0-9_]*$" + }, + "type": "object" + }, + "expect": { + "maxLength": 4000, + "minLength": 1, + "type": "string" + }, + "samples": { + "maximum": 9, + "minimum": 1, + "type": "integer" + } + }, + "required": [ + "expect" + ], + "type": "object" + }, + "id": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "maskSelectors": { + "default": [], + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": "array" + }, + "maxDiffRatio": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "optional": { + "default": false, + "type": "boolean" + }, + "pixel": { + "default": true, + "type": "boolean" + }, + "pixelThreshold": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/pixelThreshold" + }, + "readySelector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "route": { + "default": "/", + "type": "string" + }, + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "source": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "type": { + "const": "existing", + "type": "string" + }, + "url": { + "format": "uri", + "type": "string" + } + }, + "required": [ + "type", + "url" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "port": { + "exclusiveMinimum": 0, + "maximum": 65535, + "type": "integer" + }, + "root": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "type": { + "const": "static", + "type": "string" + } + }, + "required": [ + "type", + "root" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "args": { + "default": [], + "items": { + "type": "string" + }, + "type": "array" + }, + "command": { + "minLength": 1, + "type": "string" + }, + "cwd": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot", + "default": "." + }, + "env": { + "additionalProperties": { + "pattern": "^[A-Za-z_][A-Za-z0-9_]*$", + "type": "string" + }, + "default": {}, + "type": "object" + }, + "readyUrl": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/source/anyOf/0/properties/url" + }, + "type": { + "const": "command", + "type": "string" + } + }, + "required": [ + "type", + "command", + "readyUrl" + ], + "type": "object" + } + ] + }, + "steps": { + "default": [], + "items": { + "anyOf": [ + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "click", + "type": "string" + } + }, + "required": [ + "type", + "selector" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "input", + "type": "string" + }, + "value": { + "type": "string" + } + }, + "required": [ + "type", + "selector", + "value" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "hover", + "type": "string" + } + }, + "required": [ + "type", + "selector" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "type": { + "const": "scroll", + "type": "string" + }, + "x": { + "type": "number" + }, + "y": { + "type": "number" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "durationMs": { + "exclusiveMinimum": 0, + "maximum": 60000, + "type": "integer" + }, + "selector": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/defaults/properties/selector" + }, + "state": { + "enum": [ + "visible", + "hidden" + ], + "type": "string" + }, + "type": { + "const": "wait", + "type": "string" + }, + "url": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/pages/items/properties/source/anyOf/0/properties/url" + } + }, + "required": [ + "type" + ], + "type": "object" + } + ] + }, + "type": "array" + }, + "storageState": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/baselineRoot" + }, + "viewports": { + "items": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/id" + }, + "minItems": 1, + "type": "array" + } + }, + "required": [ + "id", + "source" + ], + "type": "object" + }, + "type": "array" + }, + "viewports": { + "default": [ + { + "deviceScaleFactor": 1, + "height": 720, + "id": "desktop", + "width": 1280 + }, + { + "deviceScaleFactor": 1, + "height": 844, + "id": "mobile", + "width": 390 + } + ], + "items": { + "additionalProperties": false, + "properties": { + "deviceScaleFactor": { + "default": 1, + "exclusiveMinimum": 0, + "maximum": 4, + "type": "number" + }, + "height": { + "$ref": "#/properties/acceptanceOverride/properties/visual/properties/viewports/items/properties/width" + }, + "id": { + "maxLength": 100, + "pattern": "^[A-Za-z0-9][A-Za-z0-9_-]*$", + "type": "string" + }, + "width": { + "exclusiveMinimum": 0, + "type": "integer" + } + }, + "required": [ + "id", + "width", + "height" + ], + "type": "object" + }, + "minItems": 1, + "type": "array" + } + }, + "type": "object" + } + }, + "type": "object" +}
2 tool updates
v0.6.2- Changed
run_task1 field changed- added
Input schema / properties / idempotencyKeyAdded value: +{ + "type": "string" +}
- Changed
verify_task1 field changed- added
Input schema / properties / idempotencyKeyAdded value: +{ + "type": "string" +}
1 tool update
v0.5.8- Changed
run_task2 fields changed- added
Input schema / properties / modelSourceAdded value: +{ + "enum": [ + "default", + "custom" + ], + "type": "string" +} - changed
Input schema / properties / reasoningLevel / enumPrevious value: -[ - "低", - "中", - "高", - "low", - "medium", - "high" -]New value: +[ + "低", + "中", + "高", + "low", + "medium", + "high", + "max", + "极高", + "xhigh", + "最大", + "关闭思考", + "on", + "off" +]
3 tool updates
v0.5.3- Added
approve_visual_baseline - Added
prepare_visual_baseline - Changed
run_task2 fields changed- added
Input schema / properties / allowCreateProjectAdded value: +{ + "type": "boolean" +} - changed
Input schema / requiredPrevious value: -[ - "projectPath", - "task" -]New value: +[ + "task" +]
2 tool updates
v0.3.1- Added
continue_task - Changed
run_task3 fields changed- added
Input schema / properties / designSystemAdded value: +{ + "minLength": 1, + "type": "string" +} - added
Input schema / properties / planDocAdded value: +{ + "minLength": 1, + "type": "string" +} - added
Input schema / properties / reasoningLevelAdded value: +{ + "enum": [ + "低", + "中", + "高", + "low", + "medium", + "high" + ], + "type": "string" +}
4 tool updates
v0.1.9- Changed
get_task_report2 fields changed- removed
Input schema / properties / round / exclusiveMinimumRemoved value: -0 - added
Input schema / properties / round / minimumAdded value: +0
- Changed
rework_task1 field changed- removed
Input schema / properties / roundRemoved value: -{ - "exclusiveMinimum": 0, - "type": "integer" -}
- Changed
run_task2 fields changed- added
Input schema / properties / modeAdded value: +{ + "enum": [ + "Work", + "Code", + "Design" + ], + "type": "string" +} - added
Input schema / properties / modelAdded value: +{ + "minLength": 1, + "type": "string" +}
- Changed
verify_task2 fields changed- added
Input schema / properties / baselineRef / minLengthAdded value: +1 - added
Input schema / properties / checksModeAdded value: +{ + "enum": [ + "append", + "replace" + ], + "type": "string" +}
8 tool updates
v0.1.0- First observed
cancel_task - First observed
get_profiles - First observed
get_task_report - First observed
list_tasks - First observed
query_task - First observed
rework_task - First observed
run_task - First observed
verify_task
TDQS
Scored across 13 tools
Each tool maps to a distinct lifecycle action: dispatch (run_task), wait (wait_task/wait_any), inspect (query_task/get_task_report/list_tasks), resume (continue_task), rework (rework_task), cancel, verify, and baseline approval. The only mild overlap is among read-only status tools (query_task vs wait_task vs wait_any), but their blocking, snapshot, and group semantics are clearly described.
All names use lower_snake_case with a verb-first pattern (get_task_report, list_tasks, run_task, verify_task, rework_task, etc.). wait_any is the only verb-only name that slightly breaks the verb_noun convention, but the set remains highly predictable.
13 tools is well within the ideal 3-15 range and each tool corresponds to a real orchestration need: dispatch, polling/waiting, inspection, cancellation, recovery, verification, baseline approval, and adapter inspection. No tool feels redundant or out of scope.
The surface covers the full task lifecycle from run_task through wait/query/cancel/continue/rework/verify/report and includes visual-baseline prepare/approve plus profile inspection. Minor gaps exist, such as no reject/discard operation for a prepared visual baseline and no direct standalone plan-doc retrieval tool, but these are workable via existing meta outputs.
Maintenance
Related MCP Connectors
Hand tasks, bugs and finished work to AI coding agents, and get back a write-up with evidence.
AI-powered spec-to-task decomposition and execution orchestration for coding agents.
AI work orchestration for plans, tasks, teams, and coding-agent dispatch.
Project registry, behavioral specs, and engineering threads for AI coding agent workflows.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceCoordinates parallel AI coding agents by providing task ownership, scoped file locks, handoffs, and verification workflows.MIT
- AlicenseNot gradedqualityBmaintenanceEnables a model in Claude Code, Codex, or Grok to allocate tasks to other agent CLIs on the machine, choosing the lane, model, and reasoning effort, and have a different model verify the result blind.Apache 2.0
- AlicenseAqualityAmaintenanceEnables orchestrating external AI agents to execute development tasks with asynchronous management, automated acceptance verification, failure rework loops, and optional visual baseline validation.119 npmApache 2.0
- AlicenseCqualityAmaintenanceEnables multiple CLI-based AI agents to collaborate as a coordinated team through shared task queues, shared memory, and a message bus, with DAG orchestration, rate-limit avoidance, parallel dispatching, and long-task management.632MIT