Claude-codex-coop
Claude-Codex-coop
English · 中文
你是否也遇到过这些情况?
一个 AI 说得头头是道,你却拿不准它是不是在"一本正经地胡说";
想让 Claude 和 Codex 互相把关,只能在两个窗口之间来回复制粘贴,两边说法不同时也不知道该信谁。
Claude-Codex-coop 让两个 AI 在同一个对话里替你互相把关。
在 Claude 桌面版或 Codex 桌面版里照常聊天,当前的 AI 就是主协调:它先形成自己的判断,需要时自动调用另一边的 AI 做审查、找反例或实现,逐项核对对方的依据,再给你一个综合回答,并说清哪些已证实、哪些存疑。
在哪边聊天 | 主协调 | 被调用执行任务的协作 AI |
Claude 桌面版 | Claude | Codex |
Codex 桌面版 | Codex | Claude |
预览版。Windows 10/11 已实测;macOS 为实验性支持(仅脚本安装,尚未在真机验证)。
和官方 codex-plugin-cc 有什么不同
OpenAI 的 codex-plugin-cc 让你在 Claude Code 里调用 Codex,提供命令、自然语言委派和可选的审查门禁。本项目的侧重点不同:
双向:同时装进 Claude 与 Codex 桌面版,在哪边聊天,哪边就能调用另一边。
自动协作,并按协议复核:开启后按任务需要自动调用;主协调先给初判,结构化交接,逐项核对依据,分歧转成可执行的检验,最后标出已证实与存疑之处。适用于代码,也适用于写作、分析等需要"防胡说"的场景。
可视化面板:随时看到对方的模型、额度、完整回复和实时进度。
Related MCP server: phone-a-friend-mcp-server
功能
自动协作:打开一次,之后值得协作的任务会自动调用另一边;问候、状态查询等不会触发。
协作协议:主协调独立判断后,把目标、已知、初判、具体问题、验收标准交给对方;对方按分析、决策、审查、实施四种模式交付;每个任务最多两轮。
侧边面板:对方的模型、推理强度、实时状态、账号额度与本轮 token;完整回复(思考过程与命令步骤默认折叠);手动指定下一次调用的模型与推理强度;背景可跟随 App 浅色/深色,或选择极光、星河、地平线动态背景;界面中英文随系统语言,也可在外观菜单切换。
按任务授权:分析、决策、审查只读;实施任务按你给出的授权修改文件,两边规则相同。
安装
前提:Windows 10/11;Python 3.10+,并且在 PowerShell 里能运行 python --version(Codex 只从 PATH 启动插件,脚本安装也需要);已登录的 Claude 桌面版与 Codex 桌面版。
方式一:插件市场(推荐)
在 PowerShell 中为 Claude 安装(终端与 Claude 桌面版的本地会话共用同一份插件设置,装一次两边都能用):
claude plugin marketplace add youagainchen/claude-codex-coop
claude plugin install ai-coop@claude-codex-coop也可以在 Claude Code 终端会话里输入 /plugin marketplace add youagainchen/claude-codex-coop;添加之后,在桌面版 Code 标签页点输入框旁的 + → Plugins → Add plugin 也能找到并安装它。
在 PowerShell 中为 Codex 安装。只装了 Codex 桌面版时 codex 不在 PATH 里,先定位 App 自带的那一份:
$codex = (Get-ChildItem "$env:LOCALAPPDATA\OpenAI\Codex\bin\*\codex.exe" | Sort-Object LastWriteTime | Select-Object -Last 1).FullName
& $codex plugin marketplace add youagainchen/claude-codex-coop
& $codex plugin add ai-coop@claude-codex-coop已单独安装 Codex CLI 的话,直接运行 codex plugin marketplace add … 和 codex plugin add … 即可。
方式二:安装脚本
git clone https://github.com/youagainchen/claude-codex-coop.git
cd claude-codex-coop
.\scripts\install.ps1 -Target Both没有 Git 时,可在 GitHub 页面下载 ZIP,解压后在解压出的仓库文件夹里打开 PowerShell,运行 .\scripts\install.ps1 -Target Both。若 PowerShell 提示禁止运行脚本,改用 powershell -ExecutionPolicy Bypass -File .\scripts\install.ps1 -Target Both。
只装一边:
-Target Claude或-Target Codex;禁止协作方修改工作区文件:加-ReadOnly。升级:
git pull后再次运行.\scripts\install.ps1 -Target Both。
macOS(实验性)
需要 Python 3.10+(brew install python 或 python.org 安装包;系统自带的 /usr/bin/python3 通常是 3.9,版本不够)。
git clone https://github.com/youagainchen/claude-codex-coop.git
cd claude-codex-coop
sh scripts/install.sh --target both只装一边:
--target claude或--target codex;禁止协作方修改工作区文件:加--read-only。登录 Claude CLI:
sh scripts/claude-login.sh(面板里的登录按钮会打开“终端”窗口)。尚未在真机验证:两个 App 自带 CLI 的位置是按 Windows 版推断的,找不到时安装脚本会直接报错退出,不改动任何文件;可以先单独安装 Claude Code / Codex CLI。遇到问题欢迎提 issue。
安装后完全退出并重新打开两个 App。
使用
在任意一边的对话里说 "打开 AI Coop",面板出现在右侧,自动协作开启。
第一次使用时,如果面板提示 Claude CLI 未登录,点 "登录 Claude CLI",在浏览器里用与 Claude 桌面版相同的账号授权即可(只需一次)。
正常描述任务。可以先试一句:"帮我审查这个项目的 README,请另一边找出遗漏,并核对你们的分歧"。面板里出现调用记录,就说明协作已经生效。
指定对方的模型或推理强度:在面板"下一次调用"里选"手动指定",或直接在对话里说,例如"让 Codex 用 gpt-6-sol、high 审查这个方案"。
关闭:点面板右上角的开关,或说"关闭 AI Coop"。
数据与隐私
发给协作 AI 的内容:任务包,以及当前工作区的只读快照:文件清单(最多 400 项)和常见说明文件(
AGENTS.md、CLAUDE.md、README.md等)。快照会跳过疑似密钥或凭据的文件(
.env*、*.pem、*.key、id_rsa*、*credentials*、*secret*等),以及.git、.ssh、node_modules等目录。在项目根目录放一个
.ai-coop-context(每行一个相对路径),可以指定随文件清单一起附带哪些说明文件;密钥类文件即使写进去也会被跳过。
本地记录:每次调用的请求、回复与事件流保存在
~/.ai-coop/runs/。账号额度与模型列表(额度为预览功能,仅用于面板显示):使用两边 CLI 已有的登录查询额度(
api.anthropic.com、chatgpt.com);Claude 的可用模型来自 Anthropic 的模型接口,Codex 的模型来自本机 Codex CLI。额度接口并非公开文档接口,两边升级后可能失效;失效时面板只是不显示额度,不影响协作。
卸载
插件市场安装:在 PowerShell 中运行
claude plugin uninstall ai-coop@claude-codex-coop和codex plugin remove ai-coop@claude-codex-coop;只装了 Codex 桌面版时,先按上面的方法定位$codex,再运行& $codex plugin remove ai-coop@claude-codex-coop。脚本安装:
.\scripts\uninstall.ps1 -Target Both,加-Purge会同时删除~/.ai-coop整个目录(运行记录、偏好和升级时留下的插件备份);macOS 用sh scripts/uninstall.sh --target both(--purge同理)。
常见问题
协作一直失败:运行
.\scripts\doctor.ps1,确认输出中的codex_exe和claude_exe都有路径;任一为null表示没找到对应的 CLI,可用环境变量AI_COOP_CODEX_EXE/AI_COOP_CLAUDE_EXE指定。macOS 上改为在仓库目录运行python3 server/mcp_server.py --self-test。找不到 Python(Windows):安装 Python 3.10+ 并勾选加入 PATH(Codex 只从 PATH 启动插件,两种安装方式都要求 PATH 上有
python);Claude 一侧的脚本安装也可以用环境变量AI_COOP_PYTHON_EXE指定解释器。找不到 Python(macOS):安装脚本会依次尝试
python3.14…python3.10、python3和 Homebrew、python.org 的默认位置,选第一个 3.10+;也可以用环境变量AI_COOP_PYTHON指定。两个 App 都用这个安装时选定的解释器启动插件,与 App 的 PATH 无关。
工作原理
插件是一个本地 MCP 服务(server/mcp_server.py,只依赖 Python 标准库):主协调调用 start_workflow 后,服务在后台启动另一边的 CLI(codex exec --json 或 claude -p --output-format stream-json),事件实时写入面板,结束后结果交回主协调。调用 Claude 时固定使用 Claude CLI 登录的账号。面板(server/sidebar.py)只监听 127.0.0.1,端口随机,每次启动生成一个访问令牌。
开发测试:python -m unittest discover -s tests
许可
代码采用 MIT 许可,见 LICENSE。面板的"极光"背景改编自 nimitz 的 Shadertoy 作品,出处与许可见 NOTICE.md。
Available Tools
15 toolsget_model_catalog读取模型目录BRead-only
识别主对话 AI,并读取另一端协作 AI 的候选模型、推理强度和已保存设置。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| primary_agent | No | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false. The description adds data scope by naming candidate models, reasoning intensity, and saved settings, but does not add further behavioral context such as auth needs, rate limits, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Its brevity is efficient, though it leaves important usage and parameter context unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Read-only annotations and the workspace schema cover safety and path context, and the description names the returned data. However, the lack of usage routing and primary_agent parameter detail leaves gaps for selecting among sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; workspace is described in the schema, but primary_agent has no schema description. The description hints at '主对话 AI' but does not explain the auto/codex/claude enum values or how primary_agent affects the catalog retrieval.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states specific verbs (识别/读取) and resources (主对话 AI、候选模型、推理强度、已保存设置). It does not explicitly differentiate from siblings like show_partner_selector or set_partner_preferences, but the core purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not, or alternative tool is mentioned. The agent must infer the appropriate context from the name and sibling tools, with no routing guidance in the definition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_statusBRead-only
查询后台协作任务状态;完成后返回最终报告。
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: the response shape changes with task state, returning the final report only on completion. It omits polling cadence, terminal/error states, and whether partial progress is ever returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the action and then the completion behavior; there is no filler. It is lean to the point of under-specification, but nothing is padded or wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only status check with annotations covering safety, the description covers purpose and the key return behavior. It is still incomplete on how to obtain run_id and what non-terminal responses contain, which matters for a polling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required run_id param has no description in either the schema or the description text. The phrase about a background collaboration task weakly implies run_id identifies that task, but the agent gets no format, no source (list_runs/start_workflow), and no validity guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: queries the status of a background collaboration task and returns the final report once done. That is clear enough to act on, but it does not differentiate itself from sibling list_runs, which an agent could easily confuse with this lookup-by-id tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as list_runs or start_workflow. The agent must infer that run_id comes from a previously started run, and nothing tells it when polling is appropriate versus when to list runs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usage_summaryARead-only
读取 ai-coop 单次或最近多次任务的 Codex/Claude token 与费用统计;不代表账户剩余额度。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| run_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful interpretive context beyond that: the figures do not represent remaining account quota, which prevents a plausible misuse of the returned data. It still says nothing about granularity, currency, or freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with zero filler, front-loading the resource and scope and reserving the trailing clause for the important non-quota caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values, yet it only says 'token 与费用统计' without units, currency, or breakdown, and both parameters are undocumented in the schema. Adequate to call the tool, but an agent cannot predict the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter is described in the schema. The description only implicitly gestures at the two modes ('单次' vs '最近多次'), leaving the agent to guess that run_id selects one run and limit bounds a recent-run list; it does not document the limit's default, range, or the meaning of run_id's format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (读取) and resource (Codex/Claude token 与费用统计) with a scope covering single or recent multiple runs. It is distinguishable from siblings like list_runs and get_run_status, though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '单次或最近多次任务' implies the two usage modes (a specific run vs. recent history), which maps loosely to the run_id/limit split, and the closing caveat clarifies what this is not. However, there is no explicit when-to-use guidance or named alternative tool for related questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
health_check检查 AI Coop 连接BRead-only
检查 Codex/Claude CLI、工作区、运行目录和写入模式,不启动模型任务。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the meaningful constraint that no model task is launched, reinforcing read-only behavior, but says nothing about what the checks report, how failures surface, or whether it touches the filesystem.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence listing the checked items first and the non-behavior last — well front-loaded with no filler. It is a touch terse given no return-value information, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description would ideally hint at what the health check returns (per-component status, errors). It names what is checked but not what an agent gets back, leaving a moderate gap for a diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one optional parameter and schema coverage is 100%, so the schema fully documents workspace including its dynamic per-chat semantics. Baseline 3 applies since the description adds nothing about the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (检查/check) with four named resources (Codex/Claude CLI, workspace, run directory, write mode), which is far more specific than a generic 'health check'. It implicitly separates itself from the start_* workflow siblings by scoping to diagnostics, though it never names an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause '不启动模型任务' implies this is a pre-flight/diagnostic call rather than an execution call, which is useful routing context. However it gives no explicit when-to-use guidance and does not distinguish itself from the similarly diagnostic project_status sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsBRead-only
列出最近的协作运行记录。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and openWorldHint=false, so the agent knows it's a safe, local read operation. The description adds minor context that it lists recent records, but doesn't disclose pagination behavior, sorting order, or what 'recent' means—limitations that matter for a list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no waste, but it is not overly front-loaded or structured—it simply states the purpose without additional details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool lists runs, an agent needs to know the default limit, whether results are paginated, and how 'recent' is defined. The description provides none of this, and with no output schema, it leaves significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the single parameter (limit) is documented only by its type and constraints (integer, 1-50) without explanatory text. The description does not mention the parameter at all, so it fails to compensate for the schema's lack of description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (列出/list) and resource (最近的协作运行记录/recent collaborative run records), which is clear enough for an agent to understand the tool's purpose. However, it does not explicitly differentiate from siblings like get_run_status, which also concerns runs, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as get_run_status for specific runs. The description only states what the tool does, not the context or conditions for its use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_sidebar准备 AI Coop 侧栏A
准备当前工作区的本地协作控制页(显示协作伙伴模型、推理强度与运行状态),返回 URL。Codex 用 open_in_codex 在 right 面板打开;Claude 桌面版用内置浏览器面板 preview_start 打开。此工具本身不会打开面板。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | Yes | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| primary_agent | No | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, so the agent already knows this performs some non-destructive write. The description adds the key behavioral fact that the tool only prepares a page and returns a URL rather than opening a panel, which is genuinely useful context beyond the annotations, though it does not say whether the prepared page persists or how long the URL is valid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool produces, followed by per-client routing and the explicit non-behavior. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema tool, the description covers the return value (URL), the client-specific next step, and the side-effect boundary. The only residual gap is the unexplained primary_agent enum, which is more a schema/documentation issue than a description omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: workspace is documented in the schema, but primary_agent (enum auto/codex/claude) has no schema description. The description mentions Codex and Claude desktop usage context, which loosely hints at what primary_agent selects, but never explains the parameter or the 'auto' default directly, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: it prepares a local collaboration control page for the current workspace and returns a URL. It explicitly negates the obvious confusion ('此工具本身不会打开面板'), which separates it from panel-opening tools and makes its scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: Codex should open the returned URL with open_in_codex in the right panel, Claude desktop should use the built-in browser panel preview_start. It also states what this tool does NOT do, so the agent knows the handoff point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_statusARead-only
读取项目 Git 状态和差异统计;非 Git 目录则列出顶层文件。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and openWorldHint, so the safety profile is already covered. The description adds fallback behavior for non-Git directories listing top-level files, which is useful context beyond annotations, though it omits details like permissions or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the primary function and fallback. It is appropriately sized with zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a simple read-only status tool with an optional workspace parameter; the description covers both Git and non-Git cases. It lacks return format details, but no output schema exists and the function is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single workspace parameter is fully documented in the schema. The description adds no parameter-specific meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (读取) and resources (Git 状态和差异统计), plus a fallback for non-Git directories. Distinguishes from sibling read_project_file by focusing on project-level status rather than file contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for checking project Git status, but does not explicitly state when to use this versus read_project_file or other status tools. No exclusions or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_project_fileBRead-only
读取所选工作区内的 UTF-8 文本文件,禁止越界访问。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| max_chars | No | ||
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety and closed-world behavior are covered. The description adds that only UTF-8 text files are read and that out-of-workspace access is forbidden, which is useful behavioral context, but it omits error handling, truncation, and permission details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It immediately states the action, file type, and scope constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read tool with low schema description coverage and no output schema, the description is too sparse. It does not explain parameter semantics such as path or max_chars behavior, leaving the agent to infer important usage details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with 'path' and 'max_chars' undocumented in the schema. The description does not compensate: it does not explain path format, what max_chars does, or the 30000 default limit, and only vaguely references the workspace parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: reading UTF-8 text files within the selected workspace. It also clarifies the access boundary ('禁止越界访问'), but does not differentiate itself from any sibling tool, as the listed siblings are unrelated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The intended use is only implied by the tool name and the phrase '读取所选工作区内的 UTF-8 文本文件'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_collaboration_mode打开或关闭 AI CoopA
打开后,当前宿主中的后续实质任务会自动交给另一端 AI 协作;用户无需再调用 workflow。
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this as state-changing but non-destructive, and the description adds meaningful context beyond that: the mode persists and alters the host's future task routing automatically. It does not disclose scope (session vs. global) or how to revert, but the core behavioral trait is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the effect is stated first. It is efficient, though covering only the enabled branch makes it slightly one-sided rather than fully structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation toggle with no output schema and one undocumented parameter, the description explains the primary effect but omits the disable behavior, persistence scope, and any prerequisite (e.g., a partner being selected via set_partner_preferences). Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One boolean parameter with 0% schema description coverage. The description describes the enabled behavior (打开后...) but never names the parameter or explains the disabled state, which matters because the default is true. It partially compensates but leaves the off-path undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete effect: once enabled, subsequent substantive tasks in the current host are auto-delegated to the other AI, removing the need to call workflow. This distinguishes it from the start_workflow sibling by implication. It is slightly indirect in that it never explicitly frames itself as a persistent on/off mode setter, but the intent is recoverable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage (turn this on so you don't have to keep invoking workflow), which is a real routing hint against the start_workflow/start_consultation siblings. However it never states when NOT to use it, nor what disabling does, so the guidance is only implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_partner_preferences保存协作 AI 设置B
保存另一端 AI 的自动/手动选模、模型 ID 和推理强度,供之后的原生聊天协作调用使用。
| Name | Required | Description | Default |
|---|---|---|---|
| partner_model | No | ||
| primary_agent | No | auto | |
| partner_effort | No | medium | |
| partner_selection | No | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, closed-world write. The description adds that the values are persisted for later native chat calls, which is useful persistence context, but it does not say whether saving overwrites prior values, what scope (session vs global) the settings apply to, or what happens if fields are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence led by the verb, with no filler. It is dense but appropriate for a small setter; only the missing primary_agent detail keeps it from being fully earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, all-optional mutation with no output schema, the description covers most parameters and the persistence purpose, but omits one parameter entirely and gives no indication of write semantics or scope. Annotations carry the safety profile, so the remaining gap is moderate rather than severe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It names three of the four inputs (selection mode, model ID, reasoning effort), but primary_agent is never explained, and no format or default semantics (e.g. what 'auto' resolves to) are supplied beyond the raw enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (保存/save) and resource (partner AI preferences: selection mode, model ID, reasoning effort), and its intended downstream use. It is distinguishable from nearby siblings like show_partner_selector (read) and set_collaboration_mode, though it doesn't explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause '供之后的原生聊天协作调用使用' gives the purpose/context for saving, which implies when the settings matter, but there is no explicit when-to-call guidance, no prerequisites, and no mention of alternatives such as set_collaboration_mode or show_partner_selector.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_partner_selector显示协作 AI 选择条ARead-only
在当前对话中显示一个紧凑的协作 AI 模型与推理强度选择条;选择会保存并用于之后的 start_workflow 调用。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| primary_agent | No | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already state readOnlyHint=true and openWorldHint=false, lowering the burden. The description adds useful behavioral context beyond the annotations: it says the selector appears in the current conversation and that the selection persists and affects subsequent start_workflow calls. There is no direct contradiction with the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no wasted words. The display action is front-loaded, followed immediately by the key behavioral consequence for start_workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, two-parameter, read-only UI selector with no output schema, the description covers the core action and downstream effect. It remains incomplete on routing versus sibling preference tools and does not clarify how the mentioned reasoning-strength selector maps to the available input parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%; 'workspace' is documented in the schema, while 'primary_agent' has no schema description. The description mentions model and reasoning-strength selection, but the schema only exposes 'primary_agent', and the enum values (auto, codex, claude) are not explained. It adds some conceptual context but does not fully compensate for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('显示') and resource ('协作 AI 模型与推理强度选择条') in a clear context ('在当前对话中'). It also names a downstream dependency ('start_workflow'), but it does not differentiate itself from sibling tools such as set_partner_preferences, set_collaboration_mode, or open_sidebar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by saying the selection will be saved and used in later start_workflow calls, which suggests it should be invoked before a workflow. However, it gives no explicit when-to-use guidance, no exclusions, and does not mention alternative sibling tools for setting preferences or collaboration mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_codex_implementationB
按用户已批准的决策让 Codex 修改项目并验证。服务需显式启用 --allow-write。
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| task_id | No | ||
| reviewer | No | ||
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| codex_model | No | gpt-5.6-sol | |
| human_owner | No | ||
| codex_effort | No | medium | |
| approved_decision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=false; the description adds the operationally critical fact that the service must be launched with --allow-write or the call cannot write, and that a verification step follows the modification. It still does not say what is modified, whether changes are reversible, or what happens on partial failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the enabling requirement; no filler. The second sentence is terse enough to be slightly cryptic, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a write-capable, 8-parameter tool with no output schema and near-zero schema documentation, yet the description omits what the tool returns, how success/failure is reported, and what most parameters mean. For a mutating tool the description leaves too much unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13% (only workspace is documented), and the description adds meaning for none of the remaining seven parameters (task, task_id, reviewer, codex_model, codex_effort, human_owner). It only loosely gestures at approved_decision, so it fails to compensate for the coverage gap on an 8-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action on a specific resource: letting Codex modify the project and then verify, driven by an approved decision. It is clearly distinct from the other start_* siblings (start_workflow, start_consultation, start_model_debate), though it never explicitly names those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 按用户已批准的决策 implies a precondition (a decision must already be approved) and the --allow-write sentence implies an environment prerequisite, which is real usage guidance. However, it never says when to prefer this over the other start_* tools or how to obtain an approved decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_consultationCRead-only
后台启动一次 Codex 或 Claude 的独立只读咨询。
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | 独立审稿人 | |
| task | Yes | ||
| agent | Yes | ||
| model | No | ||
| effort | No | ||
| context | No | ||
| task_id | No | ||
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| human_owner | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description usefully adds that the run is background (async, presumably returning a task_id) and independent/isolated, but it omits lifecycle details (how to retrieve results, whether it blocks, authorization/owner semantics) despite a 9-param surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight, front-loaded sentence with no filler. However it is arguably under-specified rather than richly concise given the 9-param surface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-output-schema tool with near-zero parameter documentation, one sentence is far from sufficient. An agent cannot determine the role/effort/model semantics or how to retrieve the async result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 11% (just the workspace path), yet the tool has 9 parameters including role, task, model, effort (9-value enum), context, task_id and human_owner that are entirely undocumented. The description mentions only 'agent' implicitly via Codex/Claude, leaving most parameter meaning to guesswork.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (start) plus resource (consultation) and scopes it as background and read-only, naming the two supported agents (Codex/Claude). This clearly distinguishes it from write/implementation siblings, though it doesn't explicitly name those siblings (e.g. start_codex_implementation, start_model_debate) to route the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this over the many sibling start_* tools. The 'background' and 'read-only consultation' framing implies a use case, but no explicit when/when-not conditions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_model_debateBRead-only
后台启动 Claude 与 Codex 的独立方案、交叉质询和证据裁决;不会修改项目。
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| rounds | No | ||
| context | No | ||
| task_id | No | ||
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| codex_model | No | gpt-5.6-terra | |
| human_owner | No | ||
| claude_model | No | sonnet | |
| codex_effort | No | medium | |
| claude_effort | No | medium |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, and '不会修改项目' is consistent with that rather than adding new information. '后台启动' (background launch) is the one genuinely additive trait: it tells the agent the call is asynchronous and returns before results exist, though it omits how to retrieve those results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, front-loading the verb and scope before the non-mutation guarantee. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter async orchestration tool with no output schema, the description is far too thin: it omits how the background run is tracked or fetched, what the effort/round parameters mean, and what inputs (task vs context vs workspace) are required or optional.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 10 parameters and only ~10% schema description coverage (just workspace), the description must carry the load, and it does not. 'rounds', 'task_id', 'context', 'human_owner', and both 'claude_effort'/'codex_effort' enums are never explained; only the two model roles are implied by naming Claude and Codex.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (启动/start) plus the resource: a background multi-model debate between Claude and Codex with independent proposals, cross-examination, and evidence adjudication. That distinguishes it from single-model siblings like start_codex_implementation or start_consultation, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives no when-to-use or when-not-to-use guidance and never contrasts with the obvious alternatives (start_consultation, start_workflow, start_codex_implementation). Only an implicit signal that it is a heavyweight, non-mutating analysis run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_workflow让协作 AI 执行一轮工作C
当前聊天 AI 自动担任主协调者,把本轮任务交给另一端 AI;完成后主 AI读取结果、继续追问或向用户汇总。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | decide | |
| task | Yes | ||
| context | No | ||
| task_id | No | ||
| workspace | No | 当前原生聊天所打开工作区的绝对路径;每次调用按当前聊天动态选择,不绑定安装目录。 | |
| partner_model | No | ||
| primary_agent | No | auto | |
| partner_effort | No | ||
| approved_decision | No | ||
| partner_selection | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so safety profile is covered by structured data. The description adds genuinely useful behavioral context beyond that: the current AI becomes the coordinator, delegates to a partner, then reads results and may loop with follow-ups. It still omits blocking vs async behavior, timeout/failure handling, and any auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the coordination model and the delegation step. Nothing is wasted, though it is arguably too terse for a 10-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters, 4 enums, near-zero schema description coverage, and no output schema, the description is not complete enough. It explains the orchestration loop but leaves parameter intent, return behavior, and failure modes entirely to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 10 parameters and only 10% schema description coverage, the description carries the burden and largely fails. It alludes to 'task' being handed to the partner AI, but says nothing about mode, approved_decision, partner_model, partner_effort, partner_selection, primary_agent, or context — nine parameters whose meaning and enum values remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys an orchestration action — the current chat AI acts as coordinator and dispatches this round's task to a partner AI — but it never names the tool's actual verb/resource clearly, and it does not distinguish start_workflow from near-identical siblings such as start_consultation, start_model_debate, and start_codex_implementation. An agent can guess it starts a delegation round, but cannot tell which of several similar 'start_*' tools to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the alternatives, no prerequisites, and no exclusions. The description describes the internal flow after the call (read result, follow up, summarize) but gives the agent no selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.1.0- First observed
get_model_catalog - First observed
get_run_status - First observed
get_usage_summary - First observed
health_check - First observed
list_runs - First observed
open_sidebar - First observed
project_status - First observed
read_project_file - First observed
set_collaboration_mode - First observed
set_partner_preferences - First observed
show_partner_selector - First observed
start_codex_implementation - First observed
start_consultation - First observed
start_model_debate - First observed
start_workflow
TDQS
Scored across 15 tools
The start_* tools represent distinct workflows (workflow, consultation, debate, implementation), and the setup/status tools are mostly distinct. However, set_collaboration_mode and start_workflow both concern handing off work to the partner AI, and start_consultation vs start_model_debate could be confused without reading descriptions carefully.
Tool names are consistently snake_case and mostly follow a verb_noun pattern such as get_model_catalog, set_partner_preferences, and start_workflow. A few names break the pattern (health_check, project_status) but overall the convention is predictable.
At 15 tools, the set is within the expected range for a collaboration-orchestration server and each tool appears to serve a distinct role in setup, execution, or monitoring. No obvious redundant or filler tools are present.
The surface covers setup, model selection, workflow starts, status polling, run listing, usage summaries, and implementation runs. A notable gap is the lack of an explicit cancel/stop tool for background collaboration runs, which could leave agents unable to abort a running task.
Maintenance
Related MCP Connectors
Adversarial behavioural-bias engine — audits your decisions for cognitive biases via your own AI.
Convene a panel of expert AI personas to debate any decision from every side.
Let your AI sessions talk to each other — messaging, tasks, sessions, and alerts
Multiple AIs peer-review and debate your question, then return one fact-checked answer.
Related MCP Servers
- AlicenseAqualityDmaintenanceMulti-AI Consensus Tool: Query multiple AI models in parallel, synthesize responses for better accuracy, and reduce AI bias through ensemble decision-making.131MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI-to-AI consultation for critical thinking and complex reasoning via OpenRouter, allowing one AI to delegate tasks to another AI model.10MIT
- FlicenseNot gradedqualityCmaintenanceEnables two local LLMs (Qwen and Llama) to communicate via MCP, allowing one model to consult the other for second opinions or additional reasoning.1-
- AlicenseNot gradedqualityAmaintenanceEnables multi-model LLM council reviews and parallel sidecar conversations within Claude, allowing Claude to orchestrate structured reviews from various AI models and fold their responses back into the session.745 npm2MIT