Skip to main content
Glama

peer-agents-mcp

MCP 服务器,让其他 AI 编码工具(Codex、Claude、Cursor 等)能够调用 Grok CLIAntigravity CLI 作为同行评审员和协作者。

功能

该服务器将本地的 grokagy(Antigravity)CLI 封装在一个简洁的 Model Context Protocol (MCP) 接口之后。

任何支持 MCP 的代理现在都可以:

  • 将代码更改、计划、错误或问题发送给 Grok 或 Antigravity

  • 接收结构化的同行反馈

  • 运行带会话记忆的多轮评审/调试/规划会话

  • 通过在同一个任务上运行两个 CLI 来获得独立意见

主代理(Codex、Claude 等)保持控制权。当它想要第二个(或不同的)意见时,只需将特定任务委托给这些同行。

Related MCP server: xAI Grok MCP Bridge

核心思想

与其让一个模型做所有事情,你的主编码代理可以将 Grok 和 Antigravity 作为同行 使用:

  • Grok 负责大多数编码工作(评审、规划、调试、实现批判)

  • Antigravity 负责大上下文、通用知识或多模态任务

根据请求类型自动进行智能路由。

可用工具

工具

用途

路由到

peer_review_diff

评审统一 diff 或补丁

Grok(通常)

peer_plan

创建实现计划

Grok

peer_debug

从日志/堆栈跟踪中诊断故障

Grok

peer_verify

检查测试/构建输出的安全性

Grok

peer_ask

通用有依据的问答

Antigravity

peer_debate

独立比较方案 A 与方案 B

Grok

peer_turn

继续多轮同行会话

同一同行

peer_turn_async

长时间运行的后续轮次(后台任务)

同一同行

peer_implement_async

冷启动 Grok 实现交接(任务)

Grok

peer_review_diff_async

长时间运行的 diff 评审(后台任务)

Grok

peer_debug_async

长时间运行的调试交接(后台任务)

Grok

peer_job_status

轮询后台任务(包含 progress

peer_job_cancel

取消后台任务

peer_jobs_gc

垃圾回收旧的终态任务

peer_compare

对两个 CLI 进行底层并排调用

两者

其他会话工具:peer_summarizepeer_transcriptpeer_list_sessionspeer_resetpeer_health

所有路由工具都通过 files 参数接受完整文件内容,并通过 diff 接受 diff。切勿发送摘要——请发送实际内容。

何时使用同步与异步

Grok 对普通 diff 的同步评审通常需要 3–6 分钟。Grok 子进程超时为 GROK_TURN_TIMEOUT_MS(默认 6 分钟)。提高子进程超时不会提高宿主 MCP 客户端的等待时间。如果宿主先放弃,Codex 将永远不会在该工具响应中看到 durationAdvisory / continuationHint / nativeSessionId——请使用 *_async

服务器不会静默地将同步工具转换为任务。

使用同步(peer_review_diffpeer_planpeer_turn、…)

使用异步(*_async + peer_job_status

适合 ~80k 提示字符的普通 diff / 计划

~80k 字符(~20k token),接近 120k 上限,或被截断

默认 / 中等风险

risk_level=highfocus=security--effort high

后续“重新检查这一个文件”

实现交接(peer_implement_async

宿主可以等待 ~6 分钟

宿主 MCP 超时 ≤ 2–3 分钟;日志巨大;多次尝试调试

当提示被截断、估计 token 超过 ~20k(~80k 字符)、risk_level=high / focus=security,或消耗了 stub 自动继续时,同步 Grok/路由结果可能包含附加的 durationAdvisory。示例:

{
  "durationAdvisory": "Grok sync reviews of this size often take 3–6 minutes. If your MCP client times out sooner, use peer_review_diff_async / peer_turn_async and poll peer_job_status."
}

截断的 peer_review_diff 仍然同步运行(会前置 tools-against-cwd 指令)。下次请优先使用 peer_review_diff_async;不要将截断视为硬性拒绝。

长时间运行的异步任务

大型实现交接可能超过 MCP 客户端的同步工具超时。请使用异步路径,而不是阻塞在 peer_turn 上:

  1. 使用 peer_implement_async(冷启动)或 peer_turn_async(现有会话)开始工作。

  2. 在同行运行期间继续本地工作。

  3. 30–60 秒 轮询一次 peer_job_status(避免过于频繁的轮询)。

  4. statusrunning 时,可选的 progress 可能包含 textSnippetlastThoughteventCount(Grok streaming-json / agy stream-json)。

  5. statussucceeded 时,读取 result,如有需要可继续使用 peer_turn

  6. 使用 peer_job_cancel 停止由该 MCP 进程拥有的排队/运行中的任务。

终态状态:succeededfailedtimed_outcancelledorphaned

幂等性:使用相同的 idempotency_key 重试会返回同一个任务(运行中或粘性终态)。在 timed_out / cancelled / failed 之后,请使用新的键重试该工作。

任务和已完成的结果存储在 ~/.peer-agents/jobs/ 下。活动的提供方进程不会在 MCP 服务器重启后存活;非终态任务在 hydrate 时被标记为 orphaned(除非会话已经提交了该操作,这种情况下会恢复为 succeeded)。

超过 7 天 的终态任务会在 hydrate 时被垃圾回收(可通过 PEER_AGENTS_JOB_GC_MAX_AGE_MS 覆盖),或通过 peer_jobs_gc 手动回收。

异步任务使用与同步轮次不同的超时:

  • PEER_AGENTS_JOB_TIMEOUT_MS — 默认 30 分钟1800000

  • GROK_JOB_TIMEOUT_MS / ANTIGRAVITY_JOB_TIMEOUT_MS — 可选的按提供方覆盖

  • PEER_AGENTS_JOB_GC_MAX_AGE_MS — 终态任务保留时间(默认 7 天)

  • PEER_AGENTS_GROK_TRANSPORTheadless(默认)或 acp,用于热进程池

  • PEER_AGENTS_GROK_ACP_MAX_CLIENTS — 最大并发 ACP 进程数(默认 4)

  • PEER_AGENTS_GROK_ACP_IDLE_MS — ACP 轮次之间的空闲回收(默认 max(5 min, GROK_TURN_TIMEOUT_MS + 60s))。空闲不是任务生命周期;进行中的 session/prompt 会忽略它。

在任务运行期间保持 MCP 服务器进程存活。

Grok 传输:headless 与 ACP

headless(默认)

acp

调用方式

每轮 grok --prompt-file

每个 cwd 一个长期存活的 grok agent stdio

延迟

每轮冷启动

热进程;多轮复用进程 + 会话

CLI 功能

沙箱、worktree、生成的 --session-id;工具循环上没有 --json-schema;始终使用 --output-format streaming-json

子集(--always-approve);通过提示获得结构化发现。空闲是轮次之间的后备机制,不是任务生命周期。

启用

(默认)

PEER_AGENTS_GROK_TRANSPORT=acp

对于需要严格沙箱的一次性评审,优先使用 headless。当你要运行大量后续 peer_turn 并希望降低进程启动成本时,优先使用 acp

其他代理如何使用它

Codex、Claude 或任何其他 MCP 客户端通过 stdio 连接到该服务器。连接后,代理可以像调用任何其他工具一样调用同行工具。

典型流程:

  1. 你的代理准备一个 diff、错误日志或任务描述。

  2. 它调用 peer_review_diffpeer_planpeer_debug 等。

  3. 服务器以 headless 模式调用相应的 CLI。

  4. 同行响应返回时带有 sessionId

  5. 你的代理之后可以使用该 sessionId 通过 peer_turn 进行后续跟进。

这让你获得持久、有上下文的同行对话,而无需主代理自己管理 CLI 调用。

前提条件

  • Node.js ≥ 18

  • grok CLI(或设置 GROK_COMMAND

  • agy CLI(Antigravity,或设置 ANTIGRAVITY_COMMAND

两个 CLI 都必须已在你的机器上完成身份验证并能正常工作。

安装与使用

git clone https://github.com/Rakeen70210/peer-agents-mcp
cd peer-agents-mcp
npm install
npm run build

直接运行:

node dist/index.js

MCP 客户端配置

将其添加到你的客户端的 MCP 服务器配置中(典型 stdio 设置示例):

{
  "mcpServers": {
    "peer-agents": {
      "command": "node",
      "args": ["/absolute/path/to/peer-agents-mcp/dist/index.js"],
      "env": {
        "GROK_COMMAND": "/home/you/.grok/bin/grok",
        "ANTIGRAVITY_COMMAND": "/home/you/.local/bin/agy"
      }
    }
  }
}

环境变量

  • GROK_COMMAND — grok 二进制文件的路径(默认:grok

  • ANTIGRAVITY_COMMAND — agy 二进制文件的路径(默认:agy

  • GROK_ARGS / ANTIGRAVITY_ARGS — 额外 CLI 参数的 JSON 数组

  • ANTIGRAVITY_CONVERSATIONS_DIR — 覆盖用作回退 session-id 捕获的 agy 会话存储(默认:~/.gemini/antigravity-cli/conversations

  • PEER_AGENTS_WORKTREE_DIR — DIY Grok git worktree 的父目录(默认:~/.peer-agents/worktrees

  • PEER_AGENTS_STORAGE_DIR — 会话持久化位置(默认:~/.peer-agents/sessions

  • PEER_AGENTS_ENABLED_PROVIDERS — 同行 CLI 的逗号分隔白名单(grokantigravity)。当宿主是 Grok 时,单独使用 antigravity,这样同行永远不会重新进入 Grok。

  • PEER_AGENTS_DISABLED_PROVIDERS — 逗号分隔黑名单(如果设置了 PEER_AGENTS_ENABLED_PROVIDERS,则忽略此项)

  • GROK_TURN_TIMEOUT_MS — headless 和 ACP 的 Grok 同步超时(默认 6 分钟 / 360000)。除每次调用的 timeoutMs 外,唯一的超时来源。Grok 不会读取 PEER_AGENTS_TURN_TIMEOUT_MS

  • PEER_AGENTS_TURN_TIMEOUT_MS — 仅作为 Antigravity 同步回退(默认 300 秒)。不会限制 Grok。 请取消设置它,或显式设置 GROK_TURN_TIMEOUT_MS。希望使用旧的 120 秒 Grok 超时的操作员必须设置 GROK_TURN_TIMEOUT_MS=120000

  • ANTIGRAVITY_TURN_TIMEOUT_MS — 可选的 Antigravity 同步覆盖(默认 300 秒)

  • PEER_AGENTS_GROK_ACP_IDLE_MS — ACP 轮次之间的空闲回收(默认 max(5 min, GROK_TURN_TIMEOUT_MS + 60s))。进行中的 session/prompt 会忽略空闲(promptDepth)。不是 30 分钟的任务生命周期。

  • PEER_AGENTS_JOB_TIMEOUT_MS — 异步任务超时(默认 30 分钟)

  • GROK_JOB_TIMEOUT_MS / ANTIGRAVITY_JOB_TIMEOUT_MS — 可选的按提供方异步覆盖

  • PEER_AGENTS_MAX_PROMPT_CHARS — 提示大小的安全限制(默认 120000);截断时会前置 tools-against-cwd 指令并继续同步轮次

多轮同行会话

每次路由调用都会返回一个 sessionId。使用 peer_turn 继续对话:

  • 告诉同行发生了什么变化

  • 附加新的 diff 或文件

  • 请它重新评审或检查你的修复

会话会持久化到磁盘,因此它们能在 MCP 服务器重启后继续存在。

当第一轮捕获到对话/会话 id 时,Grok 和 Antigravity 的多轮轮次优先使用 原生 CLI 恢复;否则 MCP 会将最近的对话记录重新 hydrate 到提示中。

Grok CLI 集成(1.0.x+)

Grok 同行轮次在底层使用现代 headless 标志(调用者无需传递这些标志):

关注点

行为

大型提示词

始终使用 --prompt-file(避免 argv 限制)

多轮对话

可用时使用 --resume <nativeSessionId>;冷启动在生成前创建 --session-id;回退到 MCP 转录重水合

审查者 / 批评者

--sandbox read-only--always-approve,拒绝编辑工具,无网络搜索,--no-plan--no-subagents

规划者

--sandbox read-only--permission-mode plan--always-approve

实施者

--sandbox workspace--always-approve

权限

审查者/批评者/规划者使用 --always-approve(审查者绝不使用 --permission-mode default),这样非 TTY 的 MCP 子进程不会等待点击。保留只读沙箱、--disallowed-tools search_replace,write 以及 --deny 破坏性 bash。

peer_implement_async

默认通过 git worktree add + --cwd 实现 git worktree 隔离(use_worktree: false 可选择退出)。Grok 1.0 headless 忽略 --worktree

审查结果

尽力从最终文本中解析 findings JSON;纯文本也是有效的审查。Grok headless 不会传递 --json-schema(该标志在 1.0.5 上会中止工具循环)。

风险 / 安全

提高 --effort;额外的自校验 --rules(Grok 1.0 移除了 --check)。高 effort 是使用 durationAdvisory / *_async 的理由,而非放弃 --effort high 的理由。

专家

为安全审查 / 架构规划打包 --agent

同步 + 异步输出

在 Grok headless 上始终使用 --output-format streaming-json(如果 CLI 仍输出单个对象,则回退到 json-envelope)。在 peer_job_status 上使用 progress

ACP 池(选择加入)

PEER_AGENTS_GROK_TRANSPORT=acp 热进程复用。空闲是轮次之间的后备机制,而非作业生命周期;进行中的 session/prompt 会忽略它。

支出遥测

在结果上使用 metricsusagenum_turnsstopReason,以及存在时的 cost)

Antigravity CLI 集成(agy 1.1.8+)

Antigravity peer 轮次在底层使用 print 模式(调用方不传递这些标志):

关注点

行为

调用

agy -p … --print-timeout … --dangerously-skip-permissions --output-format json --disable-slash-commands

多轮对话

--conversation <id> 取自 json 的 conversation_id*.db 的目录快照仅作为回退)

审查结果

为审查者/批评者提供 --json-schema 结构化 findings;structured_output 映射到 structured

审查者 / 批评者

--sandbox

规划者

--sandbox --mode plan

实施者

--mode accept-edits

风险

--effort 来自与 Grok 相同的风险/复杂度映射

工作区

当设置了仓库路径时使用 --add-dir <cwd>

代理

提供时可选 --agent

Slash 命令/技能

始终使用 --disable-slash-commands,这样 peer 提示词无法展开 /commands

异步进度

在异步作业上使用 --output-format stream-json + 在 peer_job_status 上使用 progress

健康检查

优先使用 agy models;回退到一次简短的 pong 轮次

支出遥测

来自 json 信封的 metricsusagenum_turns

agy 仍然没有 worktree、--prompt-file 或 ACP 传输。同步轮次保持 --output-format json;只有后台作业使用流式输出。

设计说明

  • 服务器本身从不修改你的仓库——它只运行你已有的 CLI。

  • 会话转录中的用户消息从调用方的视角进行标记(通常是 "Codex")。

  • 支持幂等键,因此使用相同键的重复调用是安全的。

  • 当输入看起来过于单薄(缺少文件、diff 等)时,会返回上下文质量提示。

  • 实施交接默认使用隔离的 git worktree(--cwd 进入其中),这样 peer 不会破坏脏的主树。

许可证

MIT(或按仓库中的指定)。

Available Tools

13 tools
peer_askC

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. General knowledge or grounded Q&A — routes to Antigravity.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionYesRequired. The decision or question, plus what constraints and tradeoffs the answer must address.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
contextNoBackground the peer needs: prior decisions, relevant code paths, docs links, or constraints.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
taskNoHuman-readable session label: what you are trying to achieve, affected behavior, and specific concerns for the peer.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the routing behavior to Antigravity and mentions automatic staging of binaries. However, it omits details on mutation, authentication, rate limits, or side effects beyond the routing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences but dense and run-on. It mixes instructions, conditions, and routing info in a stream-like manner. Could be better organized with bullet points or clearer separation of purpose vs. usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, no output schema, and complex routing behavior, the description should explain what the tool returns or the outcome. It lacks any mention of return value, response format, or post-call state, leaving the agent uncertain about what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all parameters. The description adds minimal extra meaning (e.g., 'Binary attachments are staged to disk...' in the files parameter is already covered by schema). Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description focuses on pre-call instructions rather than stating the tool's purpose. It vaguely mentions 'General knowledge or grounded Q&A — routes to Antigravity', but the primary verb and resource are unclear. It does not effectively distinguish from sibling tools like peer_debate or peer_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit before-calling steps (read files, attach content, set task) and mentions routing to Antigravity for certain queries. However, it lacks a clear 'when to use' vs 'when not to use' and does not reference specific sibling tools as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_compareA

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Low-level dual-CLI comparison (prefer phase tools for routing).

ParametersJSON Schema
NameRequiredDescriptionDefault
messageYesExact question both peers must answer independently.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
taskYesShort label for this comparison session.
providersNo
diffNoFull unified diff or patch output. Never substitute a prose summary for the actual diff.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
modeNo
systemNo
parallelNo
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility. It mentions 'dual-CLI comparison' but does not explain the internal behavior, such as whether it modifies state, runs external processes, or returns results. The focus is on input preparation rather than tool effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a few short sentences with clear front-loading of critical upfront instructions. Every sentence adds value, with no redundant or verbose phrases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no output schema, and no annotations, the description lacks detail on return values, completion behavior, or what happens after invocation. It explains preparation well but omits post-call context, leaving the agent unsure of the tool's overall operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description adds some context, e.g., for 'files' it says 'use correct file extensions...pass base64 or data-URI content.' However, it largely reiterates schema descriptions for parameters like message and repo_path, not adding substantial new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it is a 'low-level dual-CLI comparison' and instructs to attach files and diffs. It clearly indicates a comparison function, distinguishing it from sibling tools like peer_ask or peer_debate, but does not fully articulate the specific verb-resource relationship.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit instructions: 'Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries.' It also advises to 'prefer phase tools for routing,' guiding when not to use this tool. This gives clear usage context and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_debateA

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Independently compare Plan A vs Plan B without cross-contamination.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesWhat decision this debate must resolve and what success looks like.
plan_aYesFull Plan A: steps, tradeoffs, risks, and verification approach.
plan_bYesFull Plan B: steps, tradeoffs, risks, and verification approach.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
risk_levelNoUse high for auth, payments, migrations, concurrency, and public API changes.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden for behavioral disclosure. It explains the need for independent comparison and idempotency key for retries, but does not describe side effects, output format, or behavior under error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences) and front-loaded with 'Before calling:', which is clear. However, it packs multiple instructions into a single paragraph, slightly reducing readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description does not explain what the tool returns or how results are presented. It covers preparation well but lacks information on the tool's output, which is needed for complete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, baseline 3. The description adds meaning by instructing how to use risk_level (e.g., high for auth, payments) and advising on setting task with specific concerns. This goes beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is for independently comparing Plan A vs Plan B, which aligns with the tool name 'peer_debate'. It specifies the input requirements (files, diffs, task) but does not explicitly distinguish from sibling tools like peer_compare.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides prerequisites and instructions for preparing inputs (read files, attach contents, set task with goals). However, it lacks explicit guidance on when to use this tool versus alternatives (e.g., peer_compare) and does not mention when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_debugA

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Route a debugging request after failures.

ParametersJSON Schema
NameRequiredDescriptionDefault
error_logYesRequired. Full stderr, stack trace, assertion text, and failing test output — not a one-line summary.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
attempted_fixesNoEverything already tried and why each failed. Required when failed_attempts > 0.
failed_attemptsNoHow many fix attempts have already failed on this bug.
diffNoFull unified diff or patch output. Never substitute a prose summary for the actual diff.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
taskNoHuman-readable session label: what you are trying to achieve, affected behavior, and specific concerns for the peer.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It does not mention whether the tool is read-only or destructive, what side effects occur (e.g., file stageing), or any required permissions. The only hint is the presence of an idempotency key, but its significance is not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear guidance starting with 'Before calling', and each sentence serves a purpose. It is not overly verbose, though it could be slightly more concise by reducing repetitions (e.g., 'Pass complete diffs/logs' and later 'Full unified diff or patch output'). Overall, it is well-organized and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters and no output schema, the description covers the input side well but fails to explain what the tool returns or how to interpret the response. It does not address expected behavior after routing the request (e.g., synchronous vs. asynchronous, result format). This leaves the agent without a complete picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters with descriptions, so the baseline is 3. The description adds valuable emphasis on how parameters should be used (e.g., 'Pass complete diffs/logs — never prose summaries', 'Set task with goals, affected behavior...'), which goes beyond the schema by providing behavioral instructions that improve correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is for routing a debugging request after failures, and it distinguishes itself from sibling tools like peer_ask or peer_plan by specifying the context of debugging. The verb 'route' combined with 'debugging request' defines a specific action and resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the user when to use the tool ('after failures') and provides preparatory instructions ('Before calling: read relevant source files...'). It does not explicitly state when not to use it or list alternatives, but the context is clear enough for an agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_healthA

Check whether Grok and Antigravity CLIs are responsive

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It states the action (check responsiveness) but lacks details on side effects (presumably none), required permissions, or output format. The verb 'check' implies read-only, but the description does not explicitly confirm safety or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no filler. Every word contributes to the purpose. Front-loaded with the key action and target.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description covers the essential behavior. However, it could briefly mention how the result is returned (e.g., boolean status) to avoid ambiguity. Still, the simplicity makes it mostly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters (schema coverage 100% vacuously). Description adds no parameter info because none exist. Baseline for 0 parameters is 4, and the description does not need to elaborate further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Check' and the specific resources 'Grok and Antigravity CLIs', making the tool's purpose immediately clear. It distinguishes itself from sibling tools (e.g., peer_ask, peer_debate) which involve querying or discussing, while this is a simple health check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage (when you need to check CLI responsiveness) but provides no explicit guidance on when to use alternatives or when not to use this tool. Given the many sibling tools, some contextual tips would help, but the use case is straightforward.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_list_sessionsB

List persisted peer sessions, optionally filtered by repo path

ParametersJSON Schema
NameRequiredDescriptionDefault
repo_pathNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description adds minimal behavioral context beyond 'List' implying read-only. Does not disclose auth requirements, rate limits, or any side effects, though listing is inherently low-risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, 10 words, front-loaded with verb and resource. No wasted text; everything earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and only one parameter, the description omits return format, pagination, or any additional context about sessions. Incomplete for a tool with no other structured context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description partially compensates by stating 'repo_path' is an optional filter, but does not explain what the path refers to (e.g., local filesystem vs. remote repository).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action ('List') and the resource ('persisted peer sessions') with a specific optional filter ('by repo path'), which distinguishes it from sibling tools like peer_ask or peer_reset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, such as checking existing sessions before starting a new one or debugging. Missing when-not-to-use scenarios or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_planA

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Route an implementation planning request to the best peer model(s).

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesRequired. Goal, success criteria, affected modules, and what 'done' looks like.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
constraintsNoHard limits: API compatibility, performance budgets, forbidden approaches, deadlines, out-of-scope.
repo_summaryNoHow the repo is structured today — key modules, patterns, and entry points relevant to this task.
risk_levelNoUse high for auth, payments, migrations, concurrency, and public API changes.
complexityNocomplex when the change spans multiple modules, data paths, or deployment steps.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses only that the tool routes to a peer model, but lacks details on side effects, permissions, rate limits, return values, or whether state is modified. This is insufficient for a tool with 8 parameters and no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, concise and efficient. The first sentence sets a precondition, the second gives guidance, and the third states the purpose. However, the core purpose is last, not front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters, 3 required, and no output schema, the description is brief and does not explain return values, error handling, or what happens after routing. It lacks completeness for a planning tool that delegates to another model.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions. The description adds context for `files` (attach full contents, never prose summaries) and `task` (goals, affected behavior), but does not mention other parameters like `repo_path`, `constraints`, `repo_summary`, `risk_level`, `complexity`, or `idempotency_key`. It adds moderate value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states this tool routes an implementation planning request to the best peer model, using verbs like 'Route' and 'planning request'. It distinguishes from sibling tools that focus on asking, debating, or debugging, though it doesn't explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit pre-call instructions: read source files, attach full contents via `files`, pass diffs/logs, and set `task` with goals. It also prohibits prose summaries, offering clear guidance on proper usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_resetC

Clear a session transcript or delete the session entirely

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.
expected_versionNo
keep_metadataNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Destructive operation is implied but no details on irreversibility, auth requirements, or what exactly gets cleared/deleted. No annotations to supplement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Too short and vague to convey essential context. Single sentence fails to explain tool behavior sufficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Missing output schema, no annotations, 4 parameters mostly undocumented; description incomplete for safe effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 25% (only idempotency_key described). Description adds no parameter meaning; does not compensate for missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it clears or deletes a session, but doesn't specify which condition triggers which action. It distinguishes from siblings as reset/delete operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus other peer tools, no context about prerequisites or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_review_diffB

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Route a diff review to the best peer model(s) based on focus and risk.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffYesRequired. Full unified diff (`git diff`, `git diff --cached`, or patch file). Do not summarize.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
focusNoPrimary review lens. Pair with a detailed diff, related files, and a rich `task` describing risks.
risk_levelNoUse high for auth, payments, migrations, concurrency, and public API changes.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
taskNoHuman-readable session label: what you are trying to achieve, affected behavior, and specific concerns for the peer.
needs_speedNoPrefer a faster peer when true; still include full context.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior but only mentions 'route to best peer model' without explaining routing logic, side effects, output format, or idempotency implications. The idempotency_key parameter suggests retry safety, but the description is silent on this.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero wasted words. It front-loads the most critical instruction ('Before calling: read...') and provides crisp, actionable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters (3 required) and no output schema, the description covers only diff, files, and task. It omits context for repo_path, focus, risk_level, needs_speed, and idempotency_key, leaving the agent to rely solely on schema descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description reinforces 'files' and 'task' usage (e.g., full contents, specific concerns) but adds no new meaning beyond the schema. It does not elaborate on repo_path, focus, risk_level, or other parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool routes a diff review to peer models based on focus and risk. However, it does not differentiate from sibling tools like peer_debate or peer_compare, leaving ambiguity about when each is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit prerequisites (read files, attach full contents, pass diffs, set task) and a 'never prose summaries' rule. It gives clear context for use but lacks exclusion criteria or comparison to the extensive list of peer sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_summarizeC

Return the rolling session summary and unresolved issues

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully convey behavioral traits. It states 'Return' implying read-only, but does not confirm safety, auth needs, or whether the summary is stateful. The term 'rolling' introduces ambiguity about session state management.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, achieving conciseness. However, it is so brief that it underspecifies the tool, sacrificing clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description should provide more context about the return format, the nature of the summary, and how unresolved issues are defined. It fails to do so, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'session_id' has no schema description (0% coverage), and the tool description does not explain its purpose or format. The description adds no value beyond the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a rolling session summary and unresolved issues, which is distinct from sibling tools like peer_ask or peer_debate. However, it does not elaborate on what 'rolling' means or the context of sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks any mention of use cases, prerequisites, or when to avoid it, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_transcriptC

Export recent transcript turns

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
max_turnsNo
formatNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden but only states 'export', implying a read operation. It does not disclose permission requirements, mutability, or side effects. No information on output format or behavior for edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise at 4 words, front-loading the core action. However, conciseness sacrifices necessary detail, making it under-specified for reliable use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters and no output schema or annotations, the description lacks information on 'recent' semantics, output structure, pagination, or error handling. Incomplete for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%; description adds no meaning to parameters beyond their names. The required session_id, optional max_turns, and enum format are not explained. No examples or clarifications provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'export' and resource 'transcript turns', indicating functionality. It distinguishes from siblings like peer_list_sessions (list sessions) and peer_summarize (summarize) by focusing on exporting recent turns. However, it lacks explicit differentiation from similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like peer_turn, peer_list_sessions, or peer_summarize. Missing context about prerequisites or scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_turnC

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Follow up in an existing routed peer session (use session_id from a prior result).

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYessession_id from results.<cli>.sessionId in a prior routed tool response.
messageYesWhat changed since the last turn, what you fixed, and what you want re-checked.
diffNoFull unified diff or patch output. Never substitute a prose summary for the actual diff.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.
expected_versionNoPass version from the last turn to avoid stale-session races.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It tells the agent what to prepare but not what the tool does internally (e.g., how it processes the turn, side effects, response format). The behavioral impact is largely hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively short but contains extraneous instructions that are more procedural than definitional. The mismatch between 'task' and 'message' wastes some clarity. Overall, it is acceptably concise but could be more focused.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a peer review tool with 6 parameters and no output schema, the description is incomplete. It fails to explain what the tool returns, how to continue the session, or what agents should expect after calling peer_turn. Key contextual gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds some context by instructing to attach full contents via `files` and never use prose summaries for diffs. However, it incorrectly refers to a 'task' parameter that does not exist in the schema, which detracts from semantic clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description implies the tool is for following up in an existing peer session, but it does not explicitly state that it sends a new turn in a peer review. The phrase 'Set `task` with goals, affected behavior, and specific concerns' conflicts with the schema, which uses 'message' instead of 'task'. This mismatch reduces clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides procedural instructions (read files, attach diffs) but does not explain when to use this tool versus sibling tools like peer_ask or peer_debate. There is no guidance on alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

peer_verifyB

Before calling: read relevant source files and attach full contents via files. Pass complete diffs/logs — never prose summaries. Set task with goals, affected behavior, and specific concerns. Route verification of tests/build output.

ParametersJSON Schema
NameRequiredDescriptionDefault
test_outputYesRequired. Complete test runner or build output, including failures, skips, and timing if relevant.
repo_pathYesAbsolute path to the repository root the peer should work in (e.g. /home/user/my-app).
diffNoFull unified diff or patch output. Never substitute a prose summary for the actual diff.
filesNoChanged source files and binary attachments (screenshots, PDFs). Use correct file extensions for images/PDFs and pass base64 or data-URI content.
taskNoHuman-readable session label: what you are trying to achieve, affected behavior, and specific concerns for the peer.
risk_levelNoUse high for auth, payments, migrations, concurrency, and public API changes.
idempotency_keyYesStable key for this operation (e.g. review-auth-jwt-1). Reuse the same key when retrying after timeout.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses prerequisites and input requirements but fails to mention side effects, return values, error conditions, or any behavioral impacts beyond 'route verification.' This is insufficient for safe tool invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with four sentences and no wasted words. It is front-loaded with critical instructions, making it efficient for quick reading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having 7 parameters and no output schema, the description omits key context: what the tool returns, error handling, idempotency key usage, and prerequisites like repo path accessibility. It feels incomplete for a tool that likely performs significant actions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description reinforces proper usage (e.g., 'never prose summaries' for diff) but does not add new semantic meaning beyond the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Route verification of tests/build output,' which clearly indicates the tool's purpose. However, it does not explicitly differentiate from sibling tools like peer_review_diff or peer_debug, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit instructions on how to prepare inputs ('read relevant source files and attach full contents via files', 'pass complete diffs/logs — never prose summaries', 'set task with goals'). It implies usage for verification tasks but lacks explicit when-not-to-use guidance or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 13 tool updatesv0.1.0
    • First observedpeer_ask
    • First observedpeer_compare
    • First observedpeer_debate
    • First observedpeer_debug
    • First observedpeer_health
    • First observedpeer_list_sessions
    • First observedpeer_plan
    • First observedpeer_reset
    • First observedpeer_review_diff
    • First observedpeer_summarize
    • First observedpeer_transcript
    • First observedpeer_turn
    • First observedpeer_verify

TDQS

B3.2/5.0
Disambiguation3/5

While each tool has a distinct purpose, the descriptions share extensive boilerplate text (e.g., 'Before calling: read relevant source files...'), making it harder for an agent to quickly differentiate between tools like peer_ask, peer_compare, and peer_debate. The specific routing information at the end helps but requires careful reading.

Naming Consistency5/5

All tools follow the snakename pattern with the consistent prefix 'peer', using varied but appropriate verbs/verb phrases. No mixing of conventions like camelCase or different prefixes.

Tool Count5/5

13 tools is well within the ideal 3-15 range for a server focused on peer agent interactions. Each tool covers a distinct operation without unnecessary bloat or missing essentials.

Completeness4/5

The tool set covers a broad range of peer agent workflows: asking, comparing, debating, debugging, planning, reviewing, verifying, and session management. Minor gaps like explicit session creation are implicitly handled via peer_turn, so no critical missing operations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Rakeen70210/peer-agents-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server