grok-build-mcp-server
grok-build-mcp-server
一个 MCP stdio 服务器,它将 Grok Build CLI(grok)作为工具暴露出来,供 Claude Code、Cursor、VS Code 或任何其他 MCP 客户端调用。
Claude Code ──stdio/MCP──▶ grok-build-mcp-server ──spawn──▶ grok CLI ──▶ xAI API它是一个轻量级进程包装器。它不重新实现代理逻辑,也不直接与 xAI API 通信——所有智能都保留在 grok CLI 中。这个服务器增加的是忠实的参数构建、健壮的进程监督以及干净的 MCP 格式输出。
状态:0.2.2。 工具表面已完整。该服务器在前台或后台分离模式下运行真正的无头 Grok 代理,在运行时流式传输进度,按请求停止运行,审查 git 差异,在网络上研究问题,列出这些运行创建的会话,并报告会话、使用情况和成本。请参阅 CHANGELOG.md 了解已发布的内容,以及 ROADMAP.md 了解已考虑和拒绝的内容。
进度
长时间运行的代理在运行时是可见的,而不是静默等待后出现一堵文本墙。当您的客户端发送 progressToken 时,服务器使用 --output-format streaming-json 运行 Grok,并为每个事件转发一个通知:
#5 list_dir .
#6 read_file README.md
#7 read_file — completed
#8 thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)进度跟踪的是代理正在做什么,而不是它处于哪个阶段。推理和响应文本被合并,因此令牌流不会淹没您的客户端,而工具调用会在发生时报告。支持 resetTimeoutOnProgress 的客户端不会在运行中途超时。
不发送 progressToken 的客户端将使用更便宜的非流式路径,并且无需为此付出任何代价。
Related MCP server: Claude Code MCP Bridge
要求
Grok Build CLI 1.0.0 或更新版本,已认证(
grok models应成功)Node.js 22 或更新版本
如果 grok 不在您的 PATH 中,请在注册服务器时将 GROK_BINARY 设置为其完整路径。
安装
Claude Code
claude mcp add grok-build -- npx -y grok-build-mcp-server然后,在 Claude Code 中:
> use the grok-build check toolcheck 报告已解析的二进制文件、CLI 版本、是否已认证以及活动的权限上限。如果一切正常,其余功能将正常工作。
任何其他 MCP 客户端
该服务器通过 stdio 使用 MCP 协议,并且不接受任何自己的参数:
{
"mcpServers": {
"grok-build": {
"command": "npx",
"args": ["-y", "grok-build-mcp-server"]
}
}
}VS Code 和 Cursor 接受本页顶部的安装徽章,这些徽章携带的正是该配置。
从 MCP Registry 安装的客户端将此服务器识别为 io.github.Nuruvala/grok-build-mcp-server。注册表条目与 npm 版本相同的标签发布,并指向相同的包。
如果 npx 找不到服务器
npx 首先将裸包名解析为本地项目。如果您的 MCP 客户端的工作目录是此仓库的检出——或任何其他 package.json 名为 grok-build-mcp-server 的项目——那么 npx -y grok-build-mcp-server 会运行本地入口点,找不到它,并以 command not found 失败。将其安装到自己的目录并注册该路径:
npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-server权限
通过此服务器启动的 Grok 运行默认是只读的:--permission-mode plan 配合 --sandbox read-only。在您明确允许之前,无法修改您的文件。
权限是一个上限,在注册服务器时设置一次,而不是每次调用时提示。三个级别:
级别 |
|
| 允许的操作 |
|
|
| 读取和推理。无编辑 |
|
|
| 在工作目录内进行编辑 |
|
|
| 无人值守的完全批准 |
要让 Grok 进行编辑:
claude mcp add grok-build \
-e GROK_MCP_PERMISSION_CEILING=write \
-e GROK_MCP_DEFAULT_PERMISSION=write \
-- npx -y grok-build-mcp-server仅当您已经以完全批准模式运行 MCP 客户端,并且希望委托的 Grok 运行同样无人值守时,才使用 full。它授予生成的 grok 进程与您相同的权限。
请求超过上限的调用将被拒绝,而不是静默降级——一个被限制的运行会报告成功但实际未做任何更改,这比一个明确的错误更糟糕。
环境变量
变量 | 默认值 | 用途 |
|
|
|
|
| 任何调用可以请求的最高级别 |
|
| 当调用未请求级别时使用的级别 |
|
| 当调用省略模型时使用的模型。 |
|
| 当调用省略推理努力时使用的努力级别。 |
|
| 单次运行的挂钟时间 |
|
| 后台作业记录 |
|
| 同时存活的后台运行数。 |
|
|
|
| off | 同时输出 |
Grok 自身的变量(XAI_API_KEY、GROK_HOME、GROK_DISABLE_AUTOUPDATER)会原封不动地传递给子进程。
工具
工具 | 只读 | 用途 |
| 取决于上限 | 运行无头 Grok 代理。提示、会话恢复/继续/分叉、模型、努力、工具允许/拒绝 |
| 始终 | 审查 git 差异:工作树、针对某个引用的合并基础差异,或单个提交 |
| 始终 | 在网络上研究问题,并报告实际使用了哪些搜索和来源 |
| 始终 | 轮询后台运行,或列出最近的运行 |
| 否 | 终止后台运行的进程树 |
| 始终 | 列出、搜索和查找此机器上的 Grok 会话 |
| 是 | 服务器版本、已解析的二进制文件、 |
| 是 |
|
review
差异在进程内收集并嵌入到提示中,因此模型无需花费轮次重新发现它要审查的内容。
> review my working tree with grok-build
> review the diff against origin/main目标是 uncommitted、base: "<ref>"(一个合并基础差异,因此您在分支之后落在基础分支上的提交不会归因于您),或 commit: "<sha>"。如果未指定,它会自动检测:当您的分支领先时使用上游差异,否则使用工作树——并且它会说明选择了哪个,而不是静默猜测。
review 始终是只读的,无论 GROK_MCP_PERMISSION_CEILING 允许什么。它不接受 permission、write 或 yolo 参数,因为审查代码时编辑被审查的代码从来都不是想要的。
传递 structured: true 以获取机器可读的发现(severity、file、line、summary、rationale),这些发现位于 _meta.findings 中,并在您看到之前经过验证。
两种不同的事情可能出错,它们会被不同地报告,而不是混为一谈:
运行从未完成——它被中断,或结束而未产生发现。没有审查,因此调用是
isError: true,并且_meta.findingsComplete为false。正文首先说明原因,引用 CLI 自身的理由,并指出适合实际原因的修复方法。运行完成但其输出无法验证。 调用仍然成功,返回原始文本加上
_meta.parseError——降级的审查胜过失败的审查。
您永远不会得到的是模型编造出的看似合理的发现。--json-schema 约束模型发出的每条消息,因此当它仍在读取时,它无法说“我正在工作”,除非以发现的形式——如果不加检查,它确实会这样做。该模式携带一个必需的 status 字段,以将该叙述排除在结果之外,并且永远不会通过模式匹配从部分响应中挽救任何内容。
对大型目标的结构化审查确实会以这种方式失败,且频率不低。这种失败是设计上的响亮。
一个试图使用 shell 的审查会被拒绝,而不是被终止。在无头模式下,一个不可批准的工具请求会取消整个运行,而 CLI 仍然以 0 退出,因此 review 直接拒绝 shell 和编辑工具——模型被告知“不”,并完成其审查,而不是在句子中间死亡。
websearch
> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depthnumResults(1–50)和 searchDepth(basic 或 full)塑造提示——grok CLI 没有针对这两者的标志,并且这两个参数也不假装有。它们确实有效:同一个问题在 basic 下进行了一次搜索,跨越两个页面,而在 full 下进行了六次搜索,跨越三个页面,成本是前者的两倍半。
结果告诉您实际查找了什么,而不仅仅是模型写了什么:
[1 web search, 9 sources]其中 _meta 携带 webSearches、webToolCalls、searchQueries、sources、sourceCount、pagesOpened 和 searchPerformed。这比听起来更重要。Grok 可以通过网络搜索或 X 进行研究,当网络不可用时,它会静默地使用第二种方式——自信地回答,引用 x.com,成功退出。散文部分无法让您区分。因此,一个搜索了 X 而不是网络的运行会在其第一行说明,并单独报告 xSearches,而一个没有任何返回的运行是一个错误,而不是来自模型自身记忆的看似自信的答案:
No search ran. The answer below is the model's own prior knowledge, not current sources.searchPerformed 意味着来源返回了——而不是尝试了搜索。一个开始但从未返回,或返回空结果集的搜索,会被如实报告。
与 review 类似,websearch 始终为只读,不接受 permission、write 或 yolo 参数。它从不传递 --disable-web-search。
后台运行、status 和 stop
长时间运行的代理不必占用你的客户端。向 grok、review 或 websearch 传递 background: true,调用会立即返回一个 runId,同时一个分离的工作进程将任务运行至完成:
> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that run运行属于机器,而非此服务器:即使你的 MCP 客户端断开连接、服务器重启或关闭编辑器,它也会继续运行。记录保存在 GROK_MCP_STATE_DIR 下,每个运行对应一个目录。
对已完成运行的 status 返回与同步调用相同的结果——相同的文本、相同的元数据、相同的错误标志。后台是工具调用的传输方式,而非另一种实现。当运行处于活动状态时,你可以获取其状态、已用时间、两个进程 ID 以及进度日志的尾部;waitMs 最多阻塞两分钟,并在进度通知到达时转发它们。超时等待不是错误。
两种不诚实的情况被设计排除。工作进程不再存在的运行会被报告为 abandoned 而非仍在运行——机器重启或某些东西杀死了它。提前完成的运行也会被标记为:
mfk2p1x9-3ac71f0b completed (cut off: cancelled) grok 4m 12s refactor the parser在获得 runId 之前仍会进行验证:超过 GROK_MCP_PERMISSION_CEILING 的请求,或一对矛盾的会话标志,会被拒绝为失败的调用,而不是被接受然后在无人监控的进程中失败。
stop 提前结束运行。它向工作进程的整个进程组(工作进程及其生成的 grok 进程)发送 SIGTERM 信号,如果不够则发送 SIGKILL。停止已完成的运行不是错误,停止在调用到达前刚刚完成的运行也不是错误。
无法杀死进程树的停止会被报告为失败,而非已停止的运行。 如果没有可发送信号的对象,或杀死被拒绝,或进程树在 SIGKILL 后仍然存活,运行会保持 running 状态,调用返回一个包含 pid 的错误。一个 cancelled 记录与一个活动进程并存,这会是更整洁的答案,但也是无用的。
你在运行中途停止的通常已经产生了一些值得保留的内容,部分结果和会话 ID 都会被保留:
Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.
The run was cancelled mid-flight, but it recorded a session before it ended:
grok -r 01a010e2-478c-73d2-bce9-23552245c64dGrok 仅在运行到达终点时报告会话 ID,而停止的运行永远不会到达终点——因此该 ID 是从 CLI 自身的会话存储中读取的,而非重建的。_meta.sessionIdSource 告诉你你拥有的是哪种。如果同一目录中的两个运行都可能匹配,你会得到候选 ID 且没有恢复命令:恢复错误的会话会继续别人的工作。
sessions
每次 Grok 运行都会在磁盘上留下一个会话,此服务器报告的每个会话 ID 都可以稍后恢复——从任何目录,由你在终端中或通过另一个工具调用。
> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiter会话从 $GROK_HOME/sessions(默认为 ~/.grok/sessions)读取,这是 CLI 自身的存储,因此它们能在此服务器、你的 MCP 客户端以及你的机器重启后继续存在。传递 id 获取单个会话,query 对标题、首个提示和 ID 进行不区分大小写的搜索,cwd 限定到一个项目,limit 限制列表数量。
刚刚完成的运行还没有标题——Grok 稍后会填充它们(如果有的话)——因此行会回退到会话的第一个提示,titleSource 告诉你正在查看的是哪个。每一行都带有 resumeCommand,每个 grok 和 review 结果也是如此:
grok -r 01a00c8d-970c-7531-8a12-31dac582c22b搜索仅限本地。grok sessions search 还会查询远程索引;此工具不会,因此仅存在于服务器端的会话不会出现。
开发
npm install
npm run build # tsc -> dist/
npm run dev # tsx src/index.ts
npm test # node --test via tsx
npm run test:coverage # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run formatdocs/api-reference.md — 每个工具的参数、结果文本、
_meta键以及每个键被设置的确切条件。docs/security.md — 注册此服务器授权了什么,每个权限级别实际授予了什么,以及什么离开了你的机器。
docs/engineering.md — 代码编写方式:架构、函数式 TypeScript 规则、错误和效果纪律、测试和覆盖率策略、提交工作流。
CLAUDE.md — 项目背景以及此服务器依赖的已验证的
grokCLI 行为。ROADMAP.md — 里程碑、验收标准以及经过评估并被拒绝的想法。
发布
在 package.json 中提升 version,将 CHANGELOG.md 的 Unreleased 部分移到新版本标题下,提交,然后:
git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0.github/workflows/release.yml 运行完整门控,如果标签和 package.json 不一致则拒绝发布,将打包的 tarball 安装到临时目录,并对安装的二进制文件执行真实的 initialize,然后发布 同一个文件 并创建 GitHub 发布。
没有需要管理的发布凭据。认证是 npm trusted publishing:工作流交换一个短期 OIDC 令牌,npm 自行生成来源证明。信任是针对此仓库和此工作流的 文件名 注册的,因此重命名 release.yml 会破坏发布——而 npm 直到尝试发布时才会检查配置,此时的症状是 ENEEDAUTH,而不是任何能指出原因的信息。
许可证
MIT — 参见 LICENSE。
Available Tools
8 toolscheckCheck Grok Build readinessARead-onlyIdempotent
Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grokRun Grok BuildA
Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`. | |
| deny | No | Repeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`. | |
| yolo | No | Shorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request. | |
| agent | No | Named subagent to run, passed as `--agent`. | |
| allow | No | Repeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| rules | No | Extra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt. | |
| tools | No | Internal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`. | |
| write | No | Shorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| prompt | Yes | The task for Grok to perform. Passed verbatim as `grok -p`. | |
| resume | No | Resume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| sessionId | No | Create a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| permission | No | Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change. | |
| forkSession | No | UUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`. | |
| continueSession | No | Continue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request. | |
| disallowedTools | No | Internal tool ids to block, passed as `--disallowed-tools`. | |
| disableWebSearch | No | Pass `--disable-web-search`. `false` is not a request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
helpGrok CLI helpARead-onlyIdempotent
Show the grok CLI help text. Runs grok --help and returns its stdout.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviewReview a git diffARead-only
Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Repository to review. Defaults to the current working directory. | |
| base | No | Review the merge-base diff against this ref. Mutually exclusive with commit and uncommitted. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| commit | No | Review this commit. Mutually exclusive with base and uncommitted. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| structured | No | Return machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request. | |
| uncommitted | No | Review the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request. | |
| instructions | No | Extra reviewer guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionsList Grok sessionsARead-onlyIdempotent
List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Exact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match. | |
| cwd | No | Keep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`). | |
| limit | No | Maximum rows to return. Default 20. Ignored when `id` is set. | |
| query | No | Case-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusPoll a background runARead-onlyIdempotent
Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.
| Name | Required | Description | Default |
|---|---|---|---|
| tail | No | Bytes of progress.log to include for a live run. Default 8192. | |
| limit | No | Maximum rows to return in list mode. Default 20. Ignored when `runId` is set. | |
| runId | No | Id of a background run to inspect. Omit to list recent runs. | |
| waitMs | No | Block up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stopStop a background runADestructiveIdempotent
Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | The runId returned by a background `grok`, `review`, or `websearch` call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
websearchSearch the web with Grok BuildARead-only
Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| query | Yes | The question to research. Passed as the body of a web-search-shaped prompt. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| numResults | No | Prompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag. | |
| searchDepth | No | Prompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag. | |
| instructions | No | Extra researcher guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.4- Changed
grok2 fields changed- changed
Input schema / properties / cwd / descriptionPrevious value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`." - changed
Input schema / properties / permission / descriptionPrevious value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
8 tool updates
v0.2.2- First observed
check - First observed
grok - First observed
help - First observed
review - First observed
sessions - First observed
status - First observed
stop - First observed
websearch
TDQS
Scored across 8 tools
Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.
All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.
Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.
The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.
Maintenance
Related MCP Connectors
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Give any MCP-compatible AI assistant a builder for live, hosted web tools and workflows.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables sandboxed file operations via MCP tools, resources, and prompts, with a Claude CLI client and Groq-powered web UI for file CRUD, search, code review, and documentation generation.MIT
- FlicenseNot gradedqualityDmaintenanceExposes Claude Code's file editing, command execution, and test running capabilities as composable MCP tools for any MCP-compatible host, enabling code operations via a stateless bridge.-
- FlicenseAqualityBmaintenanceEnables using the xAI Grok CLI as an MCP sub-agent for code review, asking questions, and continuing conversations within MCP hosts like Claude Code.4-
- AlicenseNot gradedqualityBmaintenanceEnables Codex to use Grok Build CLI as a controlled subagent via MCP tools for independent investigation, review, and isolated implementation tasks.5MIT