Skip to main content
Glama

grok-build-mcp-server

npm MCP Registry CI Node License

VS Codeにインストール Cursorにインストール

これは、Grok Build CLI(grok)を、Claude Code、Cursor、VS Code、その他のMCPクライアントから呼び出せるツールとして公開するMCP stdioサーバーです。

Claude Code  ──stdio/MCP──▶  grok-build-mcp-server  ──spawn──▶  grok CLI  ──▶  xAI API

これは軽量なプロセスラッパーです。エージェントロジックを再実装するものではなく、xAI APIと直接通信することもありません。すべてのインテリジェンスはgrok CLIに委ねられています。このサーバーが追加するのは、忠実な引数構築、堅牢なプロセス監視、そしてクリーンなMCP形式の出力です。

ステータス: 0.2.2. ツールの表面は完成しています。このサーバーは、フォアグラウンドまたはバックグラウンドでデタッチされた実際のヘッドレスGrokエージェントを実行し、実行中に進捗をストリーミングし、リクエストに応じて実行を停止し、git差分をレビューし、Web上で質問を調査し、それらの実行によって作成されたセッションを一覧表示し、セッション、使用量、コストを報告します。出荷された内容についてはCHANGELOG.mdを、検討されて却下された内容についてはROADMAP.mdを参照してください。

進捗

長時間のエージェント実行は、画面上で実際に進行状況が確認できるため、沈黙の待機が一気にテキストの壁で終わることはありません。クライアントがprogressTokenを送信すると、サーバーはGrokを--output-format streaming-jsonで実行し、イベントごとに通知を転送します。

#5  list_dir .
#6  read_file README.md
#7  read_file — completed
#8  thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)

進捗は、エージェントがどのフェーズにいるかではなく、何をしているかを追跡します。推論と応答テキストは統合されるため、トークンストリームがクライアントにあふれることはなく、ツール呼び出しは発生時に報告されます。resetTimeoutOnProgressをサポートするクライアントは、実行中にタイムアウトしません。

progressTokenを送信しないクライアントは、より安価な非ストリーミングパスを取得し、これに対して何も支払いません。

Related MCP server: Claude Code MCP Bridge

必要条件

  • Grok Build CLI 1.0.0以上、認証済み(grok modelsが成功すること)

  • Node.js 22以上

grokがPATHにない場合は、サーバーを登録する際にGROK_BINARYをフルパスに設定してください。

インストール

Claude Code

claude mcp add grok-build -- npx -y grok-build-mcp-server

次に、Claude Code内で:

> use the grok-build check tool

checkは、解決されたバイナリ、CLIバージョン、認証されているかどうか、アクティブな許可上限を報告します。正常であれば、残りも機能します。

その他のMCPクライアント

サーバーはstdio経由でMCPを話し、それ自体の引数は取りません。

{
  "mcpServers": {
    "grok-build": {
      "command": "npx",
      "args": ["-y", "grok-build-mcp-server"]
    }
  }
}

VS CodeとCursorは、このページ上部のインストールバッジを受け付けます。これらのバッジには、まさにその構成が含まれています。

MCP Registryからインストールするクライアントは、このサーバーをio.github.Nuruvala/grok-build-mcp-serverとして認識します。レジストリエントリはnpmリリースと同じタグから公開され、同じパッケージを指しています。

npxがサーバーを見つけられない場合

npxは、ベアパッケージ名をローカルプロジェクトから最初に解決します。MCPクライアントの作業ディレクトリがこのリポジトリのチェックアウト、またはpackage.jsonがgrok-build-mcp-serverという名前の他の何かである場合、npx -y grok-build-mcp-serverはローカルのエントリポイントを実行しようとし、それを見つけられず、command not foundで失敗します。独自の場所にインストールし、そのパスを登録してください。

npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-server

権限

このサーバーを介して起動されたGrok実行は、デフォルトで読み取り専用です: --permission-mode planと--sandbox read-only。あなたが許可するまで、何もファイルを変更できません。

権限は上限であり、呼び出しごとにプロンプトを表示するのではなく、サーバーを登録するときに一度だけ設定します。3つのレベルがあります。

レベル

--permission-mode

--sandbox

許可される内容

read-only (デフォルト)

plan

read-only

読み取りと推論。編集は不可。

write

acceptEdits

workspace

作業ディレクトリ内での編集

full

bypassPermissions

off

無人での完全承認

Grokに編集を許可するには:

claude mcp add grok-build \
  -e GROK_MCP_PERMISSION_CEILING=write \
  -e GROK_MCP_DEFAULT_PERMISSION=write \
  -- npx -y grok-build-mcp-server

fullは、すでにMCPクライアントを完全承認で実行しており、委任されたGrok実行も同様に無人にしたい場合にのみ使用してください。これにより、起動されたgrokプロセスに、あなたと同じ権限が付与されます。

上限を超えるリクエストは、黙ってダウングレードされるのではなく、拒否されます。制限された実行は成功を報告しながら何も変更しないため、明確なエラーよりも悪質です。

環境変数

変数

デフォルト

目的

GROK_BINARY

grok

grok実行可能ファイルへのパス

GROK_MCP_PERMISSION_CEILING

read-only

任意の呼び出しが要求できる最高レベル

GROK_MCP_DEFAULT_PERMISSION

read-only

呼び出しが何も要求しない場合に使用されるレベル

GROK_MCP_DEFAULT_MODEL

grok-4.6

呼び出しがモデルを省略した場合のモデル。noneはCLIに委ねる

GROK_MCP_DEFAULT_EFFORT

high

呼び出しが努力量を省略した場合の推論努力。noneはCLIに委ねる

GROK_MCP_TIMEOUT_MS

1800000

1回の実行の壁時計

GROK_MCP_STATE_DIR

$XDG_STATE_HOME/grok-mcp

バックグラウンドジョブのレコード

GROK_MCP_MAX_CONCURRENT_RUNS

4

同時に生存できるバックグラウンド実行数。offで上限なし

GROK_MCP_LOG_LEVEL

info

debug、info、warn、error。ログはstderrに出力されます

STRUCTURED_CONTENT_ENABLED

オフ

_metaに加えてstructuredContentも出力する

Grok自身の変数(XAI_API_KEY、GROK_HOME、GROK_DISABLE_AUTOUPDATER)は、子プロセスにそのまま渡されます。

ツール

ツール

読み取り専用

目的

grok

上限による

ヘッドレスGrokエージェントを実行。プロンプト、セッションの再開/継続/フォーク、モデル、努力、ツールの許可/拒否

review

常に

git差分をレビュー: ワーキングツリー、マージベースの差分(リファレンスとの比較)、または単一のコミット

websearch

常に

Web上で質問を調査し、実際に使用された検索とソースを報告

status

常に

バックグラウンド実行をポーリング、または最近のものを一覧表示

stop

いいえ

バックグラウンド実行のプロセスツリーを終了

sessions

常に

このマシン上のGrokセッションを一覧表示、検索、および参照

check

はい

サーバーバージョン、解決されたバイナリ、grok version、認証、許可上限、実行デフォルト

help

はい

grok --help のパススルー

review

差分はプロセス内で収集され、プロンプトに埋め込まれるため、モデルは何をレビューしようとしているのかを再発見するためにターンを費やす必要がありません。

> review my working tree with grok-build
> review the diff against origin/main

ターゲットはuncommitted、base: "<ref>"(マージベースの差分。ブランチ作成後にベースに取り込まれたコミットはあなたのものとして扱われない)、またはcommit: "<sha>"です。何も指定しない場合は自動検出されます。ブランチが先行している場合は上流の差分、それ以外の場合はワーキングツリーが使用され、黙って推測するのではなく、どちらを選択したかを明示します。

reviewは、GROK_MCP_PERMISSION_CEILINGで許可されている設定に関わらず、常に読み取り専用です。レビュー対象のコードを編集するレビューが必要とされることは決してないため、permission、write、yoloの引数は取りません。

structured: trueを渡すと、_meta.findingsに機械可読な調査結果(severity、file、line、summary、rationale)が、検証後に表示されます。

2つの異なる問題が発生する可能性があり、それらは混同されることなく、異なる方法で報告されます。

  • 実行が完了しなかった — 中断されたか、調査結果を生成せずに終了した。レビューがないため、呼び出しはisError: trueとなり、_meta.findingsCompleteはfalseです。本文は原因を先頭に、CLI自身の理由を引用し、実際の原因に適合する修正方法を挙げます。

  • 実行は完了したが、出力が検証に合格しない。 呼び出しは依然として成功し、生のテキストに加えて_meta.parseErrorを返します。品質が低下したレビューは、失敗したレビューよりも優れています。

絶対に発生しないのは、モデルがでっち上げたもっともらしい調査結果です。--json-schemaはモデルが発するすべてのメッセージを制約するため、モデルがまだ読み取り中の間は、調査結果の形以外で「作業中」と言う方法はありません。そして、チェックを怠ると、まさにそれをやってしまいます。スキーマには必須のstatusフィールドがあり、その説明を結果から除外します。部分的な応答からパターンマッチングで何かを救出することは決してありません。

大規模なターゲットに対する構造化レビューは、この方法でかなりの頻度で失敗します。この失敗は設計上明白です。

シェルに手を伸ばすレビューは拒否され、強制終了されるのではありません。ヘッドレスモードでは、承認不可能なツールリクエストはCLIが終了コード0で終了する間に実行全体をキャンセルするため、reviewはシェルおよび編集ツールを完全に拒否します。モデルには「いいえ」と伝えられ、文の途中で死ぬ代わりにレビューを完了します。

websearch

> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depth

numResults(1~50)とsearchDepth(basicまたはfull)はプロンプトを形成します。grok CLIにはこれらのフラグはなく、どちらのパラメータもそのように偽装しません。これらは実際に機能します。同じ質問をbasicで行うと2ページにわたって1回の検索が行われ、fullで行うと3ページにわたって6回の検索が行われ、コストは2.5倍になりました。

結果は、モデルが書いたものだけでなく、実際に何が検索されたかを示します。

[1 web search, 9 sources]

_metaにはwebSearches、webToolCalls、searchQueries、sources、sourceCount、pagesOpened、searchPerformedが含まれます。これは聞こえ以上に重要です。GrokはWeb検索またはXを通じて調査を行うことができ、Webが利用できない場合は静かに後者を行います。自信満々に答え、x.comを引用し、正常終了します。散文ではそれを見分ける方法がありません。したがって、Xを検索してWebを検索しなかった実行は、最初の行でその旨を述べ、xSearchesを別途報告します。また、何も返ってこなかった実行は、モデル自身の記憶から得た自信満々に見える回答ではなく、エラーとして扱われます。

No search ran. The answer below is the model's own prior knowledge, not current sources.

searchPerformedは、ソースが返ってきたことを意味します。検索が試行されたことではありません。開始されたが何も返ってこなかった検索、または空の結果セットを返した検索は、そのまま報告されます。

review と同様に、websearch は常に読み取り専用であり、permission、write、yolo のいずれの引数も受け取りません。--disable-web-search を渡すこともありません。

バックグラウンド実行、status、stop

長いエージェント実行はクライアントを占有する必要はありません。grok、review、 または websearch に background: true を渡すと、呼び出しはすぐに runId を返し、その間、切り離されたワーカープロセスが ジョブを完了まで実行します:

> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that run

実行はこのサーバーではなくマシンに属します。MCP クライアントが切断しても、 サーバーが再起動しても、またはエディタを閉じても実行は継続します。レコードは GROK_MCP_STATE_DIR 配下、 実行ごとに 1 つのディレクトリに格納されます。

完了した実行に対する status は、同期呼び出しが返したであろうものを返します — 同じテキスト、 同じメタデータ、同じエラーフラグ。バックグラウンドはツール呼び出しのためのトランスポートであり、 別の実装ではありません。実行が進行中の場合、その状態、経過時間、両方のプロセス ID、および 進行ログの末尾を取得できます。waitMs は最大 2 分間ブロックし、進行通知が届くたびに転送します。 待機がタイムアウトしてもエラーではありません。

2 種類の偽りは構造的に排除されています。ワーカープロセスが存在しなくなった実行は、 まだ実行中としてではなく abandoned として報告されます — マシンが再起動したか、何かがそれを 強制終了したかのいずれかです。そして、早期に終了した実行はそのようにラベル付けされます:

mfk2p1x9-3ac71f0b  completed (cut off: cancelled)  grok  4m 12s  refactor the parser

runId を受け取る前に検証は依然として行われます:GROK_MCP_PERMISSION_CEILING を超えるリクエスト、 または矛盾するセッションフラグの組み合わせは、受け入れられてから誰も監視していないプロセスで 失敗するのではなく、失敗した呼び出しとして拒否されます。

stop は実行を早期に終了します。ワーカーのプロセスグループ全体 — ワーカーとそれが起動した grok プロセス — に SIGTERM を送信し、それで不十分な場合は SIGKILL を送信します。すでに終了した 実行を停止してもエラーにはなりません。呼び出しが届く直前に終了した実行を停止しても同様です。

プロセスツリーを強制終了できなかった停止は、停止した実行としてではなく、失敗として報告されます。 シグナルを送る対象がない場合、強制終了が拒否された場合、またはツリーが SIGKILL を生き延びた場合、 実行は running のままとなり、呼び出しは pid を名指ししたエラーを返します。生存中のプロセスの隣に cancelled レコードがあるのは、よりすっきりした答えであり、役に立たない答えでもあります。

途中で停止した実行は、通常すでに保存する価値のあるものを生成しており、部分的な結果と セッション ID の両方が保持されます:

Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.

The run was cancelled mid-flight, but it recorded a session before it ended:
  grok -r 01a010e2-478c-73d2-bce9-23552245c64d

Grok は実行が終了に達したときのみセッション ID を報告しますが、停止された実行は終了に達することはありません — そのため、その ID は再構築されるのではなく、CLI 自身のセッションストアから読み戻されます。_meta.sessionIdSource はどちらの方法で取得したかを示します。同じディレクトリ内の 2 つの実行が両方とも一致する可能性がある場合、 候補 ID が返され、再開コマンドは返されません:誤ったセッションを再開すると、他人の作業を継続することになります。

sessions

すべての Grok 実行はディスク上にセッションを残し、このサーバーが報告するすべてのセッション ID は 後で再開できます — 任意のディレクトリから、ターミナルで直接、または別のツール呼び出しによって。

> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiter

セッションは $GROK_HOME/sessions (デフォルトは ~/.grok/sessions) から読み取られます。これは CLI 自身の ストアであり、このサーバー、MCP クライアント、マシンの再起動後も存続します。1 つのセッションには id を、 タイトル、最初のプロンプト、ID に対する大文字小文字を区別しない検索には query を、1 つのプロジェクトに スコープを絞るには cwd を、リストの上限を設定するには limit を渡します。

終了したばかりの実行にはまだタイトルがありません — Grok は後で埋めることもあります — そのため、行は セッションの最初のプロンプトにフォールバックし、titleSource はどちらを見ているかを示します。すべての 行に resumeCommand が含まれ、すべての grok と review の結果にも含まれます:

grok -r 01a00c8d-970c-7531-8a12-31dac582c22b

検索はローカルのみです。grok sessions search はリモートインデックスも参照しますが、このツールは 参照しないため、サーバー側にのみ存在するセッションは表示されません。

開発

npm install
npm run build          # tsc -> dist/
npm run dev            # tsx src/index.ts
npm test               # node --test via tsx
npm run test:coverage  # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run format
  • docs/api-reference.md — 各ツールのパラメータ、結果テキスト、_meta キー、およびそれぞれが設定される正確な条件。

  • docs/security.md — このサーバーを登録することで何が許可されるか、各権限レベルが 実際に何を付与するか、そして何がマシンから外部に出るか。

  • docs/engineering.md — ここでのコードの書き方:アーキテクチャ、関数型 TypeScript のルール、エラーとエフェクトの規律、テストとカバレッジの方針、コミットワークフロー。

  • CLAUDE.md — プロジェクトの背景と、このサーバーが依存する検証済みの grok CLI の動作。

  • ROADMAP.md — マイルストーン、受け入れ基準、そして評価され却下されたアイデア。

リリース

package.json の version を上げ、CHANGELOG.md の Unreleased セクションを 新しいバージョンの見出しの下に移動し、コミットしてから:

git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0

.github/workflows/release.yml は完全なゲートを実行し、タグと package.json が一致しない場合は公開を拒否し、パッケージ化された tarball をスクラッチディレクトリに インストールし、インストールされたバイナリに対して実際の initialize を実行してから、その同じファイル を 公開し、GitHub リリースを作成します。

管理すべき公開用の資格情報はありません。認証は npm trusted publishing です:ワークフローは 短命の OIDC トークンを交換し、npm は独自に来歴証明を生成します。信頼は このリポジトリとこのワークフローの ファイル名 に対して登録されているため、release.yml の名前を変更すると公開が壊れます — そして npm は公開が試行されるまで設定をチェックせず、そのときの 症状は原因を特定できるものではなく ENEEDAUTH です。

ライセンス

MIT — LICENSE を参照してください。

Available Tools

8 tools
checkCheck Grok Build readinessA
Read-onlyIdempotent

Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokRun Grok BuildA

Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`.
denyNoRepeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`.
yoloNoShorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request.
agentNoNamed subagent to run, passed as `--agent`.
allowNoRepeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
rulesNoExtra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt.
toolsNoInternal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`.
writeNoShorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
promptYesThe task for Grok to perform. Passed verbatim as `grok -p`.
resumeNoResume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
sessionIdNoCreate a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
permissionNoPermission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change.
forkSessionNoUUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`.
continueSessionNoContinue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request.
disallowedToolsNoInternal tool ids to block, passed as `--disallowed-tools`.
disableWebSearchNoPass `--disable-web-search`. `false` is not a request.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

helpGrok CLI helpA
Read-onlyIdempotent

Show the grok CLI help text. Runs grok --help and returns its stdout.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reviewReview a git diffA
Read-only

Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Repository to review. Defaults to the current working directory.
baseNoReview the merge-base diff against this ref. Mutually exclusive with commit and uncommitted.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
commitNoReview this commit. Mutually exclusive with base and uncommitted.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
structuredNoReturn machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request.
uncommittedNoReview the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request.
instructionsNoExtra reviewer guidance, appended verbatim to the prompt.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sessionsList Grok sessionsA
Read-onlyIdempotent

List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoExact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match.
cwdNoKeep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`).
limitNoMaximum rows to return. Default 20. Ignored when `id` is set.
queryNoCase-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusPoll a background runA
Read-onlyIdempotent

Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
tailNoBytes of progress.log to include for a live run. Default 8192.
limitNoMaximum rows to return in list mode. Default 20. Ignored when `runId` is set.
runIdNoId of a background run to inspect. Omit to list recent runs.
waitMsNoBlock up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stopStop a background runA
DestructiveIdempotent

Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesThe runId returned by a background `grok`, `review`, or `websearch` call.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

websearchSearch the web with Grok BuildA
Read-only

Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
queryYesThe question to research. Passed as the body of a web-search-shaped prompt.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
numResultsNoPrompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag.
searchDepthNoPrompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag.
instructionsNoExtra researcher guidance, appended verbatim to the prompt.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.2.4
    • Changedgrok2 fields changed
      • changedInput schema / properties / cwd / description
        Previous value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`."
      • changedInput schema / properties / permission / description
        Previous value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
  2. 8 tool updatesv0.2.2
    • First observedcheck
    • First observedgrok
    • First observedhelp
    • First observedreview
    • First observedsessions
    • First observedstatus
    • First observedstop
    • First observedwebsearch

TDQS

A4.5/5.0

Scored across 8 tools

Disambiguation5/5

Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.

Naming Consistency4/5

All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.

Tool Count5/5

Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.

Completeness5/5

The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers