grok-build-mcp-server
grok-build-mcp-server
Сервер MCP stdio, который предоставляет CLI Grok Build (grok) в виде инструментов, которые можно вызывать из Claude Code, Cursor, VS Code или любого другого MCP-клиента.
Claude Code ──stdio/MCP──▶ grok-build-mcp-server ──spawn──▶ grok CLI ──▶ xAI APIЭто тонкая обёртка процесса. Она не перереализует логику агента и не общается напрямую с xAI API — весь интеллект остаётся в CLI grok. Что добавляет этот сервер — это точное построение аргументов, надёжный контроль процессов и чистый вывод в формате MCP.
Статус: 0.2.2. Поверхность инструментов завершена. Сервер запускает реальных безголовых агентов Grok в фоновом режиме или в фоне, передаёт прогресс во время их работы, останавливает выполнение по запросу, просматривает git-диффы, исследует вопросы в интернете, выводит список сессий, созданных этими запусками, и сообщает о сессиях, использовании и стоимости. См. CHANGELOG.md о том, что было выпущено, и ROADMAP.md о том, что было рассмотрено и отклонено.
Progress
Длительный запуск агента виден в процессе, а не молчаливое ожидание, заканчивающееся стеной текста. Когда ваш клиент отправляет progressToken, сервер запускает Grok с --output-format streaming-json и пересылает уведомление на каждое событие:
#5 list_dir .
#6 read_file README.md
#7 read_file — completed
#8 thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)Прогресс отслеживает, что делает агент, а не в какой фазе он находится. Текст рассуждений и ответов объединяется, чтобы поток токенов не заливал ваш клиент, а вызовы инструментов сообщаются по мере их выполнения. Клиенты, поддерживающие resetTimeoutOnProgress, не будут тайм-аутиться во время выполнения.
Клиент, который не отправляет progressToken, получает более дешёвый нестриминговый путь и ничего за это не платит.
Related MCP server: Claude Code MCP Bridge
Requirements
CLI Grok Build 1.0.0 или новее, аутентифицирован (
grok modelsдолжен выполняться успешно)Node.js 22 или новее
Если grok нет в вашем PATH, установите GROK_BINARY на полный путь при регистрации сервера.
Install
Claude Code
claude mcp add grok-build -- npx -y grok-build-mcp-serverЗатем в Claude Code:
> use the grok-build check toolcheck сообщает разрешённый бинарник, версию CLI, аутентифицированы ли вы, и активный потолок разрешений. Если он доволен, остальное будет работать.
Any other MCP client
Сервер говорит на MCP через stdio и не принимает собственных аргументов:
{
"mcpServers": {
"grok-build": {
"command": "npx",
"args": ["-y", "grok-build-mcp-server"]
}
}
}VS Code и Cursor принимают значки установки в верхней части этой страницы, которые несут именно эту конфигурацию.
Клиенты, устанавливающие из MCP Registry, знают этот сервер как io.github.Nuruvala/grok-build-mcp-server. Запись в реестре публикуется с того же тега, что и npm-релиз, и указывает на тот же пакет.
If npx cannot find the server
npx сначала разрешает голое имя пакета относительно локального проекта. Если рабочая директория вашего MCP-клиента — это клон этого репозитория или чего-то ещё, чей package.json называется grok-build-mcp-server, то npx -y grok-build-mcp-server запускает локальную точку входа, не находит её и завершается с ошибкой command not found. Установите его в отдельное место и зарегистрируйте этот путь:
npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-serverPermissions
Запуски Grok через этот сервер по умолчанию только для чтения: --permission-mode plan с --sandbox read-only. Ничто не может изменить ваши файлы, пока вы не разрешите.
Разрешение — это потолок, устанавливаемый один раз при регистрации сервера, а не запрос при каждом вызове. Три уровня:
Level |
|
| What it allows |
|
|
| Чтение и рассуждения. Без изменений |
|
|
| Изменения внутри рабочей директории |
|
|
| Полное одобрение без участия пользователя |
Чтобы разрешить Grok вносить изменения:
claude mcp add grok-build \
-e GROK_MCP_PERMISSION_CEILING=write \
-e GROK_MCP_DEFAULT_PERMISSION=write \
-- npx -y grok-build-mcp-serverИспользуйте full только если вы уже запускаете свой MCP-клиент с полным одобрением и хотите, чтобы делегированный запуск Grok был таким же неавтоматизированным. Он предоставляет порождённому процессу grok те же полномочия, что и у вас.
Вызов, запрашивающий больше потолка, отклоняется, а не молча понижается — ограниченный запуск сообщил бы об успехе, ничего не меняя, что хуже, чем чёткая ошибка.
Environment variables
Variable | Default | Purpose |
|
| Путь к исполняемому файлу |
|
| Максимальный уровень, который может запросить любой вызов |
|
| Уровень, используемый, когда вызов не запрашивает никакого |
|
| Модель, когда вызов опускает её. |
|
| Усилие рассуждения, когда вызов опускает его. |
|
| Стенное время для одного запуска |
|
| Записи фоновых заданий |
|
| Фоновые запуски одновременно. |
|
|
|
| off | Также выводить |
Собственные переменные Grok (XAI_API_KEY, GROK_HOME, GROK_DISABLE_AUTOUPDATER) передаются дочернему процессу без изменений.
Tools
Tool | Read-only | Purpose |
| by ceiling | Запустить безголового агента Grok. Подсказка, возобновление/продолжение/форк сессии, модель, усилие, разрешение/запрет инструментов |
| always | Просмотреть git-дифф: рабочее дерево, дифф merge-base относительно ссылки или отдельный коммит |
| always | Исследовать вопрос в интернете и сообщить, какие поиски и источники были фактически использованы |
| always | Опрашивать фоновый запуск или вывести список последних |
| no | Завершить дерево процессов фонового запуска |
| always | Вывести список, искать и просматривать сессии Grok на этой машине |
| yes | Версия сервера, разрешённый бинарник, |
| yes | Передача |
review
Дифф собирается в процессе и встраивается в подсказку, чтобы модель не тратила шаги на повторное обнаружение того, что она должна просмотреть.
> review my working tree with grok-build
> review the diff against origin/mainЦели: uncommitted, base: "<ref>" (дифф merge-base, так что коммиты, попавшие на базу после вашего ответвления, не приписываются вам) или commit: "<sha>". Если не указано, определяется автоматически: дифф с upstream, когда ваша ветка впереди, иначе рабочее дерево — и он сообщает, что выбрал, а не угадывает молча.
review всегда только для чтения, независимо от того, что разрешает GROK_MCP_PERMISSION_CEILING. Он не принимает аргументов permission, write или yolo, потому что ревью, которое редактирует просматриваемый код, никогда не является желаемым.
Передайте structured: true для машиночитаемых результатов (severity, file, line, summary, rationale) в _meta.findings, проверенных перед тем, как вы их увидите.
Две разные вещи могут пойти не так, и они сообщаются по-разному, а не смешиваются:
Запуск никогда не завершился — он был прерван или закончился без выдачи результатов. Ревью нет, поэтому вызов имеет
isError: trueи_meta.findingsCompleteравноfalse. Тело начинается с объяснения причины, цитируя собственную причину CLI, и называет исправление, соответствующее фактической причине.Запуск завершился, но его вывод не пройдёт валидацию. Вызов всё равно успешен, возвращая сырой текст плюс
_meta.parseError— деградированное ревью лучше, чем неудачное.
Чего вы никогда не получите — это правдоподобно выглядящего результата, который модель выдумала. --json-schema ограничивает каждое сообщение, которое модель выдаёт, поэтому, пока она ещё читает, у неё нет способа сказать «Я работаю», кроме как в форме результата — и если не контролировать, она делает именно это. Схема содержит обязательное поле status, чтобы убрать это повествование из ваших результатов, и ничего никогда не извлекается из частичного ответа с помощью сопоставления шаблонов.
Структурированные ревью больших целей действительно терпят неудачу таким образом с некоторой регулярностью. Сбой громкий по замыслу.
Ревью, которое тянется к оболочке, отклоняется, а не убивается. В безголовом режиме неодобряемый запрос инструмента отменяет весь запуск, в то время как CLI всё равно завершается с кодом 0, поэтому review категорически запрещает оболочку и инструменты редактирования — модели говорят «нет», и она завершает своё ревью, а не умирает на полуслове.
websearch
> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depthnumResults (1–50) и searchDepth (basic или full) формируют подсказку — CLI grok не имеет флагов ни для того, ни для другого, и ни один параметр не притворяется иначе. Они работают: один и тот же вопрос, заданный на basic, выполнил один поиск по двум страницам, а на full — шесть поисков по трём, что в два с половиной раза дороже.
Результат сообщает вам, что было фактически найдено, а не только то, что написала модель:
[1 web search, 9 sources]с _meta, содержащим webSearches, webToolCalls, searchQueries, sources, sourceCount, pagesOpened и searchPerformed. Это важнее, чем кажется. Grok может исследовать через веб-поиск или через X, и когда веб недоступен, он тихо сделает второе — отвечая уверенно, цитируя x.com, успешно завершаясь. Проза не даёт вам возможности понять. Поэтому запуск, который искал в X, а не в вебе, сообщает об этом в первой строке и отдельно сообщает xSearches, а запуск, в котором ничего не вернулось, является ошибкой, а не уверенно выглядящим ответом из собственной памяти модели:
No search ran. The answer below is the model's own prior knowledge, not current sources.searchPerformed означает, что источники вернулись — а не то, что поиск был предпринят. Поиск, который начался и никогда не вернулся, или вернул пустой набор результатов, сообщается как то, чем он был.
Подобно review, websearch всегда только для чтения и не принимает аргументов permission, write или yolo. Он никогда не передаёт --disable-web-search.
Фоновые запуски, status и stop
Длительный запуск агента не обязан занимать ваш клиент. Передайте background: true в grok, review или websearch, и вызов сразу вернёт runId, пока отдельный рабочий процесс выполняет задачу до конца:
> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that runЗапуск привязан к машине, а не к этому серверу: он продолжается, если ваш MCP-клиент отключается, сервер перезапускается или вы закрываете редактор. Записи хранятся в GROK_MCP_STATE_DIR, по одному каталогу на запуск.
status для завершённого запуска возвращает то же, что вернул бы синхронный вызов — тот же текст, те же метаданные, тот же флаг ошибки. Фоновый режим — это транспорт для вызова инструмента, а не вторая реализация. Пока запуск активен, вы получаете его состояние, прошедшее время, оба идентификатора процесса и хвост журнала прогресса; waitMs блокирует до двух минут и пересылает уведомления о прогрессе по мере поступления. Тайм-аут ожидания не является ошибкой.
Два вида нечестности исключены архитектурно. Запуск, чей рабочий процесс больше не существует, помечается как abandoned, а не как всё ещё работающий — машина перезагрузилась или кто-то его убил. А запуск, завершившийся раньше времени, помечается соответствующим образом:
mfk2p1x9-3ac71f0b completed (cut off: cancelled) grok 4m 12s refactor the parserВалидация всё ещё выполняется до того, как вы получите runId: запрос, превышающий GROK_MCP_PERMISSION_CEILING, или противоречивая пара флагов сессии отклоняется как неудачный вызов, а не принимается, а затем терпит неудачу в процессе, за которым никто не наблюдает.
stop завершает запуск досрочно. Он отправляет сигнал SIGTERM всей группе процессов рабочего — рабочему и порождённому им процессу grok, а затем SIGKILL, если этого недостаточно. Остановка уже завершённого запуска не является ошибкой, равно как и остановка того, который завершился за мгновение до вашего вызова.
Остановка, которая не смогла убить дерево процессов, сообщается как сбой, а не как остановленный запуск. Если нечего сигнализировать, или убийство отклонено, или дерево переживает SIGKILL, запуск остаётся в статусе running, а вызов возвращает ошибку с указанием pid. Запись cancelled рядом с живым процессом была бы более аккуратным ответом, но бесполезным.
Запуск, который вы остановили на полпути, обычно уже произвёл что-то стоящее, и частичный результат, и идентификатор сессии сохраняются:
Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.
The run was cancelled mid-flight, but it recorded a session before it ended:
grok -r 01a010e2-478c-73d2-bce9-23552245c64dGrok сообщает идентификатор сессии только когда запуск достигает своего конца, чего остановленный запуск никогда не делает — поэтому этот идентификатор читается из собственного хранилища сессий CLI, а не восстанавливается. _meta.sessionIdSource сообщает вам, какой из них у вас есть. Если два запуска в одном каталоге могут соответствовать, вы получаете идентификаторы кандидатов и никакой команды возобновления: возобновление неправильной сессии продолжает чужую работу.
sessions
Каждый запуск Grok оставляет сессию на диске, и каждый идентификатор сессии, сообщаемый этим сервером, может быть возобновлён позже — из любого каталога, вами в терминале или другим вызовом инструмента.
> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiterСессии читаются из $GROK_HOME/sessions (по умолчанию ~/.grok/sessions), что является собственным хранилищем CLI, поэтому они переживают перезапуски этого сервера, вашего MCP-клиента и вашей машины. Передайте id для одной сессии, query для поиска без учёта регистра по заголовкам, первым запросам и идентификаторам, cwd для ограничения одним проектом и limit для ограничения списка.
Только что завершённый запуск ещё не имеет заголовка — Grok заполняет их позже, если вообще заполняет, поэтому строки возвращаются к первому запросу сессии, а titleSource сообщает вам, на что вы смотрите. Каждая строка содержит resumeCommand, и каждый результат grok и review тоже:
grok -r 01a00c8d-970c-7531-8a12-31dac582c22bПоиск только локальный. grok sessions search также обращается к удалённому индексу; этот инструмент этого не делает, поэтому сессия, существующая только на сервере, не появится.
Разработка
npm install
npm run build # tsc -> dist/
npm run dev # tsx src/index.ts
npm test # node --test via tsx
npm run test:coverage # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run formatdocs/api-reference.md — параметры каждого инструмента, текст результата, ключи
_metaи точные условия, при которых каждый из них установлен.docs/security.md — что регистрация этого сервера разрешает, что на самом деле даёт каждый уровень разрешений и что покидает вашу машину.
docs/engineering.md — как здесь пишется код: архитектура, правила функционального TypeScript, дисциплина ошибок и эффектов, политика тестирования и покрытия, рабочий процесс коммитов.
CLAUDE.md — предыстория проекта и проверенное поведение CLI
grok, от которого зависит этот сервер.ROADMAP.md — вехи, критерии приёмки и идеи, которые были оценены и отвергнуты.
Релиз
Увеличьте version в package.json, переместите раздел Unreleased из CHANGELOG.md под новый заголовок версии, закоммитьте, затем:
git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0.github/workflows/release.yml выполняет полную проверку, отказывается публиковать, если тег и package.json не совпадают, устанавливает упакованный tarball в отдельный каталог и запускает реальный initialize против установленного бинарника, затем публикует тот же самый файл и создаёт релиз на GitHub.
Нет учётных данных для публикации, которыми нужно управлять. Аутентификация — это npm trusted publishing: рабочий процесс обменивается краткосрочным токеном OIDC, а npm самостоятельно генерирует подтверждение происхождения. Доверие зарегистрировано для этого репозитория и имени файла этого рабочего процесса, поэтому переименование release.yml ломает публикацию — и npm не проверяет конфигурацию до тех пор, пока не будет предпринята попытка публикации, где симптомом является ENEEDAUTH, а не что-то, что называет причину.
Лицензия
MIT — см. LICENSE.
Available Tools
8 toolscheckCheck Grok Build readinessARead-onlyIdempotent
Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grokRun Grok BuildA
Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`. | |
| deny | No | Repeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`. | |
| yolo | No | Shorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request. | |
| agent | No | Named subagent to run, passed as `--agent`. | |
| allow | No | Repeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| rules | No | Extra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt. | |
| tools | No | Internal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`. | |
| write | No | Shorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| prompt | Yes | The task for Grok to perform. Passed verbatim as `grok -p`. | |
| resume | No | Resume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| sessionId | No | Create a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| permission | No | Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change. | |
| forkSession | No | UUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`. | |
| continueSession | No | Continue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request. | |
| disallowedTools | No | Internal tool ids to block, passed as `--disallowed-tools`. | |
| disableWebSearch | No | Pass `--disable-web-search`. `false` is not a request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
helpGrok CLI helpARead-onlyIdempotent
Show the grok CLI help text. Runs grok --help and returns its stdout.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviewReview a git diffARead-only
Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Repository to review. Defaults to the current working directory. | |
| base | No | Review the merge-base diff against this ref. Mutually exclusive with commit and uncommitted. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| commit | No | Review this commit. Mutually exclusive with base and uncommitted. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| structured | No | Return machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request. | |
| uncommitted | No | Review the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request. | |
| instructions | No | Extra reviewer guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionsList Grok sessionsARead-onlyIdempotent
List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Exact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match. | |
| cwd | No | Keep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`). | |
| limit | No | Maximum rows to return. Default 20. Ignored when `id` is set. | |
| query | No | Case-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusPoll a background runARead-onlyIdempotent
Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.
| Name | Required | Description | Default |
|---|---|---|---|
| tail | No | Bytes of progress.log to include for a live run. Default 8192. | |
| limit | No | Maximum rows to return in list mode. Default 20. Ignored when `runId` is set. | |
| runId | No | Id of a background run to inspect. Omit to list recent runs. | |
| waitMs | No | Block up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stopStop a background runADestructiveIdempotent
Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | The runId returned by a background `grok`, `review`, or `websearch` call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
websearchSearch the web with Grok BuildARead-only
Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| query | Yes | The question to research. Passed as the body of a web-search-shaped prompt. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| numResults | No | Prompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag. | |
| searchDepth | No | Prompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag. | |
| instructions | No | Extra researcher guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.4- Changed
grok2 fields changed- changed
Input schema / properties / cwd / descriptionPrevious value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`." - changed
Input schema / properties / permission / descriptionPrevious value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
8 tool updates
v0.2.2- First observed
check - First observed
grok - First observed
help - First observed
review - First observed
sessions - First observed
status - First observed
stop - First observed
websearch
TDQS
Scored across 8 tools
Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.
All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.
Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.
The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.
Maintenance
Related MCP Connectors
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Give any MCP-compatible AI assistant a builder for live, hosted web tools and workflows.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables sandboxed file operations via MCP tools, resources, and prompts, with a Claude CLI client and Groq-powered web UI for file CRUD, search, code review, and documentation generation.MIT
- FlicenseNot gradedqualityDmaintenanceExposes Claude Code's file editing, command execution, and test running capabilities as composable MCP tools for any MCP-compatible host, enabling code operations via a stateless bridge.-
- FlicenseAqualityBmaintenanceEnables using the xAI Grok CLI as an MCP sub-agent for code review, asking questions, and continuing conversations within MCP hosts like Claude Code.4-
- AlicenseNot gradedqualityBmaintenanceEnables Codex to use Grok Build CLI as a controlled subagent via MCP tools for independent investigation, review, and isolated implementation tasks.5MIT