grok-build-mcp-server
grok-build-mcp-server
Un servidor MCP stdio que expone la CLI de Grok Build
(grok) como herramientas que puedes invocar desde Claude Code, Cursor, VS Code o cualquier otro cliente MCP.
Claude Code ──stdio/MCP──▶ grok-build-mcp-server ──spawn──▶ grok CLI ──▶ xAI APIEs un envoltorio ligero de procesos. No reimplementa la lógica del agente ni se comunica directamente con la API de xAI
— toda la inteligencia reside en la CLI grok. Lo que este servidor añade es una construcción fiel de argumentos,
una supervisión robusta de procesos y una salida limpia con formato MCP.
Estado: 0.2.2. La superficie de herramientas está completa. El servidor ejecuta agentes Grok headless reales en primer plano o en segundo plano de forma desatendida, transmite el progreso mientras se ejecutan, detiene una ejecución cuando se solicita, revisa diferencias de git, investiga preguntas en la web, lista las sesiones que crearon esas ejecuciones, e informa sobre sesiones, uso y coste. Consulta CHANGELOG.md para ver lo publicado y ROADMAP.md para lo que se consideró y rechazó.
Progreso
Una ejecución larga de un agente es visible mientras ocurre, en lugar de una espera silenciosa que termina en un muro de texto.
Cuando tu cliente envía un progressToken, el servidor ejecuta Grok con --output-format streaming-json
y reenvía una notificación por evento:
#5 list_dir .
#6 read_file README.md
#7 read_file — completed
#8 thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)El progreso rastrea lo que el agente está haciendo, no en qué fase se encuentra. El razonamiento y el texto de respuesta se
fusionan para que un flujo de tokens no inunde tu cliente, mientras que las llamadas a herramientas se informan a medida que
ocurren. Los clientes que admiten resetTimeoutOnProgress no agotarán el tiempo de espera a mitad de la ejecución.
Un cliente que no envía progressToken obtiene la ruta más económica sin streaming y no paga nada por
ello.
Related MCP server: Claude Code MCP Bridge
Requisitos
CLI de Grok Build 1.0.0 o superior, autenticada (
grok modelsdebe funcionar)Node.js 22 o superior
Si grok no está en tu PATH, establece GROK_BINARY con su ruta completa al registrar el servidor.
Instalación
Claude Code
claude mcp add grok-build -- npx -y grok-build-mcp-serverLuego, en Claude Code:
> use the grok-build check toolcheck informa el binario resuelto, la versión de la CLI, si estás autenticado y el límite de permisos activo.
Si todo está correcto, el resto funcionará.
Cualquier otro cliente MCP
El servidor habla MCP a través de stdio y no acepta argumentos propios:
{
"mcpServers": {
"grok-build": {
"command": "npx",
"args": ["-y", "grok-build-mcp-server"]
}
}
}VS Code y Cursor aceptan las insignias de instalación en la parte superior de esta página, que llevan exactamente esa configuración.
Los clientes que se instalan desde el Registro MCP conocen este
servidor como io.github.Nuruvala/grok-build-mcp-server. La entrada del registro se publica desde la misma
etiqueta que el lanzamiento de npm y apunta al mismo paquete.
Si npx no encuentra el servidor
npx resuelve un nombre de paquete simple contra el proyecto local primero. Si el directorio de trabajo de tu
cliente MCP es una copia de este repositorio — o de cualquier otro cuyo package.json se llame
grok-build-mcp-server — npx -y grok-build-mcp-server ejecuta el punto de entrada local, no encuentra
ninguno y falla con command not found. Instálalo en su propio lugar y registra esa ruta:
npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-serverPermisos
Las ejecuciones de Grok lanzadas a través de este servidor son de solo lectura por defecto: --permission-mode plan con
--sandbox read-only. Nada puede modificar tus archivos hasta que tú lo indiques.
El permiso es un límite máximo, establecido una vez al registrar el servidor, en lugar de una solicitud en cada llamada. Tres niveles:
Nivel |
|
| Qué permite |
|
|
| Lectura y razonamiento. Sin ediciones |
|
|
| Ediciones dentro del directorio de trabajo |
|
|
| Aprobación total sin supervisión |
Para permitir que Grok haga ediciones:
claude mcp add grok-build \
-e GROK_MCP_PERMISSION_CEILING=write \
-e GROK_MCP_DEFAULT_PERMISSION=write \
-- npx -y grok-build-mcp-serverUsa full solo si ya ejecutas tu cliente MCP con aprobación total y deseas que la ejecución delegada de Grok
esté igualmente desatendida. Concede al proceso grok generado la misma autoridad que tienes tú.
Una llamada que solicite más del límite máximo es rechazada, no degradada silenciosamente — una ejecución limitada informaría éxito sin cambiar nada, lo cual es peor que un error claro.
Variables de entorno
Variable | Por defecto | Propósito |
|
| Ruta al ejecutable |
|
| Nivel más alto que cualquier llamada puede solicitar |
|
| Nivel usado cuando una llamada no solicita ninguno |
|
| Modelo cuando una llamada lo omite. |
|
| Esfuerzo de razonamiento cuando una llamada lo omite. |
|
| Tiempo máximo de reloj para una sola ejecución |
|
| Registros de trabajos en segundo plano |
|
| Ejecuciones en segundo plano activas a la vez. |
|
|
|
| desactivado | También emite |
Las variables propias de Grok (XAI_API_KEY, GROK_HOME, GROK_DISABLE_AUTOUPDATER) se transmiten al
proceso hijo sin cambios.
Herramientas
Herramienta | Solo lectura | Propósito |
| según límite | Ejecutar un agente Grok headless. Prompt, reanudar/continuar/bifurcar sesión, modelo, esfuerzo, permitir/denegar herramientas |
| siempre | Revisar un diff de git: árbol de trabajo, diff de base de fusión contra una referencia, o un solo commit |
| siempre | Investigar una pregunta en la web e informar qué búsquedas y fuentes utilizó realmente |
| siempre | Consultar una ejecución en segundo plano o listar las recientes |
| no | Terminar el árbol de procesos de una ejecución en segundo plano |
| siempre | Listar, buscar y consultar las sesiones de Grok en esta máquina |
| sí | Versión del servidor, binario resuelto, |
| sí | Paso directo de |
review
El diff se recopila en el proceso y se incrusta en el prompt, para que el modelo no gaste turnos redescubriendo lo que se supone que debe revisar.
> review my working tree with grok-build
> review the diff against origin/mainLos objetivos son uncommitted, base: "<ref>" (un diff de base de fusión, por lo que los commits que llegaron a la base
después de que creaste tu rama no se te atribuyen), o commit: "<sha>". Si no se proporciona ninguno, se detecta
automáticamente: el diff ascendente cuando tu rama está adelantada, de lo contrario el árbol de trabajo — y dice
cuál eligió en lugar de adivinarlo en silencio.
review es siempre de solo lectura, independientemente de lo que permita GROK_MCP_PERMISSION_CEILING. No acepta
ningún argumento permission, write o yolo, porque una revisión que edita el código bajo revisión nunca es
lo que se deseaba.
Pasa structured: true para obtener hallazgos legibles por máquina (severity, file, line, summary,
rationale) en _meta.findings, validados antes de que los veas.
Dos cosas diferentes pueden salir mal, y se informan de manera diferente en lugar de difuminarse:
La ejecución nunca terminó — se interrumpió o finalizó sin producir sus hallazgos. No hay revisión, por lo que la llamada es
isError: truey_meta.findingsCompleteesfalse. El cuerpo comienza con el motivo, citando la razón de la propia CLI, y nombra la solución que se ajusta a la causa real.La ejecución terminó pero su salida no se validará. La llamada aún tiene éxito, devolviendo el texto sin procesar más un
_meta.parseError— una revisión degradada es mejor que una fallida.
Lo que nunca obtendrás es un hallazgo de apariencia plausible que el modelo inventó. --json-schema
restringe cada mensaje que el modelo emite, por lo que mientras aún está leyendo no tiene forma de decir "estoy
trabajando" excepto en la forma de un hallazgo — y sin control, hace exactamente eso. El esquema
lleva un campo status requerido para mantener esa narración fuera de tus resultados, y nunca se
recupera nada de una respuesta parcial mediante coincidencia de patrones.
Las revisiones estructuradas de objetivos grandes fallan de esta manera con cierta regularidad. El fallo es ruidoso por diseño.
Una revisión que intenta acceder a un shell es rechazada, no terminada. En modo headless, una solicitud de herramienta
no aprobable cancela toda la ejecución mientras la CLI aún sale con 0, por lo que review niega las herramientas de shell
y edición directamente — se le dice no al modelo y termina su revisión en lugar de morir a mitad de la frase.
websearch
> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depthnumResults (1–50) y searchDepth (basic o full) dan forma al prompt — la CLI grok no tiene
banderas para ninguno, y ningún parámetro finge lo contrario. Funcionan: la misma pregunta hecha en basic
hizo una búsqueda en dos páginas, y en full hizo seis búsquedas en tres, por dos
veces y media el coste.
El resultado te dice lo que realmente se consultó, no solo lo que el modelo escribió:
[1 web search, 9 sources]con _meta que lleva webSearches, webToolCalls, searchQueries, sources, sourceCount,
pagesOpened y searchPerformed. Eso importa más de lo que parece. Grok puede investigar a través de la búsqueda
web o a través de X, y cuando la web no está disponible, hará silenciosamente lo segundo — respondiendo
con confianza, citando x.com, saliendo con éxito. La prosa no te da forma de saberlo. Por lo tanto, una ejecución que
buscó en X y no en la web lo dice en su primera línea e informa xSearches por separado, y una ejecución
donde no volvió nada en absoluto es un error en lugar de una respuesta de apariencia confiada de la propia memoria del
modelo:
No search ran. The answer below is the model's own prior knowledge, not current sources.searchPerformed significa que volvieron fuentes — no que se intentó una búsqueda. Una búsqueda que comenzó
y nunca regresó, o que devolvió un conjunto de resultados vacío, se informa como lo que fue.
Al igual que review, websearch es siempre de solo lectura y no acepta argumentos permission, write ni yolo. Nunca pasa --disable-web-search.
Ejecuciones en segundo plano, status y stop
Una ejecución larga del agente no tiene por qué ocupar tu cliente. Pasa background: true a grok, review o websearch y la llamada devuelve un runId de inmediato, mientras un proceso trabajador independiente ejecuta el trabajo hasta completarlo:
> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that runLa ejecución pertenece a la máquina, no a este servidor: continúa si tu cliente MCP se desconecta, si el servidor se reinicia o si cierras tu editor. Los registros viven en GROK_MCP_STATE_DIR, un directorio por ejecución.
status sobre una ejecución finalizada devuelve lo que habría devuelto la llamada síncrona — mismo texto, mismos metadatos, mismo indicador de error. El segundo plano es un transporte para una llamada a herramienta, no una segunda implementación de la misma. Mientras una ejecución está activa obtienes su estado, tiempo transcurrido, ambos identificadores de proceso y la cola de su registro de progreso; waitMs bloquea hasta dos minutos y reenvía las notificaciones de progreso a medida que llegan. Una espera agotada no es un error.
Dos tipos de deshonestidad quedan descartados por construcción. Una ejecución cuyo proceso trabajador ya no existe se reporta como abandoned en lugar de como aún en ejecución — la máquina se reinició o algo la mató. Y una ejecución que terminó temprano se etiqueta como tal:
mfk2p1x9-3ac71f0b completed (cut off: cancelled) grok 4m 12s refactor the parserLa validación sigue ocurriendo antes de que obtengas un runId: una solicitud por encima de GROK_MCP_PERMISSION_CEILING, o un par contradictorio de indicadores de sesión, se rechaza como una llamada fallida en lugar de aceptarse y luego fallar en un proceso que nadie está observando.
stop termina una ejecución anticipadamente. Envía una señal al grupo de procesos completo del trabajador — el trabajador y el proceso grok que generó — con SIGTERM, y luego SIGKILL si eso no es suficiente. Detener una ejecución ya finalizada no es un error, ni tampoco lo es detener una que terminó justo antes de que llegara tu llamada.
Una parada que no pudo matar el árbol de procesos se reporta como un fallo, no como una ejecución detenida. Si no hay nada a lo que enviar señal, o el asesinato es rechazado, o el árbol sobrevive a SIGKILL, la ejecución sigue marcada como running y la llamada devuelve un error que nombra el pid. Un registro cancelled junto a un proceso vivo sería la respuesta más ordenada, pero inútil.
Una ejecución que detienes a medio vuelo normalmente ya ha producido algo que vale la pena conservar, y tanto el resultado parcial como el identificador de sesión se conservan:
Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.
The run was cancelled mid-flight, but it recorded a session before it ended:
grok -r 01a010e2-478c-73d2-bce9-23552245c64dGrok solo reporta un identificador de sesión cuando una ejecución llega a su fin, cosa que una detenida nunca hace — así que ese id se lee del propio almacén de sesiones de la CLI en lugar de reconstruirse. _meta.sessionIdSource te indica cuál tienes. Si dos ejecuciones en el mismo directorio pudieran coincidir, obtienes los ids candidatos y ningún comando de reanudación: reanudar la sesión incorrecta continúa el trabajo de otra persona.
sessions
Cada ejecución de Grok deja una sesión en disco, y cada identificador de sesión que reporta este servidor puede reanudarse después — desde cualquier directorio, por ti en una terminal o mediante otra llamada a herramienta.
> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiterLas sesiones se leen de $GROK_HOME/sessions (por defecto ~/.grok/sessions), que es el propio almacén de la CLI, por lo que sobreviven a reinicios de este servidor, de tu cliente MCP y de tu máquina. Pasa id para una sesión, query para una búsqueda que no distingue mayúsculas sobre títulos, primeros avisos e ids, cwd para limitar a un proyecto, y limit para acotar la lista.
Una ejecución que acaba de terminar aún no tiene título — Grok los completa después, si acaso — así que las filas recurren al primer aviso de la sesión, y titleSource te indica cuál estás viendo. Cada fila lleva resumeCommand, y también lo lleva cada resultado de grok y review:
grok -r 01a00c8d-970c-7531-8a12-31dac582c22bLa búsqueda es solo local. grok sessions search también consulta un índice remoto; esta herramienta no lo hace, por lo que una sesión que solo exista del lado del servidor no aparecerá.
Desarrollo
npm install
npm run build # tsc -> dist/
npm run dev # tsx src/index.ts
npm test # node --test via tsx
npm run test:coverage # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run formatdocs/api-reference.md — parámetros de cada herramienta, texto de resultado, claves
_metay las condiciones exactas bajo las que se establece cada una.docs/security.md — qué autoriza registrar este servidor, qué otorga realmente cada nivel de permiso y qué sale de tu máquina.
docs/engineering.md — cómo se escribe el código aquí: arquitectura, reglas de TypeScript funcional, disciplina de errores y efectos, política de pruebas y cobertura, flujo de trabajo de confirmaciones.
CLAUDE.md — antecedentes del proyecto y el comportamiento verificado de la CLI
grokdel que depende este servidor.ROADMAP.md — hitos, criterios de aceptación e ideas que se midieron y rechazaron.
Publicación
Incrementa version en package.json, mueve la sección Unreleased de CHANGELOG.md bajo el nuevo encabezado de versión, confirma, luego:
git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0.github/workflows/release.yml ejecuta la compuerta completa, se niega a publicar si la etiqueta y package.json no coinciden, instala el tarball empaquetado en un directorio temporal y ejecuta un initialize real contra el binario instalado, luego publica ese mismo archivo y crea un lanzamiento en GitHub.
No hay credenciales de publicación que gestionar. La autenticación es npm trusted publishing: el flujo de trabajo intercambia un token OIDC de corta duración, y npm genera la atestación de procedencia por su cuenta. La confianza está registrada contra este repositorio y el nombre del archivo de este flujo de trabajo, por lo que renombrar release.yml rompe la publicación — y npm no verifica la configuración hasta que se intenta publicar, donde el síntoma es ENEEDAUTH en lugar de algo que nombre la causa.
Licencia
MIT — consulta LICENSE.
Available Tools
8 toolscheckCheck Grok Build readinessARead-onlyIdempotent
Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grokRun Grok BuildA
Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`. | |
| deny | No | Repeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`. | |
| yolo | No | Shorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request. | |
| agent | No | Named subagent to run, passed as `--agent`. | |
| allow | No | Repeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| rules | No | Extra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt. | |
| tools | No | Internal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`. | |
| write | No | Shorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| prompt | Yes | The task for Grok to perform. Passed verbatim as `grok -p`. | |
| resume | No | Resume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| sessionId | No | Create a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| permission | No | Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change. | |
| forkSession | No | UUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`. | |
| continueSession | No | Continue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request. | |
| disallowedTools | No | Internal tool ids to block, passed as `--disallowed-tools`. | |
| disableWebSearch | No | Pass `--disable-web-search`. `false` is not a request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
helpGrok CLI helpARead-onlyIdempotent
Show the grok CLI help text. Runs grok --help and returns its stdout.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviewReview a git diffARead-only
Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Repository to review. Defaults to the current working directory. | |
| base | No | Review the merge-base diff against this ref. Mutually exclusive with commit and uncommitted. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| commit | No | Review this commit. Mutually exclusive with base and uncommitted. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| structured | No | Return machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request. | |
| uncommitted | No | Review the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request. | |
| instructions | No | Extra reviewer guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionsList Grok sessionsARead-onlyIdempotent
List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Exact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match. | |
| cwd | No | Keep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`). | |
| limit | No | Maximum rows to return. Default 20. Ignored when `id` is set. | |
| query | No | Case-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusPoll a background runARead-onlyIdempotent
Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.
| Name | Required | Description | Default |
|---|---|---|---|
| tail | No | Bytes of progress.log to include for a live run. Default 8192. | |
| limit | No | Maximum rows to return in list mode. Default 20. Ignored when `runId` is set. | |
| runId | No | Id of a background run to inspect. Omit to list recent runs. | |
| waitMs | No | Block up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stopStop a background runADestructiveIdempotent
Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | The runId returned by a background `grok`, `review`, or `websearch` call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
websearchSearch the web with Grok BuildARead-only
Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| query | Yes | The question to research. Passed as the body of a web-search-shaped prompt. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| numResults | No | Prompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag. | |
| searchDepth | No | Prompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag. | |
| instructions | No | Extra researcher guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.4- Changed
grok2 fields changed- changed
Input schema / properties / cwd / descriptionPrevious value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`." - changed
Input schema / properties / permission / descriptionPrevious value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
8 tool updates
v0.2.2- First observed
check - First observed
grok - First observed
help - First observed
review - First observed
sessions - First observed
status - First observed
stop - First observed
websearch
TDQS
Scored across 8 tools
Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.
All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.
Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.
The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.
Maintenance
Related MCP Connectors
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Give any MCP-compatible AI assistant a builder for live, hosted web tools and workflows.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables sandboxed file operations via MCP tools, resources, and prompts, with a Claude CLI client and Groq-powered web UI for file CRUD, search, code review, and documentation generation.MIT
- FlicenseNot gradedqualityDmaintenanceExposes Claude Code's file editing, command execution, and test running capabilities as composable MCP tools for any MCP-compatible host, enabling code operations via a stateless bridge.-
- FlicenseAqualityBmaintenanceEnables using the xAI Grok CLI as an MCP sub-agent for code review, asking questions, and continuing conversations within MCP hosts like Claude Code.4-
- AlicenseNot gradedqualityBmaintenanceEnables Codex to use Grok Build CLI as a controlled subagent via MCP tools for independent investigation, review, and isolated implementation tasks.5MIT