Skip to main content
Glama

grok-build-mcp-server

npm MCP Registry CI Node License

Install in VS Code Install in Cursor

Un servidor MCP stdio que expone la CLI de Grok Build (grok) como herramientas que puedes invocar desde Claude Code, Cursor, VS Code o cualquier otro cliente MCP.

Claude Code  ──stdio/MCP──▶  grok-build-mcp-server  ──spawn──▶  grok CLI  ──▶  xAI API

Es un envoltorio ligero de procesos. No reimplementa la lógica del agente ni se comunica directamente con la API de xAI — toda la inteligencia reside en la CLI grok. Lo que este servidor añade es una construcción fiel de argumentos, una supervisión robusta de procesos y una salida limpia con formato MCP.

Estado: 0.2.2. La superficie de herramientas está completa. El servidor ejecuta agentes Grok headless reales en primer plano o en segundo plano de forma desatendida, transmite el progreso mientras se ejecutan, detiene una ejecución cuando se solicita, revisa diferencias de git, investiga preguntas en la web, lista las sesiones que crearon esas ejecuciones, e informa sobre sesiones, uso y coste. Consulta CHANGELOG.md para ver lo publicado y ROADMAP.md para lo que se consideró y rechazó.

Progreso

Una ejecución larga de un agente es visible mientras ocurre, en lugar de una espera silenciosa que termina en un muro de texto. Cuando tu cliente envía un progressToken, el servidor ejecuta Grok con --output-format streaming-json y reenvía una notificación por evento:

#5  list_dir .
#6  read_file README.md
#7  read_file — completed
#8  thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)

El progreso rastrea lo que el agente está haciendo, no en qué fase se encuentra. El razonamiento y el texto de respuesta se fusionan para que un flujo de tokens no inunde tu cliente, mientras que las llamadas a herramientas se informan a medida que ocurren. Los clientes que admiten resetTimeoutOnProgress no agotarán el tiempo de espera a mitad de la ejecución.

Un cliente que no envía progressToken obtiene la ruta más económica sin streaming y no paga nada por ello.

Related MCP server: Claude Code MCP Bridge

Requisitos

  • CLI de Grok Build 1.0.0 o superior, autenticada (grok models debe funcionar)

  • Node.js 22 o superior

Si grok no está en tu PATH, establece GROK_BINARY con su ruta completa al registrar el servidor.

Instalación

Claude Code

claude mcp add grok-build -- npx -y grok-build-mcp-server

Luego, en Claude Code:

> use the grok-build check tool

check informa el binario resuelto, la versión de la CLI, si estás autenticado y el límite de permisos activo. Si todo está correcto, el resto funcionará.

Cualquier otro cliente MCP

El servidor habla MCP a través de stdio y no acepta argumentos propios:

{
  "mcpServers": {
    "grok-build": {
      "command": "npx",
      "args": ["-y", "grok-build-mcp-server"]
    }
  }
}

VS Code y Cursor aceptan las insignias de instalación en la parte superior de esta página, que llevan exactamente esa configuración.

Los clientes que se instalan desde el Registro MCP conocen este servidor como io.github.Nuruvala/grok-build-mcp-server. La entrada del registro se publica desde la misma etiqueta que el lanzamiento de npm y apunta al mismo paquete.

Si npx no encuentra el servidor

npx resuelve un nombre de paquete simple contra el proyecto local primero. Si el directorio de trabajo de tu cliente MCP es una copia de este repositorio — o de cualquier otro cuyo package.json se llame grok-build-mcp-server — npx -y grok-build-mcp-server ejecuta el punto de entrada local, no encuentra ninguno y falla con command not found. Instálalo en su propio lugar y registra esa ruta:

npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-server

Permisos

Las ejecuciones de Grok lanzadas a través de este servidor son de solo lectura por defecto: --permission-mode plan con --sandbox read-only. Nada puede modificar tus archivos hasta que tú lo indiques.

El permiso es un límite máximo, establecido una vez al registrar el servidor, en lugar de una solicitud en cada llamada. Tres niveles:

Nivel

--permission-mode

--sandbox

Qué permite

read-only (por defecto)

plan

read-only

Lectura y razonamiento. Sin ediciones

write

acceptEdits

workspace

Ediciones dentro del directorio de trabajo

full

bypassPermissions

off

Aprobación total sin supervisión

Para permitir que Grok haga ediciones:

claude mcp add grok-build \
  -e GROK_MCP_PERMISSION_CEILING=write \
  -e GROK_MCP_DEFAULT_PERMISSION=write \
  -- npx -y grok-build-mcp-server

Usa full solo si ya ejecutas tu cliente MCP con aprobación total y deseas que la ejecución delegada de Grok esté igualmente desatendida. Concede al proceso grok generado la misma autoridad que tienes tú.

Una llamada que solicite más del límite máximo es rechazada, no degradada silenciosamente — una ejecución limitada informaría éxito sin cambiar nada, lo cual es peor que un error claro.

Variables de entorno

Variable

Por defecto

Propósito

GROK_BINARY

grok

Ruta al ejecutable grok

GROK_MCP_PERMISSION_CEILING

read-only

Nivel más alto que cualquier llamada puede solicitar

GROK_MCP_DEFAULT_PERMISSION

read-only

Nivel usado cuando una llamada no solicita ninguno

GROK_MCP_DEFAULT_MODEL

grok-4.6

Modelo cuando una llamada lo omite. none delega en la CLI

GROK_MCP_DEFAULT_EFFORT

high

Esfuerzo de razonamiento cuando una llamada lo omite. none delega en la CLI

GROK_MCP_TIMEOUT_MS

1800000

Tiempo máximo de reloj para una sola ejecución

GROK_MCP_STATE_DIR

$XDG_STATE_HOME/grok-mcp

Registros de trabajos en segundo plano

GROK_MCP_MAX_CONCURRENT_RUNS

4

Ejecuciones en segundo plano activas a la vez. off para sin límite

GROK_MCP_LOG_LEVEL

info

debug, info, warn, error. Los registros van a stderr

STRUCTURED_CONTENT_ENABLED

desactivado

También emite structuredContent junto con _meta

Las variables propias de Grok (XAI_API_KEY, GROK_HOME, GROK_DISABLE_AUTOUPDATER) se transmiten al proceso hijo sin cambios.

Herramientas

Herramienta

Solo lectura

Propósito

grok

según límite

Ejecutar un agente Grok headless. Prompt, reanudar/continuar/bifurcar sesión, modelo, esfuerzo, permitir/denegar herramientas

review

siempre

Revisar un diff de git: árbol de trabajo, diff de base de fusión contra una referencia, o un solo commit

websearch

siempre

Investigar una pregunta en la web e informar qué búsquedas y fuentes utilizó realmente

status

siempre

Consultar una ejecución en segundo plano o listar las recientes

stop

no

Terminar el árbol de procesos de una ejecución en segundo plano

sessions

siempre

Listar, buscar y consultar las sesiones de Grok en esta máquina

check

sí

Versión del servidor, binario resuelto, grok version, autenticación, límite de permisos, valores predeterminados de ejecución

help

sí

Paso directo de grok --help

review

El diff se recopila en el proceso y se incrusta en el prompt, para que el modelo no gaste turnos redescubriendo lo que se supone que debe revisar.

> review my working tree with grok-build
> review the diff against origin/main

Los objetivos son uncommitted, base: "<ref>" (un diff de base de fusión, por lo que los commits que llegaron a la base después de que creaste tu rama no se te atribuyen), o commit: "<sha>". Si no se proporciona ninguno, se detecta automáticamente: el diff ascendente cuando tu rama está adelantada, de lo contrario el árbol de trabajo — y dice cuál eligió en lugar de adivinarlo en silencio.

review es siempre de solo lectura, independientemente de lo que permita GROK_MCP_PERMISSION_CEILING. No acepta ningún argumento permission, write o yolo, porque una revisión que edita el código bajo revisión nunca es lo que se deseaba.

Pasa structured: true para obtener hallazgos legibles por máquina (severity, file, line, summary, rationale) en _meta.findings, validados antes de que los veas.

Dos cosas diferentes pueden salir mal, y se informan de manera diferente en lugar de difuminarse:

  • La ejecución nunca terminó — se interrumpió o finalizó sin producir sus hallazgos. No hay revisión, por lo que la llamada es isError: true y _meta.findingsComplete es false. El cuerpo comienza con el motivo, citando la razón de la propia CLI, y nombra la solución que se ajusta a la causa real.

  • La ejecución terminó pero su salida no se validará. La llamada aún tiene éxito, devolviendo el texto sin procesar más un _meta.parseError — una revisión degradada es mejor que una fallida.

Lo que nunca obtendrás es un hallazgo de apariencia plausible que el modelo inventó. --json-schema restringe cada mensaje que el modelo emite, por lo que mientras aún está leyendo no tiene forma de decir "estoy trabajando" excepto en la forma de un hallazgo — y sin control, hace exactamente eso. El esquema lleva un campo status requerido para mantener esa narración fuera de tus resultados, y nunca se recupera nada de una respuesta parcial mediante coincidencia de patrones.

Las revisiones estructuradas de objetivos grandes fallan de esta manera con cierta regularidad. El fallo es ruidoso por diseño.

Una revisión que intenta acceder a un shell es rechazada, no terminada. En modo headless, una solicitud de herramienta no aprobable cancela toda la ejecución mientras la CLI aún sale con 0, por lo que review niega las herramientas de shell y edición directamente — se le dice no al modelo y termina su revisión en lugar de morir a mitad de la frase.

websearch

> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depth

numResults (1–50) y searchDepth (basic o full) dan forma al prompt — la CLI grok no tiene banderas para ninguno, y ningún parámetro finge lo contrario. Funcionan: la misma pregunta hecha en basic hizo una búsqueda en dos páginas, y en full hizo seis búsquedas en tres, por dos veces y media el coste.

El resultado te dice lo que realmente se consultó, no solo lo que el modelo escribió:

[1 web search, 9 sources]

con _meta que lleva webSearches, webToolCalls, searchQueries, sources, sourceCount, pagesOpened y searchPerformed. Eso importa más de lo que parece. Grok puede investigar a través de la búsqueda web o a través de X, y cuando la web no está disponible, hará silenciosamente lo segundo — respondiendo con confianza, citando x.com, saliendo con éxito. La prosa no te da forma de saberlo. Por lo tanto, una ejecución que buscó en X y no en la web lo dice en su primera línea e informa xSearches por separado, y una ejecución donde no volvió nada en absoluto es un error en lugar de una respuesta de apariencia confiada de la propia memoria del modelo:

No search ran. The answer below is the model's own prior knowledge, not current sources.

searchPerformed significa que volvieron fuentes — no que se intentó una búsqueda. Una búsqueda que comenzó y nunca regresó, o que devolvió un conjunto de resultados vacío, se informa como lo que fue.

Al igual que review, websearch es siempre de solo lectura y no acepta argumentos permission, write ni yolo. Nunca pasa --disable-web-search.

Ejecuciones en segundo plano, status y stop

Una ejecución larga del agente no tiene por qué ocupar tu cliente. Pasa background: true a grok, review o websearch y la llamada devuelve un runId de inmediato, mientras un proceso trabajador independiente ejecuta el trabajo hasta completarlo:

> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that run

La ejecución pertenece a la máquina, no a este servidor: continúa si tu cliente MCP se desconecta, si el servidor se reinicia o si cierras tu editor. Los registros viven en GROK_MCP_STATE_DIR, un directorio por ejecución.

status sobre una ejecución finalizada devuelve lo que habría devuelto la llamada síncrona — mismo texto, mismos metadatos, mismo indicador de error. El segundo plano es un transporte para una llamada a herramienta, no una segunda implementación de la misma. Mientras una ejecución está activa obtienes su estado, tiempo transcurrido, ambos identificadores de proceso y la cola de su registro de progreso; waitMs bloquea hasta dos minutos y reenvía las notificaciones de progreso a medida que llegan. Una espera agotada no es un error.

Dos tipos de deshonestidad quedan descartados por construcción. Una ejecución cuyo proceso trabajador ya no existe se reporta como abandoned en lugar de como aún en ejecución — la máquina se reinició o algo la mató. Y una ejecución que terminó temprano se etiqueta como tal:

mfk2p1x9-3ac71f0b  completed (cut off: cancelled)  grok  4m 12s  refactor the parser

La validación sigue ocurriendo antes de que obtengas un runId: una solicitud por encima de GROK_MCP_PERMISSION_CEILING, o un par contradictorio de indicadores de sesión, se rechaza como una llamada fallida en lugar de aceptarse y luego fallar en un proceso que nadie está observando.

stop termina una ejecución anticipadamente. Envía una señal al grupo de procesos completo del trabajador — el trabajador y el proceso grok que generó — con SIGTERM, y luego SIGKILL si eso no es suficiente. Detener una ejecución ya finalizada no es un error, ni tampoco lo es detener una que terminó justo antes de que llegara tu llamada.

Una parada que no pudo matar el árbol de procesos se reporta como un fallo, no como una ejecución detenida. Si no hay nada a lo que enviar señal, o el asesinato es rechazado, o el árbol sobrevive a SIGKILL, la ejecución sigue marcada como running y la llamada devuelve un error que nombra el pid. Un registro cancelled junto a un proceso vivo sería la respuesta más ordenada, pero inútil.

Una ejecución que detienes a medio vuelo normalmente ya ha producido algo que vale la pena conservar, y tanto el resultado parcial como el identificador de sesión se conservan:

Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.

The run was cancelled mid-flight, but it recorded a session before it ended:
  grok -r 01a010e2-478c-73d2-bce9-23552245c64d

Grok solo reporta un identificador de sesión cuando una ejecución llega a su fin, cosa que una detenida nunca hace — así que ese id se lee del propio almacén de sesiones de la CLI en lugar de reconstruirse. _meta.sessionIdSource te indica cuál tienes. Si dos ejecuciones en el mismo directorio pudieran coincidir, obtienes los ids candidatos y ningún comando de reanudación: reanudar la sesión incorrecta continúa el trabajo de otra persona.

sessions

Cada ejecución de Grok deja una sesión en disco, y cada identificador de sesión que reporta este servidor puede reanudarse después — desde cualquier directorio, por ti en una terminal o mediante otra llamada a herramienta.

> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiter

Las sesiones se leen de $GROK_HOME/sessions (por defecto ~/.grok/sessions), que es el propio almacén de la CLI, por lo que sobreviven a reinicios de este servidor, de tu cliente MCP y de tu máquina. Pasa id para una sesión, query para una búsqueda que no distingue mayúsculas sobre títulos, primeros avisos e ids, cwd para limitar a un proyecto, y limit para acotar la lista.

Una ejecución que acaba de terminar aún no tiene título — Grok los completa después, si acaso — así que las filas recurren al primer aviso de la sesión, y titleSource te indica cuál estás viendo. Cada fila lleva resumeCommand, y también lo lleva cada resultado de grok y review:

grok -r 01a00c8d-970c-7531-8a12-31dac582c22b

La búsqueda es solo local. grok sessions search también consulta un índice remoto; esta herramienta no lo hace, por lo que una sesión que solo exista del lado del servidor no aparecerá.

Desarrollo

npm install
npm run build          # tsc -> dist/
npm run dev            # tsx src/index.ts
npm test               # node --test via tsx
npm run test:coverage  # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run format
  • docs/api-reference.md — parámetros de cada herramienta, texto de resultado, claves _meta y las condiciones exactas bajo las que se establece cada una.

  • docs/security.md — qué autoriza registrar este servidor, qué otorga realmente cada nivel de permiso y qué sale de tu máquina.

  • docs/engineering.md — cómo se escribe el código aquí: arquitectura, reglas de TypeScript funcional, disciplina de errores y efectos, política de pruebas y cobertura, flujo de trabajo de confirmaciones.

  • CLAUDE.md — antecedentes del proyecto y el comportamiento verificado de la CLI grok del que depende este servidor.

  • ROADMAP.md — hitos, criterios de aceptación e ideas que se midieron y rechazaron.

Publicación

Incrementa version en package.json, mueve la sección Unreleased de CHANGELOG.md bajo el nuevo encabezado de versión, confirma, luego:

git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0

.github/workflows/release.yml ejecuta la compuerta completa, se niega a publicar si la etiqueta y package.json no coinciden, instala el tarball empaquetado en un directorio temporal y ejecuta un initialize real contra el binario instalado, luego publica ese mismo archivo y crea un lanzamiento en GitHub.

No hay credenciales de publicación que gestionar. La autenticación es npm trusted publishing: el flujo de trabajo intercambia un token OIDC de corta duración, y npm genera la atestación de procedencia por su cuenta. La confianza está registrada contra este repositorio y el nombre del archivo de este flujo de trabajo, por lo que renombrar release.yml rompe la publicación — y npm no verifica la configuración hasta que se intenta publicar, donde el síntoma es ENEEDAUTH en lugar de algo que nombre la causa.

Licencia

MIT — consulta LICENSE.

Available Tools

8 tools
checkCheck Grok Build readinessA
Read-onlyIdempotent

Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokRun Grok BuildA

Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`.
denyNoRepeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`.
yoloNoShorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request.
agentNoNamed subagent to run, passed as `--agent`.
allowNoRepeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
rulesNoExtra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt.
toolsNoInternal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`.
writeNoShorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
promptYesThe task for Grok to perform. Passed verbatim as `grok -p`.
resumeNoResume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
sessionIdNoCreate a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
permissionNoPermission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change.
forkSessionNoUUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`.
continueSessionNoContinue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request.
disallowedToolsNoInternal tool ids to block, passed as `--disallowed-tools`.
disableWebSearchNoPass `--disable-web-search`. `false` is not a request.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

helpGrok CLI helpA
Read-onlyIdempotent

Show the grok CLI help text. Runs grok --help and returns its stdout.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reviewReview a git diffA
Read-only

Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Repository to review. Defaults to the current working directory.
baseNoReview the merge-base diff against this ref. Mutually exclusive with commit and uncommitted.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
commitNoReview this commit. Mutually exclusive with base and uncommitted.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
structuredNoReturn machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request.
uncommittedNoReview the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request.
instructionsNoExtra reviewer guidance, appended verbatim to the prompt.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sessionsList Grok sessionsA
Read-onlyIdempotent

List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoExact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match.
cwdNoKeep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`).
limitNoMaximum rows to return. Default 20. Ignored when `id` is set.
queryNoCase-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusPoll a background runA
Read-onlyIdempotent

Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
tailNoBytes of progress.log to include for a live run. Default 8192.
limitNoMaximum rows to return in list mode. Default 20. Ignored when `runId` is set.
runIdNoId of a background run to inspect. Omit to list recent runs.
waitMsNoBlock up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stopStop a background runA
DestructiveIdempotent

Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesThe runId returned by a background `grok`, `review`, or `websearch` call.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

websearchSearch the web with Grok BuildA
Read-only

Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
queryYesThe question to research. Passed as the body of a web-search-shaped prompt.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
numResultsNoPrompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag.
searchDepthNoPrompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag.
instructionsNoExtra researcher guidance, appended verbatim to the prompt.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.2.4
    • Changedgrok2 fields changed
      • changedInput schema / properties / cwd / description
        Previous value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`."
      • changedInput schema / properties / permission / description
        Previous value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
  2. 8 tool updatesv0.2.2
    • First observedcheck
    • First observedgrok
    • First observedhelp
    • First observedreview
    • First observedsessions
    • First observedstatus
    • First observedstop
    • First observedwebsearch

TDQS

A4.5/5.0

Scored across 8 tools

Disambiguation5/5

Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.

Naming Consistency4/5

All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.

Tool Count5/5

Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.

Completeness5/5

The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers