Skip to main content
Glama

grok-build-mcp-server

npm MCP Registry CI Node License

Install in VS Code Install in Cursor

Ein MCP stdio-Server, der die Grok Build-CLI (grok) als Tools bereitstellt, die Sie von Claude Code, Cursor, VS Code oder jedem anderen MCP-Client aus aufrufen können.

Claude Code  ──stdio/MCP──▶  grok-build-mcp-server  ──spawn──▶  grok CLI  ──▶  xAI API

Es ist ein dünner Prozess-Wrapper. Es implementiert keine Agentenlogik neu und kommuniziert nicht direkt mit der xAI-API – die gesamte Intelligenz bleibt in der grok-CLI. Was dieser Server hinzufügt, ist eine getreue Argumentkonstruktion, robuste Prozessüberwachung und saubere, MCP-förmige Ausgabe.

Status: 0.2.2. Die Tool-Oberfläche ist vollständig. Der Server führt echte kopflose Grok-Agenten im Vordergrund oder abgekoppelt im Hintergrund aus, streamt Fortschritte während sie laufen, stoppt einen Lauf auf Anfrage, überprüft Git-Diffs, recherchiert Fragen im Web, listet die Sitzungen, die diese Läufe erstellt haben, und meldet Sitzung, Nutzung und Kosten. Siehe CHANGELOG.md für das, was ausgeliefert wurde, und ROADMAP.md für das, was in Betracht gezogen und abgelehnt wurde.

Fortschritt

Ein langer Agentenlauf ist sichtbar, während er stattfindet, anstatt eines stillen Wartens, das in einer Textwand endet. Wenn Ihr Client einen progressToken sendet, führt der Server Grok mit --output-format streaming-json aus und leitet eine Benachrichtigung pro Ereignis weiter:

#5  list_dir .
#6  read_file README.md
#7  read_file — completed
#8  thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)

Der Fortschritt verfolgt, was der Agent tut, nicht in welcher Phase er sich befindet. Reasoning- und Antworttext werden zusammengefasst, sodass ein Token-Stream Ihren Client nicht überflutet, während Tool-Aufrufe gemeldet werden, sobald sie auftreten. Clients, die resetTimeoutOnProgress unterstützen, werden während des Laufs nicht auslaufen.

Ein Client, der keinen progressToken sendet, erhält den günstigeren Nicht-Streaming-Pfad und zahlt nichts dafür.

Related MCP server: Claude Code MCP Bridge

Anforderungen

  • Grok Build CLI 1.0.0 oder neuer, authentifiziert (grok models sollte erfolgreich sein)

  • Node.js 22 oder neuer

Wenn grok nicht in Ihrem PATH ist, setzen Sie GROK_BINARY auf den vollständigen Pfad, wenn Sie den Server registrieren.

Installieren

Claude Code

claude mcp add grok-build -- npx -y grok-build-mcp-server

Dann, in Claude Code:

> use the grok-build check tool

check meldet das aufgelöste Binary, die CLI-Version, ob Sie authentifiziert sind, und die aktive Berechtigungsobergrenze. Wenn es zufrieden ist, funktioniert der Rest.

Jeder andere MCP-Client

Der Server spricht MCP über stdio und akzeptiert keine eigenen Argumente:

{
  "mcpServers": {
    "grok-build": {
      "command": "npx",
      "args": ["-y", "grok-build-mcp-server"]
    }
  }
}

VS Code und Cursor akzeptieren die Installations-Badges oben auf dieser Seite, die genau diese Konfiguration enthalten.

Clients, die aus der MCP Registry installieren, kennen diesen Server als io.github.Nuruvala/grok-build-mcp-server. Der Registry-Eintrag wird vom selben Tag wie die npm-Veröffentlichung veröffentlicht und zeigt auf dasselbe Paket.

Wenn npx den Server nicht finden kann

npx löst einen bloßen Paketnamen zuerst gegen das lokale Projekt auf. Wenn das Arbeitsverzeichnis Ihres MCP-Clients ein Checkout dieses Repositorys ist – oder von irgendetwas anderem, dessen package.json grok-build-mcp-server heißt – führt npx -y grok-build-mcp-server den lokalen Einstiegspunkt aus, findet keinen und schlägt mit command not found fehl. Installieren Sie es an einem eigenen Ort und registrieren Sie diesen Pfad:

npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-server

Berechtigungen

Grok-Läufe, die über diesen Server gestartet werden, sind standardmäßig schreibgeschützt: --permission-mode plan mit --sandbox read-only. Nichts kann Ihre Dateien ändern, bis Sie es sagen.

Die Berechtigung ist eine Obergrenze, die einmal beim Registrieren des Servers gesetzt wird, anstatt einer Aufforderung bei jedem Aufruf. Drei Stufen:

Stufe

--permission-mode

--sandbox

Was es erlaubt

read-only (Standard)

plan

read-only

Lesen und Denken. Keine Änderungen

write

acceptEdits

workspace

Änderungen innerhalb des Arbeitsverzeichnisses

full

bypassPermissions

off

Unbeaufsichtigte vollständige Genehmigung

Um Grok Änderungen vornehmen zu lassen:

claude mcp add grok-build \
  -e GROK_MCP_PERMISSION_CEILING=write \
  -e GROK_MCP_DEFAULT_PERMISSION=write \
  -- npx -y grok-build-mcp-server

Verwenden Sie full nur, wenn Sie Ihren MCP-Client bereits mit vollständiger Genehmigung ausführen und möchten, dass der delegierte Grok-Lauf ebenso unbeaufsichtigt ist. Es gewährt dem gestarteten grok-Prozess dieselbe Autorität, die Sie haben.

Ein Aufruf, der mehr als die Obergrenze anfordert, wird abgelehnt, nicht stillschweigend herabgestuft – ein begrenzter Lauf würde Erfolg melden, während er nichts ändert, was schlimmer ist als ein klarer Fehler.

Umgebungsvariablen

Variable

Standard

Zweck

GROK_BINARY

grok

Pfad zur grok-Ausführbaren

GROK_MCP_PERMISSION_CEILING

read-only

Höchste Stufe, die ein Aufruf anfordern darf

GROK_MCP_DEFAULT_PERMISSION

read-only

Stufe, die verwendet wird, wenn ein Aufruf keine anfordert

GROK_MCP_DEFAULT_MODEL

grok-4.6

Modell, wenn ein Aufruf keins angibt. none delegiert an die CLI

GROK_MCP_DEFAULT_EFFORT

high

Reasoning-Aufwand, wenn ein Aufruf keinen angibt. none delegiert an die CLI

GROK_MCP_TIMEOUT_MS

1800000

Echtzeit für einen einzelnen Lauf

GROK_MCP_STATE_DIR

$XDG_STATE_HOME/grok-mcp

Hintergrundjob-Aufzeichnungen

GROK_MCP_MAX_CONCURRENT_RUNS

4

Gleichzeitig aktive Hintergrundläufe. off für keine Begrenzung

GROK_MCP_LOG_LEVEL

info

debug, info, warn, error. Logs gehen nach stderr

STRUCTURED_CONTENT_ENABLED

aus

Auch structuredContent zusammen mit _meta ausgeben

Grok's eigene Variablen (XAI_API_KEY, GROK_HOME, GROK_DISABLE_AUTOUPDATER) werden unverändert an den Kindprozess weitergegeben.

Tools

Tool

Schreibgeschützt

Zweck

grok

je nach Obergrenze

Führen Sie einen kopflosen Grok-Agenten aus. Prompt, Sitzung fortsetzen/weiterführen/verzweigen, Modell, Aufwand, Tool erlauben/verbieten

review

immer

Überprüfen Sie einen Git-Diff: Arbeitsbaum, einen Merge-Base-Diff gegen einen Ref, oder einen einzelnen Commit

websearch

immer

Recherchieren Sie eine Frage im Web und melden Sie, welche Suchen und Quellen tatsächlich verwendet wurden

status

immer

Fragen Sie einen Hintergrundlauf ab oder listen Sie kürzliche auf

stop

nein

Beenden Sie den Prozessbaum eines Hintergrundlaufs

sessions

immer

Listen, durchsuchen und schlagen Sie die Grok-Sitzungen auf diesem Rechner nach

check

ja

Serverversion, aufgelöstes Binary, grok version, Authentifizierung, Berechtigungsobergrenze, Laufstandards

help

ja

grok --help Durchleitung

review

Der Diff wird prozessintern gesammelt und in den Prompt eingebettet, sodass das Modell keine Züge damit verbringt, wiederzuentdecken, was es überprüfen soll.

> review my working tree with grok-build
> review the diff against origin/main

Ziele sind uncommitted, base: "<ref>" (ein Merge-Base-Diff, sodass Commits, die nach Ihrem Branch auf der Basis gelandet sind, nicht Ihnen zugeschrieben werden), oder commit: "<sha>". Ohne Angabe erkennt es automatisch: den Upstream-Diff, wenn Ihr Branch voraus ist, ansonsten den Arbeitsbaum – und es sagt, welches es gewählt hat, anstatt still zu raten.

review ist immer schreibgeschützt, unabhängig davon, was GROK_MCP_PERMISSION_CEILING erlaubt. Es akzeptiert kein permission-, write- oder yolo-Argument, denn eine Überprüfung, die den überprüften Code bearbeitet, ist nie das, was gewünscht war.

Übergeben Sie structured: true für maschinenlesbare Ergebnisse (severity, file, line, summary, rationale) auf _meta.findings, validiert bevor Sie sie sehen.

Zwei verschiedene Dinge können schiefgehen, und sie werden unterschiedlich gemeldet, anstatt verschwommen zu werden:

  • Der Lauf wurde nie beendet – er wurde abgebrochen oder endete, ohne seine Ergebnisse zu liefern. Es gibt keine Überprüfung, also ist der Aufruf isError: true und _meta.findingsComplete ist false. Der Text beginnt mit dem Warum, zitiert den eigenen Grund der CLI und nennt die Lösung, die zur tatsächlichen Ursache passt.

  • Der Lauf wurde beendet, aber seine Ausgabe wird nicht validieren. Der Aufruf ist trotzdem erfolgreich und gibt den Rohtext plus einen _meta.parseError zurück – eine degradierte Überprüfung ist besser als eine fehlgeschlagene.

Was Sie nie bekommen werden, ist ein plausibel aussehendes Ergebnis, das das Modell erfunden hat. --json-schema schränkt jede Nachricht ein, die das Modell ausgibt, sodass es während des Lesens keine Möglichkeit hat, "Ich arbeite" zu sagen, außer in Form eines Ergebnisses – und wenn man es nicht kontrolliert, tut es genau das. Das Schema trägt ein erforderliches status-Feld, um diese Erzählung aus Ihren Ergebnissen herauszuhalten, und nichts wird jemals durch Mustervergleich aus einer Teilantwort gerettet.

Strukturierte Überprüfungen großer Ziele scheitern auf diese Weise mit einiger Regelmäßigkeit. Der Fehler ist absichtlich laut.

Eine Überprüfung, die nach einer Shell greift, wird verweigert, nicht getötet. Im kopflosen Modus bricht eine nicht genehmigbare Tool-Anfrage den gesamten Lauf ab, während die CLI immer noch mit 0 beendet wird, also verweigert review die Shell- und Bearbeitungstools direkt – dem Modell wird Nein gesagt und es beendet seine Überprüfung, anstatt mitten im Satz zu sterben.

websearch

> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depth

numResults (1–50) und searchDepth (basic oder full) formen den Prompt – die grok-CLI hat keine Flags für beides, und keiner der Parameter tut so, als ob. Sie funktionieren: Dieselbe Frage, bei basic gestellt, führte eine Suche über zwei Seiten durch, und bei full sechs Suchen über drei, für zweieinhalb Mal die Kosten.

Das Ergebnis sagt Ihnen, was tatsächlich nachgeschlagen wurde, nicht nur, was das Modell geschrieben hat:

[1 web search, 9 sources]

mit _meta, das webSearches, webToolCalls, searchQueries, sources, sourceCount, pagesOpened und searchPerformed enthält. Das ist wichtiger, als es klingt. Grok kann über Websuche oder über X recherchieren, und wenn das Web nicht verfügbar ist, wird es leise das Zweite tun – selbstbewusst antworten, x.com zitieren, erfolgreich beenden. Der Text gibt Ihnen keine Möglichkeit, das zu erkennen. Ein Lauf, der X und nicht das Web durchsucht hat, sagt das in seiner ersten Zeile und meldet xSearches separat, und ein Lauf, bei dem nichts zurückkam, ist ein Fehler, anstatt einer selbstbewusst aussehenden Antwort aus dem eigenen Gedächtnis des Modells:

No search ran. The answer below is the model's own prior knowledge, not current sources.

searchPerformed bedeutet, dass Quellen zurückkamen – nicht, dass eine Suche versucht wurde. Eine Suche, die begann und nie zurückkehrte, oder eine leere Ergebnismenge zurückgab, wird als das gemeldet, was sie war.

Wie review ist websearch immer schreibgeschützt und akzeptiert weder permission, write noch yolo als Argument. Es übergibt nie --disable-web-search.

Hintergrundläufe, status und stop

Ein langer Agentenlauf muss nicht Ihren Client belegen. Übergeben Sie background: true an grok, review oder websearch, und der Aufruf gibt sofort eine runId zurück, während ein abgetrennter Worker-Prozess den Auftrag bis zum Ende ausführt:

> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that run

Der Lauf gehört zum Rechner, nicht zu diesem Server: Er läuft weiter, wenn Ihr MCP-Client die Verbindung trennt, wenn der Server neu startet oder wenn Sie Ihren Editor schließen. Aufzeichnungen liegen unter GROK_MCP_STATE_DIR, ein Verzeichnis pro Lauf.

status auf einem abgeschlossenen Lauf gibt das zurück, was der synchrone Aufruf zurückgegeben hätte – derselbe Text, dieselben Metadaten, dasselbe Fehlerflag. Hintergrund ist ein Transportmittel für einen Tool-Aufruf, nicht eine zweite Implementierung. Solange ein Lauf aktiv ist, erhalten Sie seinen Status, die verstrichene Zeit, beide Prozess-IDs und das Ende seines Fortschrittslogs; waitMs blockiert für bis zu zwei Minuten und leitet Fortschrittsmeldungen weiter, sobald sie eintreffen. Ein abgelaufenes Warten ist kein Fehler.

Zwei Arten von Unehrlichkeit sind von vornherein ausgeschlossen. Ein Lauf, dessen Worker-Prozess nicht mehr existiert, wird als abandoned gemeldet und nicht als noch laufend – der Rechner wurde neu gestartet oder etwas hat ihn beendet. Und ein Lauf, der frühzeitig beendet wurde, wird als solcher gekennzeichnet:

mfk2p1x9-3ac71f0b  completed (cut off: cancelled)  grok  4m 12s  refactor the parser

Die Validierung erfolgt immer noch, bevor Sie eine runId erhalten: Eine Anfrage über GROK_MCP_PERMISSION_CEILING oder ein widersprüchliches Paar von Session-Flags wird als fehlgeschlagener Aufruf abgelehnt, anstatt angenommen und dann in einem Prozess, den niemand beobachtet, zum Scheitern gebracht zu werden.

stop beendet einen Lauf vorzeitig. Es sendet der gesamten Prozessgruppe des Workers – dem Worker und dem von ihm gestarteten grok-Prozess – ein SIGTERM, dann SIGKILL, falls das nicht ausreicht. Das Stoppen eines bereits beendeten Laufs ist kein Fehler, ebenso wenig wie das Stoppen eines Laufs, der kurz vor Ihrem Aufruf beendet wurde.

Ein Stopp, der den Prozessbaum nicht beenden konnte, wird als Fehler gemeldet, nicht als gestoppter Lauf. Wenn es nichts zu signalisieren gibt, oder der Kill verweigert wird, oder der Baum SIGKILL überlebt, bleibt der Lauf auf running und der Aufruf gibt einen Fehler mit der PID zurück. Ein cancelled-Eintrag neben einem laufenden Prozess wäre die aufgeräumtere Antwort und die nutzlose.

Ein Lauf, den Sie mittendrin stoppen, hat normalerweise bereits etwas Nützliches hervorgebracht, und sowohl das Teilergebnis als auch die Session-ID bleiben erhalten:

Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.

The run was cancelled mid-flight, but it recorded a session before it ended:
  grok -r 01a010e2-478c-73d2-bce9-23552245c64d

Grok meldet eine Session-ID nur, wenn ein Lauf sein Ende erreicht, was ein gestoppter Lauf nie tut – daher wird diese ID aus dem eigenen Session-Store der CLI ausgelesen, anstatt rekonstruiert zu werden. _meta.sessionIdSource sagt Ihnen, welche Sie haben. Wenn zwei Läufe im selben Verzeichnis beide passen könnten, erhalten Sie die Kandidaten-IDs und keinen Wiederaufnahmebefehl: Die Wiederaufnahme der falschen Session setzt die Arbeit einer anderen Person fort.

sessions

Jeder Grok-Lauf hinterlässt eine Session auf der Festplatte, und jede Session-ID, die dieser Server meldet, kann später fortgesetzt werden – von jedem Verzeichnis aus, von Ihnen im Terminal oder durch einen anderen Tool-Aufruf.

> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiter

Sessions werden aus $GROK_HOME/sessions (Standard ~/.grok/sessions) gelesen, dem eigenen Store der CLI, sodass sie Neustarts dieses Servers, Ihres MCP-Clients und Ihres Rechners überleben. Übergeben Sie id für eine einzelne Session, query für eine groß-/kleinschreibungsunabhängige Suche über Titel, erste Prompts und IDs, cwd zur Eingrenzung auf ein Projekt und limit zur Begrenzung der Liste.

Ein gerade abgeschlossener Lauf hat noch keinen Titel – Grok füllt diese später nach, falls überhaupt – daher fallen Zeilen auf den ersten Prompt der Session zurück, und titleSource sagt Ihnen, was Sie gerade sehen. Jede Zeile trägt resumeCommand, ebenso wie jedes grok- und review-Ergebnis:

grok -r 01a00c8d-970c-7531-8a12-31dac582c22b

Die Suche ist nur lokal. grok sessions search konsultiert auch ein entferntes Index; dieses Tool tut das nicht, daher wird eine Session, die nur serverseitig existiert, nicht angezeigt.

Entwicklung

npm install
npm run build          # tsc -> dist/
npm run dev            # tsx src/index.ts
npm test               # node --test via tsx
npm run test:coverage  # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run format
  • docs/api-reference.md – Parameter jedes Tools, Ergebnistext, _meta-Schlüssel und die genauen Bedingungen, unter denen jeder gesetzt wird.

  • docs/security.md – Was die Registrierung dieses Servers autorisiert, was jede Berechtigungsstufe tatsächlich gewährt und was Ihren Rechner verlässt.

  • docs/engineering.md – Wie Code hier geschrieben wird: Architektur, funktionale TypeScript-Regeln, Fehler- und Effekt-Disziplin, Test- und Coverage-Richtlinie, Commit-Workflow.

  • CLAUDE.md – Projekthintergrund und das verifizierte grok-CLI-Verhalten, auf das dieser Server angewiesen ist.

  • ROADMAP.md – Meilensteine, Akzeptanzkriterien und die Ideen, die gemessen und verworfen wurden.

Veröffentlichung

Erhöhen Sie die version in package.json, verschieben Sie den Abschnitt Unreleased von CHANGELOG.md unter die neue Versionsüberschrift, committen Sie, dann:

git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0

.github/workflows/release.yml führt das vollständige Gate aus, verweigert die Veröffentlichung, wenn Tag und package.json nicht übereinstimmen, installiert das gepackte Tarball in ein temporäres Verzeichnis und führt eine echte initialize gegen die installierte Binärdatei aus, veröffentlicht dann dieselbe Datei und erstellt ein GitHub-Release.

Es gibt keine zu verwaltenden Veröffentlichungsanmeldeinformationen. Die Authentifizierung erfolgt über npm Trusted Publishing: Der Workflow tauscht ein kurzlebiges OIDC-Token aus, und npm generiert die Herkunftsbestätigung selbstständig. Die Vertrauensstellung ist gegen dieses Repository und den Dateinamen dieses Workflows registriert, sodass die Umbenennung von release.yml die Veröffentlichung unterbricht – und npm prüft die Konfiguration erst, wenn eine Veröffentlichung versucht wird, wobei das Symptom ENEEDAUTH ist, nichts, das die Ursache benennt.

Lizenz

MIT – siehe LICENSE.

Available Tools

8 tools
checkCheck Grok Build readinessA
Read-onlyIdempotent

Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokRun Grok BuildA

Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`.
denyNoRepeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`.
yoloNoShorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request.
agentNoNamed subagent to run, passed as `--agent`.
allowNoRepeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
rulesNoExtra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt.
toolsNoInternal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`.
writeNoShorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
promptYesThe task for Grok to perform. Passed verbatim as `grok -p`.
resumeNoResume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
sessionIdNoCreate a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
permissionNoPermission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change.
forkSessionNoUUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`.
continueSessionNoContinue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request.
disallowedToolsNoInternal tool ids to block, passed as `--disallowed-tools`.
disableWebSearchNoPass `--disable-web-search`. `false` is not a request.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

helpGrok CLI helpA
Read-onlyIdempotent

Show the grok CLI help text. Runs grok --help and returns its stdout.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reviewReview a git diffA
Read-only

Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Repository to review. Defaults to the current working directory.
baseNoReview the merge-base diff against this ref. Mutually exclusive with commit and uncommitted.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
commitNoReview this commit. Mutually exclusive with base and uncommitted.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
structuredNoReturn machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request.
uncommittedNoReview the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request.
instructionsNoExtra reviewer guidance, appended verbatim to the prompt.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sessionsList Grok sessionsA
Read-onlyIdempotent

List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoExact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match.
cwdNoKeep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`).
limitNoMaximum rows to return. Default 20. Ignored when `id` is set.
queryNoCase-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusPoll a background runA
Read-onlyIdempotent

Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
tailNoBytes of progress.log to include for a live run. Default 8192.
limitNoMaximum rows to return in list mode. Default 20. Ignored when `runId` is set.
runIdNoId of a background run to inspect. Omit to list recent runs.
waitMsNoBlock up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stopStop a background runA
DestructiveIdempotent

Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesThe runId returned by a background `grok`, `review`, or `websearch` call.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

websearchSearch the web with Grok BuildA
Read-only

Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory.
modelNoModel id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server.
queryYesThe question to research. Passed as the body of a web-search-shaped prompt.
effortNoReasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise.
maxTurnsNoMaximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research.
backgroundNoRun detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request.
numResultsNoPrompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag.
searchDepthNoPrompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag.
instructionsNoExtra researcher guidance, appended verbatim to the prompt.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.2.4
    • Changedgrok2 fields changed
      • changedInput schema / properties / cwd / description
        Previous value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`."
      • changedInput schema / properties / permission / description
        Previous value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
  2. 8 tool updatesv0.2.2
    • First observedcheck
    • First observedgrok
    • First observedhelp
    • First observedreview
    • First observedsessions
    • First observedstatus
    • First observedstop
    • First observedwebsearch

TDQS

A4.5/5.0

Scored across 8 tools

Disambiguation5/5

Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.

Naming Consistency4/5

All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.

Tool Count5/5

Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.

Completeness5/5

The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers