grok-build-mcp-server
grok-build-mcp-server
Grok Build CLI(grok)를 Claude Code, Cursor, VS Code 또는 기타 MCP 클라이언트에서 호출할 수 있는 도구로 노출하는 MCP stdio 서버입니다.
Claude Code ──stdio/MCP──▶ grok-build-mcp-server ──spawn──▶ grok CLI ──▶ xAI API이는 가벼운 프로세스 래퍼입니다. 에이전트 로직을 재구현하지 않으며 xAI API와 직접 통신하지 않습니다. 모든 지능은 grok CLI에 있습니다. 이 서버가 추가하는 것은 정확한 인수 구성, 강력한 프로세스 감독 및 깔끔한 MCP 형식의 출력입니다.
상태: 0.2.2. 도구 표면이 완성되었습니다. 서버는 실제 헤드리스 Grok 에이전트를 포그라운드 또는 백그라운드에서 분리하여 실행하고, 실행 중 진행 상황을 스트리밍하며, 요청 시 실행을 중지하고, git diff를 검토하고, 웹에서 질문을 조사하며, 해당 실행이 생성한 세션을 나열하고, 세션, 사용량 및 비용을 보고합니다. 출시된 내용은 CHANGELOG.md를, 고려 및 기각된 내용은 ROADMAP.md를 참조하세요.
진행 상황
긴 에이전트 실행은 텍스트 벽으로 끝나는 조용한 대기가 아니라 실행 중에 볼 수 있습니다. 클라이언트가 progressToken을 보내면 서버는 --output-format streaming-json으로 Grok을 실행하고 이벤트당 알림을 전달합니다.
#5 list_dir .
#6 read_file README.md
#7 read_file — completed
#8 thinking: the user asked me to list files, read README.md, then …
#10 writing: DONE
#11 finished: end_turn (2 turns)진행 상황은 에이전트가 어떤 단계에 있는지가 아니라 무엇을 하고 있는지를 추적합니다. 추론 및 응답 텍스트는 통합되어 토큰 스트림이 클라이언트를 넘치지 않도록 하는 반면, 도구 호출은 발생하는 대로 보고됩니다. resetTimeoutOnProgress를 지원하는 클라이언트는 실행 중간에 시간 초과되지 않습니다.
progressToken을 보내지 않는 클라이언트는 더 저렴한 비스트리밍 경로를 사용하며 이에 대한 비용을 지불하지 않습니다.
Related MCP server: Claude Code MCP Bridge
요구 사항
Grok Build CLI 1.0.0 이상, 인증됨 (
grok models가 성공해야 함)Node.js 22 이상
grok이 PATH에 없으면 서버를 등록할 때 GROK_BINARY를 전체 경로로 설정하세요.
설치
Claude Code
claude mcp add grok-build -- npx -y grok-build-mcp-server그런 다음 Claude Code에서:
> use the grok-build check toolcheck는 확인된 바이너리, CLI 버전, 인증 여부 및 활성 권한 상한을 보고합니다. 문제가 없으면 나머지도 작동합니다.
기타 MCP 클라이언트
서버는 stdio를 통해 MCP와 통신하며 자체 인수를 사용하지 않습니다.
{
"mcpServers": {
"grok-build": {
"command": "npx",
"args": ["-y", "grok-build-mcp-server"]
}
}
}VS Code와 Cursor는 이 페이지 상단의 설치 배지를 허용하며, 이 배지는 정확히 해당 구성을 전달합니다.
MCP 레지스트리에서 설치하는 클라이언트는 이 서버를 io.github.Nuruvala/grok-build-mcp-server로 인식합니다. 레지스트리 항목은 npm 릴리스와 동일한 태그에서 게시되며 동일한 패키지를 가리킵니다.
npx가 서버를 찾을 수 없는 경우
npx는 먼저 로컬 프로젝트에 대해 베어 패키지 이름을 확인합니다. MCP 클라이언트의 작업 디렉터리가 이 저장소의 체크아웃이거나 package.json의 이름이 grok-build-mcp-server인 다른 항목인 경우 npx -y grok-build-mcp-server는 로컬 진입점을 실행하고 찾을 수 없어 command not found 오류와 함께 실패합니다. 자체 디렉터리에 설치하고 해당 경로를 등록하세요.
npm install --prefix ~/.local/share/grok-build-mcp grok-build-mcp-server
claude mcp add grok-build -- ~/.local/share/grok-build-mcp/node_modules/.bin/grok-build-mcp-server권한
이 서버를 통해 실행되는 Grok 실행은 기본적으로 읽기 전용입니다: --permission-mode plan과 --sandbox read-only입니다. 사용자가 허용할 때까지 파일을 수정할 수 없습니다.
권한은 상한이며, 호출할 때마다 프롬프트가 표시되는 것이 아니라 서버를 등록할 때 한 번 설정됩니다. 세 가지 수준이 있습니다.
수준 |
|
| 허용되는 작업 |
|
|
| 읽기 및 추론. 편집 불가 |
|
|
| 작업 디렉터리 내 편집 |
|
|
| 무인 전체 승인 |
Grok이 편집할 수 있도록 하려면:
claude mcp add grok-build \
-e GROK_MCP_PERMISSION_CEILING=write \
-e GROK_MCP_DEFAULT_PERMISSION=write \
-- npx -y grok-build-mcp-server이미 MCP 클라이언트를 전체 승인으로 실행 중이고 위임된 Grok 실행도 동등하게 무인 상태가 되기를 원하는 경우에만 full을 사용하세요. 생성된 grok 프로세스에 사용자와 동일한 권한을 부여합니다.
상한을 초과하는 요청은 자동으로 낮춰지지 않고 거부됩니다. 제한된 실행은 아무것도 변경하지 않으면서 성공을 보고하므로 명확한 오류보다 더 나쁩니다.
환경 변수
변수 | 기본값 | 목적 |
|
|
|
|
| 모든 호출이 요청할 수 있는 최고 수준 |
|
| 호출이 아무것도 요청하지 않을 때 사용되는 수준 |
|
| 호출이 모델을 생략할 때의 모델. |
|
| 호출이 노력을 생략할 때의 추론 노력. |
|
| 단일 실행의 벽시계 시간 |
|
| 백그라운드 작업 레코드 |
|
| 동시에 활성화된 백그라운드 실행. |
|
|
|
| off |
|
Grok의 자체 변수(XAI_API_KEY, GROK_HOME, GROK_DISABLE_AUTOUPDATER)는 자식 프로세스에 그대로 전달됩니다.
도구
도구 | 읽기 전용 | 목적 |
| 상한 기준 | 헤드리스 Grok 에이전트 실행. 프롬프트, 세션 재개/계속/포크, 모델, 노력, 도구 허용/거부 |
| 항상 | git diff 검토: 작업 트리, 참조에 대한 병합 기준 diff 또는 단일 커밋 |
| 항상 | 웹에서 질문 조사 및 실제 사용된 검색어와 출처 보고 |
| 항상 | 백그라운드 실행 폴링 또는 최근 실행 목록 표시 |
| 아니요 | 백그라운드 실행의 프로세스 트리 종료 |
| 항상 | 이 머신의 Grok 세션 목록, 검색 및 조회 |
| 예 | 서버 버전, 확인된 바이너리, |
| 예 |
|
review
diff는 프로세스 내에서 수집되어 프롬프트에 포함되므로 모델이 검토해야 할 대상을 다시 발견하는 데 시간을 소비하지 않습니다.
> review my working tree with grok-build
> review the diff against origin/main대상은 uncommitted, base: "<ref>"(병합 기준 diff이므로 분기 후 베이스에 도착한 커밋은 사용자에게 귀속되지 않음) 또는 commit: "<sha>"입니다. 아무것도 제공되지 않으면 자동 감지합니다: 분기가 앞서 있을 때는 업스트림 diff, 그렇지 않으면 작업 트리이며, 조용히 추측하는 대신 어떤 것을 선택했는지 알려줍니다.
review는 GROK_MCP_PERMISSION_CEILING이 허용하는 것과 관계없이 항상 읽기 전용입니다. 검토 중인 코드를 편집하는 검토는 원하는 것이 아니므로 permission, write 또는 yolo 인수를 사용하지 않습니다.
구조화된 결과를 위해 structured: true를 전달하면 검증된 후 _meta.findings에 머신 판독 가능한 결과(severity, file, line, summary, rationale)가 제공됩니다.
두 가지 다른 문제가 발생할 수 있으며, 혼동되지 않고 다르게 보고됩니다.
실행이 완료되지 않음 — 중단되거나 결과를 생성하지 않고 종료되었습니다. 검토가 없으므로 호출은
isError: true이고_meta.findingsComplete는false입니다. 본문은 CLI 자체의 이유를 인용하여 이유를 먼저 설명하고 실제 원인에 맞는 수정 사항을 지정합니다.실행이 완료되었지만 출력의 유효성을 검사할 수 없음. 호출은 여전히 성공하며 원시 텍스트와
_meta.parseError를 반환합니다. 저하된 검토가 실패한 검토보다 낫습니다.
절대 얻지 못할 것은 모델이 지어낸 그럴듯해 보이는 결과입니다. --json-schema는 모델이 내보내는 모든 메시지를 제한하므로, 모델이 여전히 읽는 중일 때는 결과 형태 외에는 "작업 중"이라고 말할 방법이 없으며, 확인되지 않은 상태에서는 정확히 그렇게 합니다. 스키마에는 필수 status 필드가 있어 해당 내레이션을 결과에서 제외하며, 패턴 일치로 부분 응답에서 어떤 것도 구출되지 않습니다.
대규모 대상의 구조화된 검토는 이러한 방식으로 자주 실패합니다. 실패는 의도적으로 명확하게 표시됩니다.
셸에 도달하는 검토는 거부되며 종료되지 않습니다. 헤드리스 모드에서 승인할 수 없는 도구 요청은 CLI가 여전히 0으로 종료되는 동안 전체 실행을 취소하므로 review는 셸 및 편집 도구를 완전히 거부합니다. 모델은 거절당하고 문장 중간에 죽는 대신 검토를 완료합니다.
websearch
> websearch: what changed in the latest Bun release?
> search the web for how Postgres handles advisory lock contention, in depthnumResults(1–50) 및 searchDepth(basic 또는 full)는 프롬프트를 구성합니다. grok CLI에는 둘 다에 대한 플래그가 없으며, 두 매개변수 모두 그렇지 않은 척하지 않습니다. 그러나 작동합니다: basic에서 동일한 질문을 했을 때 두 페이지에 걸쳐 한 번 검색했고, full에서는 세 페이지에 걸쳐 여섯 번 검색하여 두 배 반의 비용이 들었습니다.
결과는 모델이 작성한 내용뿐만 아니라 실제로 조회된 내용을 알려줍니다.
[1 web search, 9 sources]_meta에는 webSearches, webToolCalls, searchQueries, sources, sourceCount, pagesOpened 및 searchPerformed가 포함됩니다. 이는 생각보다 중요합니다. Grok은 웹 검색이나 X를 통해 조사할 수 있으며, 웹을 사용할 수 없을 때는 조용히 두 번째 방법을 수행합니다. 자신 있게 답변하고 x.com을 인용하며 성공적으로 종료합니다. 산문만으로는 이를 구분할 방법이 없습니다. 따라서 X를 검색하고 웹을 검색하지 않은 실행은 첫 줄에 이를 명시하고 xSearches를 별도로 보고하며, 아무것도 반환되지 않은 실행은 모델 자체 메모리의 자신감 있는 답변 대신 오류로 처리됩니다.
No search ran. The answer below is the model's own prior knowledge, not current sources.searchPerformed는 소스가 반환되었음을 의미합니다. 검색이 시도되었음을 의미하는 것이 아닙니다. 시작되었지만 반환되지 않았거나 빈 결과 집합을 반환한 검색은 있었던 그대로 보고됩니다.
review와 마찬가지로 websearch는 항상 읽기 전용이며 permission, write, yolo 인자를 받지 않습니다. 또한 --disable-web-search를 전달하지 않습니다.
백그라운드 실행, status 및 stop
긴 에이전트 실행이 반드시 클라이언트를 점유할 필요는 없습니다. grok, review 또는 websearch에 background: true를 전달하면 호출이 즉시 runId를 반환하고, 분리된 워커 프로세스가 작업을 완료할 때까지 실행합니다:
> have grok refactor the parser in the background
> status
> status the run from a minute ago and wait 30s for it
> stop that run실행은 서버가 아닌 머신에 속합니다. MCP 클라이언트가 연결을 끊거나, 서버가 재시작되거나, 에디터를 닫아도 계속 실행됩니다. 레코드는 GROK_MCP_STATE_DIR 아래에 저장되며, 각 실행마다 하나의 디렉토리가 생성됩니다.
완료된 실행에 대한 status는 동기 호출이 반환했을 값과 동일합니다 — 동일한 텍스트, 동일한 메타데이터, 동일한 오류 플래그. 백그라운드는 툴 호출을 위한 전송 방식일 뿐, 툴의 두 번째 구현이 아닙니다. 실행이 진행 중인 동안에는 해당 상태, 경과 시간, 두 프로세스 ID, 그리고 진행 로그의 마지막 부분을 얻을 수 있습니다. waitMs는 최대 2분 동안 블로킹하며, 진행 알림이 도착하는 대로 전달합니다. 시간 초과된 대기는 오류가 아닙니다.
구조적으로 두 가지 부정확성은 배제됩니다. 워커 프로세스가 더 이상 존재하지 않는 실행은 여전히 실행 중인 것으로 보고되지 않고 abandoned(중단됨)로 보고됩니다. 머신이 재부팅되었거나 누군가가 프로세스를 죽인 경우입니다. 그리고 일찍 완료된 실행은 그렇게 표시됩니다:
mfk2p1x9-3ac71f0b completed (cut off: cancelled) grok 4m 12s refactor the parserrunId를 받기 전에도 검증이 이루어집니다. GROK_MCP_PERMISSION_CEILING을 초과하는 요청이나 모순되는 세션 플래그 쌍은 아무도 감시하지 않는 프로세스에서 실패한 후 수락되는 대신 실패한 호출로 거부됩니다.
stop은 실행을 조기에 종료합니다. 워커의 전체 프로세스 그룹(워커와 그것이 생성한 grok 프로세스)에 SIGTERM을 보내고, 그것으로 충분하지 않으면 SIGKILL을 보냅니다. 이미 완료된 실행을 중지하는 것은 오류가 아니며, 호출이 도착하기 직전에 완료된 실행을 중지하는 것도 오류가 아닙니다.
프로세스 트리를 종료할 수 없는 중지는 중단된 실행이 아닌 실패로 보고됩니다. 시그널을 보낼 대상이 없거나, 종료가 거부되거나, 트리가 SIGKILL에서 살아남은 경우, 실행은 running 상태로 남아 있고 호출은 PID를 명시하는 오류를 반환합니다. 살아있는 프로세스 옆에 cancelled 레코드가 있는 것이 더 깔끔한 답이겠지만, 그것은 쓸모없는 답변입니다.
진행 중간에 중지한 실행은 일반적으로 이미 보존할 가치가 있는 무언가를 생성했으며, 부분 결과와 세션 ID가 모두 보존됩니다:
Stopped run msxji60o-8f5e27c4 (grok, ran 20s).
Signalled SIGTERM to process group 1703005; the tree exited.
The run was cancelled mid-flight, but it recorded a session before it ended:
grok -r 01a010e2-478c-73d2-bce9-23552245c64dGrok은 실행이 종료될 때만 세션 ID를 보고합니다. 중지된 실행은 결코 종료되지 않으므로 — 해당 ID는 재구성되는 대신 CLI 자체 세션 저장소에서 다시 읽어옵니다. _meta.sessionIdSource는 어떤 것을 가지고 있는지 알려줍니다. 동일한 디렉토리에 있는 두 실행이 모두 일치할 수 있는 경우, 후보 ID를 제공하고 재개 명령은 제공하지 않습니다: 잘못된 세션을 재개하면 다른 사람의 작업을 계속하게 됩니다.
sessions
모든 Grok 실행은 디스크에 세션을 남기며, 이 서버가 보고하는 모든 세션 ID는 나중에 재개될 수 있습니다 — 어떤 디렉토리에서든, 터미널에서 직접 또는 다른 툴 호출을 통해.
> list my recent grok sessions
> what grok sessions did I run in this repo?
> find the grok session about the rate limiter세션은 $GROK_HOME/sessions(기본값 ~/.grok/sessions)에서 읽어옵니다. 이는 CLI 자체 저장소이므로 이 서버, MCP 클라이언트, 그리고 머신의 재시작에도 유지됩니다. 하나의 세션에 대해 id를 전달하거나, 제목, 첫 프롬프트, ID에 대해 대소문자를 구분하지 않는 검색을 위해 query를, 하나의 프로젝트로 범위를 좁히기 위해 cwd를, 목록을 제한하기 위해 limit를 전달하세요.
방금 완료된 실행에는 아직 제목이 없습니다. Grok은 나중에(있을 경우) 제목을 채워넣기 때문에, 행은 세션의 첫 번째 프롬프트로 대체되며, titleSource는 현재 보고 있는 것이 무엇인지 알려줍니다. 모든 행은 resumeCommand를 포함하며, 모든 grok 및 review 결과도 마찬가지입니다:
grok -r 01a00c8d-970c-7531-8a12-31dac582c22b검색은 로컬 전용입니다. grok sessions search는 원격 인덱스도 참조하지만, 이 툴은 그러지 않으므로 서버 측에만 존재하는 세션은 표시되지 않습니다.
개발
npm install
npm run build # tsc -> dist/
npm run dev # tsx src/index.ts
npm test # node --test via tsx
npm run test:coverage # same, with enforced coverage floors
npm run lint
npm run typecheck
npm run formatdocs/api-reference.md — 모든 툴의 매개변수, 결과 텍스트,
_meta키 및 각각이 설정되는 정확한 조건.docs/security.md — 이 서버를 등록할 때 부여되는 권한, 각 권한 수준이 실제로 허용하는 사항, 그리고 머신을 떠나는 데이터.
docs/engineering.md — 코드 작성 방법: 아키텍처, 함수형 TypeScript 규칙, 오류 및 효과 규율, 테스트 및 커버리지 정책, 커밋 워크플로.
CLAUDE.md — 프로젝트 배경 및 이 서버가 의존하는 검증된
grokCLI 동작.ROADMAP.md — 마일스톤, 승인 기준, 그리고 측정 후 기각된 아이디어들.
릴리스
package.json에서 version을 올리고, CHANGELOG.md의 Unreleased 섹션을 새 버전 제목 아래로 이동시킨 후, 커밋하고 다음을 실행하세요:
git tag -a v0.2.0 -m v0.2.0 && git push origin v0.2.0.github/workflows/release.yml은 전체 게이트를 실행하고, 태그와 package.json이 일치하지 않으면 게시를 거부하며, 패키징된 tarball을 임시 디렉토리에 설치하고 설치된 바이너리에 대해 실제 initialize를 구동한 후, 동일한 파일을 게시하고 GitHub 릴리스를 생성합니다.
관리할 게시 자격 증명이 없습니다. 인증은 npm trusted publishing을 사용합니다: 워크플로는 수명이 짧은 OIDC 토큰을 교환하고, npm이 자체적으로 출처 증명(provenance attestation)을 생성합니다. 신뢰는 이 저장소와 이 워크플로의 파일 이름에 등록되므로, release.yml의 이름을 바꾸면 게시가 중단됩니다. npm은 게시가 시도될 때까지 구성을 확인하지 않으며, 증상은 원인을 명명하지 않는 ENEEDAUTH입니다.
라이선스
MIT — LICENSE 참조.
Available Tools
8 toolscheckCheck Grok Build readinessARead-onlyIdempotent
Report grok-build-mcp-server status: version, resolved grok binary, permission ceiling, CLI readiness (grok version, grok models), and run defaults. Call this first when a grok tool behaves unexpectedly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds value by detailing exactly what is reported (version, binary, permission ceiling, CLI readiness, run defaults), giving the agent concrete expectations about the output. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence conveys all necessary information without filler. It is front-loaded with the purpose and lists specific outputs. Slightly dense but efficient; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description fully captures what the tool does and what it returns. It is self-contained: an agent reading it knows exactly when to call it and what information to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters (0 params), and schema coverage is trivially 100%. Per calibration, baseline is 4. The description has no need to explain parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and resource ('grok-build-mcp-server status'), clearly stating it outputs version, binary, permission ceiling, CLI readiness, and run defaults. It distinguishes from siblings by noting it is the first diagnostic step when a grok tool misbehaves, separating it from tools like 'grok', 'status', and 'help'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this first when a grok tool behaves unexpectedly,' providing a clear when-to-use directive. It does not mention exclusions or alternatives, but the context is sufficient for an agent to decide to invoke it for troubleshooting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grokRun Grok BuildA
Run a headless Grok Build agent (grok -p). Returns the model text plus session, usage, and cost metadata. Permission is capped by GROK_MCP_PERMISSION_CEILING; requests above it are rejected rather than silently downgraded.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: "write"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`. | |
| deny | No | Repeatable deny rules in `ToolPrefix(glob)` form, e.g. `Read(.env)`. | |
| yolo | No | Shorthand for `permission: "full"`. Ignored when `permission` is set. `false` is not a request. | |
| agent | No | Named subagent to run, passed as `--agent`. | |
| allow | No | Repeatable allow rules in `ToolPrefix(glob)` form, e.g. `Bash(npm*)`, `Write(src/**)`. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| rules | No | Extra system-prompt text, passed as `--rules`. Longer system-prompt text belongs in the prompt. | |
| tools | No | Internal tool ids to allow, passed as a single comma-joined `--tools`. Shell is `run_terminal_command`, not `bash`. | |
| write | No | Shorthand for `permission: "write"`. Ignored when `permission` is set. `false` is not a request. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| prompt | Yes | The task for Grok to perform. Passed verbatim as `grok -p`. | |
| resume | No | Resume an existing session by id or title (`--resume`). Mutually exclusive with `continueSession`. Combine with `forkSession` to fork rather than continue in place. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| sessionId | No | Create a NEW session with this UUID (`--session-id`). Cannot be combined with `resume` or `continueSession`; use `forkSession` to name a fork. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| permission | No | Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change. | |
| forkSession | No | UUID for a forked session. Requires `resume` or `continueSession`. Passed as `--fork-session --session-id`. | |
| continueSession | No | Continue the most recent session for `cwd` (`--continue`). Mutually exclusive with `resume`. `false` is not a request. | |
| disallowedTools | No | Internal tool ids to block, passed as `--disallowed-tools`. | |
| disableWebSearch | No | Pass `--disable-web-search`. `false` is not a request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds useful behavioral details: the run is headless, it returns model text plus session/usage/cost metadata, and requests above GROK_MCP_PERMISSION_CEILING are rejected rather than silently downgraded. It does not over-explain advanced semantics already covered in the schema, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and every clause earns its place: it states the command, indicates the return payload, and calls out the critical permission-boundary behavior. No fluff or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large 20-parameter tool with no output schema, the description gives essential orientation: what it does, what it returns, and the permission cap. The backing schema supplies the rest. It stops just short of a 5 because it does not summarize the long-running or side-effecting nature of an agent run beyond what annotations and schema already convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 20 parameters with detailed, self-contained descriptions, so the tool description does not need to elaborate. The description adds no parameter-specific detail beyond the permission ceiling note, but the schema carries the burden and does so well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "Run a headless Grok Build agent (`grok -p`)". It clearly distinguishes this from sibling utility tools like status, check, review, and stop by identifying it as the execution/run tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is the tool to invoke a headless Grok Build run, and it adds a meaningful note about permission ceilings. It does not explicitly name alternatives or say when not to use it, but its role as the main run tool is strongly implied and differentiated from sibling inspection/control tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
helpGrok CLI helpARead-onlyIdempotent
Show the grok CLI help text. Runs grok --help and returns its stdout.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds value by revealing the implementation detail that it runs `grok --help` and captures stdout, which is behavioral context beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The first sentence states the purpose, the second provides implementation details. Both are essential for the agent to understand the tool's behavior. Excellent front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, no output schema, and very simple behavior. The description fully captures what the tool does, how it works (runs a command), and what it returns (stdout). For a help tool, this is completely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (no parameters exist). The description mentions no arguments, which is consistent. With 0 parameters, the baseline is 4, and the description adds no further info about parameters because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs `grok --help` and returns its stdout, specifying the exact verb ('show'), resource ('Grok CLI help text'), and execution method. This distinguishes it entirely from sibling tools like `check` or `websearch`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this tool (to show the grok CLI help text), but does not provide explicit guidance on when not to use it or mention alternatives among siblings. For a tool with 0 parameters and a narrow, well-defined purpose, this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviewReview a git diffARead-only
Review a git diff with Grok Build. Targets the working tree (uncommitted), a merge-base diff against base, or a single commit. When none is specified, auto-detects: the upstream diff if the branch is ahead, otherwise the working tree. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a review that edits the code it is reviewing is never wanted. Set structured: true for machine-readable findings.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Repository to review. Defaults to the current working directory. | |
| base | No | Review the merge-base diff against this ref. Mutually exclusive with commit and uncommitted. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| commit | No | Review this commit. Mutually exclusive with base and uncommitted. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| structured | No | Return machine-readable findings via `--json-schema`. A run that stops before a final findings object fails the call with reviewIncomplete. Malformed model JSON after a normal stop degrades to raw text plus a parseError field rather than failing the call. `false` is not a request. | |
| uncommitted | No | Review the working tree (staged, unstaged, and untracked). Mutually exclusive with base and commit. `false` is not a request. | |
| instructions | No | Extra reviewer guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and destructiveHint=false, so the description reinforces this by explaining why there's no write capability ("a review that edits the code it is reviewing is never wanted") and how it ignores permission ceilings. This adds valuable context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences), efficient, and front-loaded with the core purpose. Every sentence contributes unique value: targets, auto-detection, read-only guarantee, and structured mode option.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, 100% schema coverage, no output schema, and annotations present, the description covers key behavioral aspects (read-only, auto-detection, mutual exclusivity) and provides usage patterns. It doesn't explain return values, but since there's no output schema, the tool likely streams output. A slight gap is not detailing the polling flow for background runs, but overall comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds cross-parameter relationships (mutual exclusivity), auto-detection logic, and the purpose of structured mode, which goes beyond individual parameter schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reviews a git diff using Grok Build. It specifies the three targets (uncommitted, base, commit) and auto-detection behavior, distinguishing it from sibling tools like check, grok, or sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each target mode (working tree, merge-base diff, single commit) and the auto-detection fallback. It also clearly states that review is read-only and lacks permission/write arguments, which helps the agent avoid misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionsList Grok sessionsARead-onlyIdempotent
List and search Grok Build sessions from the local store ($GROK_HOME/sessions). Search is local-only: it does not consult grok sessions search or any remote index. Pass id for a single session, query for a case-insensitive substring over title, first prompt, and id, and cwd to keep only sessions that started in that directory. A reported id resumes from any directory with grok -r <id> or the grok tool's resume argument.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Exact session id lookup. Ignores query, cwd, and limit. Falls back to a case-insensitive match. | |
| cwd | No | Keep only sessions that *started* in this directory. Resume still works from anywhere (`grok -r <id>`). | |
| limit | No | Maximum rows to return. Default 20. Ignored when `id` is set. | |
| query | No | Case-insensitive substring over title, first prompt, and id. Search is local-only: it does not consult `grok sessions search` or any remote index. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds significant behavioral context: the local-only nature, case-insensitive substring matching, parameter interactions (id ignores others, limit ignored when id set), and the ability to resume sessions from any directory using the returned id. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at about 4 sentences, front-loading the main purpose. It includes some repetition of the local-only constraint (appears in both the main description and the query parameter description), but overall it is well-structured and not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, no output schema, and good annotations, the description is largely complete. It explains the local store, parameter behavior, and usage of returned ids. It does not describe the output format, but this is mildly acceptable given the lack of output schema. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds substantial meaning beyond the schema: it explains the role of each parameter in a usage context, specifies that id ignores other parameters, and clarifies that limit is ignored when id is set. This provides a semantic understanding that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (List and search), resource (Grok Build sessions), and scope (local store at $GROK_HOME/sessions). It explicitly distinguishes from remote search by noting it does not consult any remote index, which helps differentiate it from sibling tools like 'grok sessions search'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each parameter (id for single session, query for substring search, cwd for directory filtering, limit for max rows). It also states that search is local-only and not for remote queries. However, no explicit contrast with sibling tools like 'check' or 'review' is given, though the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusPoll a background runARead-onlyIdempotent
Poll a background grok, review, or websearch run, or list recent ones. A finished run replays the original tool result — same text, same metadata, same error flag — so background is a transport, not a second implementation. A run whose worker process has vanished is reported as abandoned rather than as still running. Pass runId to inspect one run, waitMs to block until it finishes, and omit runId to list recent runs.
| Name | Required | Description | Default |
|---|---|---|---|
| tail | No | Bytes of progress.log to include for a live run. Default 8192. | |
| limit | No | Maximum rows to return in list mode. Default 20. Ignored when `runId` is set. | |
| runId | No | Id of a background run to inspect. Omit to list recent runs. | |
| waitMs | No | Block up to this many milliseconds for the run to finish. Default 0. Ignored in list mode. A timed-out wait is not an error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description adds critical behavioral details: finished runs replay the original result verbatim, abandoned runs are reported as such, and a timed-out wait is not an error. This fully informs the agent of runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: first sentence states purpose, second explains result semantics, third gives parameter usage patterns. No redundancy, front-loaded with the primary action. Extremely efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 optional parameters, no output schema, and good annotations, the description covers all necessary aspects: three operational modes, parameter interactions, special cases (abandoned, timed-out wait), and the exact replay behavior. An agent has everything needed to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented. The description enhances this by explaining how parameters interact (omitting runId triggers list mode, waitMs is ignored in list mode) and provides defaults (8192 bytes for tail, 20 limit). This integration-level meaning adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and resource ('background run') and explicitly lists the types of runs (grok, review, websearch). It distinguishes the tool from siblings like 'check', 'stop', and the run-initiating tools by making the polling/list usage obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each parameter combination (runId for inspection, waitMs for blocking, omit runId for listing). While it gives clear context and distinguishes the three modes, it does not explicitly state when not to use this tool or name alternative tools for other scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stopStop a background runADestructiveIdempotent
Terminate a background grok, review, or websearch run: the worker and the grok process it spawned. Stopping an already-finished run is not an error. A run cancelled mid-flight may still have produced a resumable session id, which the result reports.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | The runId returned by a background `grok`, `review`, or `websearch` call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey idempotent and destructive hints. The description adds critical behavioral context beyond annotations: that it terminates both the worker and the spawned grok process, that stopping a finished run is harmless, and that a cancelled run may still yield a session id. This latter point is a non-obvious side effect that an agent must know, which is valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first states what the tool does and its coverage, the second clarifies edge cases. No filler or redundant information. Every sentence adds distinct value, making it highly efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (single parameter, no output schema, no output objects), the description fully covers the tool's purpose, parameter, side effects, and edge cases. The schema and annotations are leveraged well, leaving no obvious gaps for an agent to misunderstand.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents the one parameter (runId) with a format constraint and description. Since schema description coverage is 100%, the baseline is 3. The description adds value by explicitly linking the parameter to the return values of background calls for grok/review/websearch, reinforcing its provenance and acceptable values, which warrants an above-baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Terminate') and clearly identifies the resources it acts on: a background run, the worker, and the spawned grok process. It also distinguishes from siblings by naming the three run types it applies to (grok, review, websearch), making its scope precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage guidance by listing the types of runs it applies to (grok, review, websearch). It also explains a borderline case ('stopping an already-finished run is not an error'), which helps the agent decide when to use this tool without hesitation. However, it does not explicitly state when not to use it or name alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
websearchSearch the web with Grok BuildARead-only
Research a question with Grok Build's web search. numResults and searchDepth shape the prompt only — the CLI has no flags for either. Always runs read-only (--permission-mode plan --sandbox read-only) regardless of GROK_MCP_PERMISSION_CEILING — this tool has no permission, write, or yolo argument, because a search never needs to write. Never passes --disable-web-search.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path. Working directory for the run. Passed as `--cwd`. Defaults to the current working directory. | |
| model | No | Model id to pass as `--model`. Omit to use the server default. Unknown ids are rejected by the CLI, not by this server. | |
| query | Yes | The question to research. Passed as the body of a web-search-shaped prompt. | |
| effort | No | Reasoning effort passed as `--effort`. Omit to use the server default. Values are passed through; the CLI rejects what the model does not advertise. | |
| maxTurns | No | Maximum agentic turns. Passed as `--max-turns`. Headless only. No default — a cap is how a run gets cut off mid-research. | |
| background | No | Run detached and return a runId immediately instead of waiting. Poll with the `status` tool. The run survives a restart of this MCP server. `false` is not a request. | |
| numResults | No | Prompt-level target for how many distinct sources to cite, not a backend limit. The CLI has no `--num-results` flag. | |
| searchDepth | No | Prompt-level search depth. `basic` (default) asks for one round; `full` asks for more than one, from different angles. The CLI has no `--search-depth` flag. | |
| instructions | No | Extra researcher guidance, appended verbatim to the prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing runtime behavior: it always runs with `--permission-mode plan --sandbox read-only` regardless of GROK_MCP_PERMISSION_CEILING, lacks permission/write/yolo arguments, and never passes `--disable-web-search`. This adds significant context not covered by the readOnlyHint and openWorldHint annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, zero filler. Each sentence adds unique information: research purpose, prompt-only parameters, fixed read-only behavior, and special flag avoidance. Front-loaded with the primary verb. No unnecessary words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters (with 100% schema coverage), rich annotations (readOnlyHint, openWorldHint), and no output schema, the description is complete enough. It covers the tool's safety profile, parameter effects, and constraints without needing to detail outputs. No gaps that would confuse an agent selecting or invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. However, the description adds value by clarifying that `numResults` and `searchDepth` only shape the prompt and have no CLI flags, and that `background` runs detached. It also explains `query` is the body of a web-search-shaped prompt. Not quite a 5 because it could weave in more hints about how `effort` and `model` interact with the CLI rejection logic, but still above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it researches a question using web search, with a specific verb ('research') and resource ('Grok Build's web search'). It distinguishes itself from siblings by explicitly noting it never needs to write, which sets it apart from write-oriented tools like grok or review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it always runs read-only with a fixed permission mode, never passes `--disable-web-search`, and explains that `numResults` and `searchDepth` only shape the prompt. It also indirectly suggests when not to use this tool (if write access or a different permission mode is needed), complementing the sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.4- Changed
grok2 fields changed- changed
Input schema / properties / cwd / descriptionPrevious value: -"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path."New value: +"Absolute path. Working directory for the run. Passed as `--cwd`. Use the narrowest useful path. Under `permission: \"write\"` this is also the sandbox root: the run cannot write outside it, and a refused write ends the whole run. Name an output path inside `cwd`, or use `full`." - changed
Input schema / properties / permission / descriptionPrevious value: -"Permission level for this run: `read-only`, `write`, or `full`. Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default."New value: +"Permission level for this run: `read-only` (plan mode, read-only sandbox), `write` (accepts edits, sandboxed to `cwd`), or `full` (no sandbox). Must be at or below GROK_MCP_PERMISSION_CEILING. Omit to use the server default. A tool call the sandbox refuses ends the run with `stopReason: cancelled`, so pick the level from where the run must write, not only from what it must change."
8 tool updates
v0.2.2- First observed
check - First observed
grok - First observed
help - First observed
review - First observed
sessions - First observed
status - First observed
stop - First observed
websearch
TDQS
Scored across 8 tools
Each tool maps to a clearly distinct operation: general agent run, specialized read-only review, web research, background run status, background run termination, session lookup, environment check, and CLI help. The only potential overlap is between grok, review, and websearch, but their descriptions sharply differentiate the general execution mode from the two read-only specialized modes.
All tool names are short, lowercase, single words, so there are no case or separator inconsistencies. However, the set mixes action verbs (check, help, review, stop), resource-like nouns (status, sessions), and a product name (grok), so it follows a loose CLI-subcommand style rather than a strict verb_noun naming convention.
Eight tools is well-scoped for a CLI wrapper server: core execution, two specialized read-only operations, background run lifecycle management, session inspection, diagnostics, and help. Each tool earns its place and none feels redundant.
The toolset covers the full workflow of running Grok Build headlessly, including general runs, diff reviews, web searches, background polling, cancellation, session discovery, and environment readiness checks. While session deletion/export is not exposed, session resumption is supported via the grok tool and sessions tool, so there are no dead ends.
Maintenance
Related MCP Connectors
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Give any MCP-compatible AI assistant a builder for live, hosted web tools and workflows.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables sandboxed file operations via MCP tools, resources, and prompts, with a Claude CLI client and Groq-powered web UI for file CRUD, search, code review, and documentation generation.MIT
- FlicenseNot gradedqualityDmaintenanceExposes Claude Code's file editing, command execution, and test running capabilities as composable MCP tools for any MCP-compatible host, enabling code operations via a stateless bridge.-
- FlicenseAqualityBmaintenanceEnables using the xAI Grok CLI as an MCP sub-agent for code review, asking questions, and continuing conversations within MCP hosts like Claude Code.4-
- AlicenseNot gradedqualityBmaintenanceEnables Codex to use Grok Build CLI as a controlled subagent via MCP tools for independent investigation, review, and isolated implementation tasks.5MIT