Skip to main content
Glama

grokbot-mcp

Outbound MCP server: Cursor, Claude, and VS Code call tools that drive your Grok Bot cloud agents. It wraps grokbot-client (import grokbot) against https://api2.cursor.sh.

This is not the official Grok MCP and not Forge’s inbound grokbot_mcp_gateway.py. Those let a Bot call your tools. This server is the opposite direction (IDE → Bot). Never install this MCP on a Grok Bot.

Not affiliated with Cursor or xAI. The backend is unofficial and can change when the desktop app updates. Bundled client schemas match grok-bot 0.51.0 (SCHEMA_VERSION). MIT licensed.

Package: grokbot_mcp. CLI: grokbot-mcp / python -m grokbot_mcp. Default transport is stdio. --http is Streamable HTTP (stateless=True, json_response=True). Logs go to stderr only (WARNING; httpx/httpcore silenced). PYTHONUNBUFFERED=1.

Server instructions (what the MCP tells the host): grokbot_send is fire-and-forget — it returns message_id and delivery immediately. Retry a send with the same message_id. Use grokbot_wait or grokbot_history for the reply. If the transcript is quiet, check grokbot_pending for widgets or handoff. Only create/delete agents, put_secret, or desktop wake=true when the user asked.

Auth

Sign in from the CLI (no Cursor app required). This is Cursor’s PKCE device login: the CLI prints a loginDeepControl URL, you approve it in a browser, and the process polls https://api2.cursor.sh/auth/poll until it gets an access JWT.

grokbot-mcp login              # prints a URL, opens the browser, waits up to 5 min
grokbot-mcp login --no-open    # URL only (SSH / Docker / headless)
grokbot-mcp login --print-token  # also write the JWT to stdout (avoid on shared terminals)
grokbot-mcp logout             # delete the saved file

Credentials are stored mode 0600 at ~/.config/grokbot-mcp/cursor-auth.json (or $GROKBOT_AUTH_FILE). The MCP process prefers that file, then Cursor host files, then $GROKBOT_TOKEN:

  1. grokbot-mcp login file

  2. ~/.config/cursor/auth.json, ~/.cursor/auth.json, ~/.cursor/cli-config.json

  3. else $GROKBOT_TOKEN (raw JWT)

The login JWT is a session access token, not a Cursor dashboard API key. Do not commit it. Compose mounts ~/.config/grokbot-mcp so a host login works inside the HTTP container.

HTTP also requires $GROKBOT_MCP_TOKEN (a shared secret you mint). That token is not the Cursor JWT.

Related MCP server: grok-bot-mcp

Local venv

Python 3.10+. grokbot-client is not on PyPI.

python3 -m venv .venv
source .venv/bin/activate
pip install "git+https://github.com/Kenzim/grokbot-client.git"
pip install -e ".[dev]"
grokbot-mcp login              # browser PKCE; saves ~/.config/grokbot-mcp/cursor-auth.json
grokbot-mcp                    # stdio
grokbot-mcp --http             # needs GROKBOT_MCP_TOKEN

LAN: pip install "git+https://git.stackken.com/kenzim/grokbot-client.git" or pip install -e /path/to/grokbot-client.

Copy .env.example to .env and fill values. Compose loads .env; stdio does not require it if host files or GROKBOT_TOKEN are present.

cp .env.example .env

Flags: --http, --host (default 127.0.0.1), --port (default 8080), --path (default /mcp). Env wins if a flag is omitted.

IDE config — three transports

Use one of: venv stdio, docker run -i (no -t), or HTTP URL + Bearer. Replace paths and tokens.

Cursor (~/.cursor/mcp.json)

Stdio (venv):

{
  "mcpServers": {
    "grokbot": {
      "command": "/ABS/PATH/grokbot-mcp/.venv/bin/grokbot-mcp"
    }
  }
}

Stdio (Docker). Do not pass -t — TTY breaks MCP framing. -i is required.

{
  "mcpServers": {
    "grokbot": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "GROKBOT_TOKEN",
        "-v", "${HOME}/.config/cursor:/root/.config/cursor:ro",
        "-v", "${HOME}/.cursor:/root/.cursor:ro",
        "git.stackken.com/kenzim/grokbot-mcp:latest"
      ]
    }
  }
}

HTTP (after compose is up):

{
  "mcpServers": {
    "grokbot": {
      "url": "http://127.0.0.1:8080/mcp",
      "headers": { "Authorization": "Bearer YOUR_GROKBOT_MCP_TOKEN" }
    }
  }
}

Claude Desktop

Same mcpServers object in claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/; Windows: %APPDATA%\Claude\; Linux: ~/.config/Claude/). Stdio command/args as above. HTTP uses url + headers when the host supports Streamable HTTP.

VS Code (.vscode/mcp.json)

{
  "servers": {
    "grokbot": {
      "type": "stdio",
      "command": "/ABS/PATH/grokbot-mcp/.venv/bin/grokbot-mcp"
    }
  }
}

Docker stdio: "command": "docker", "args": ["run", "-i", "--rm", ...] (no -t). HTTP:

{
  "servers": {
    "grokbot": {
      "type": "http",
      "url": "http://127.0.0.1:8080/mcp",
      "headers": { "Authorization": "Bearer YOUR_GROKBOT_MCP_TOKEN" }
    }
  }
}

Claude Code

claude mcp add grokbot -- /ABS/PATH/.venv/bin/grokbot-mcp (stdio), or claude mcp add --transport http grokbot http://127.0.0.1:8080/mcp plus an Authorization header / env for Bearer. Project file .mcp.json uses the same three shapes as Cursor.

Docker

Image: python:3.11-slim-bookworm. Build arg GROKBOT_CLIENT_GIT defaults to https://git.stackken.com/kenzim/grokbot-client.git. Public: --build-arg GROKBOT_CLIENT_GIT=https://github.com/Kenzim/grokbot-client.git. ENTRYPOINT is python -m grokbot_mcp; default CMD is empty (stdio).

Stdio container (host MCP spawns it). Interactive stdin, no TTY:

docker build -t grokbot-mcp .
docker run -i --rm \
  -e GROKBOT_TOKEN \
  -v "$HOME/.config/cursor:/root/.config/cursor:ro" \
  grokbot-mcp

Compose HTTP

cp .env.example .env   # set GROKBOT_MCP_TOKEN (required) and GROKBOT_TOKEN if no host files
docker compose up --build

Service grokbot-mcp runs --http --host 0.0.0.0, publishes 127.0.0.1:8080:8080, volume /data (cursors), optional ~/.config/cursor:ro, env_file: .env. HEALTHCHECK curls /health.

  • GET /health — unauthenticated {"ok": true}

  • /mcp — Authorization: Bearer <GROKBOT_MCP_TOKEN>; missing/wrong token → 401

  • HTTP refuses to start without GROKBOT_MCP_TOKEN

  • Rate limit: 60 requests/min per token (override GROKBOT_MCP_RATE_PER_MIN)

The HTTP process starts the in-process transcript watcher immediately. Stdio starts it lazily on first wait / pending / resource read.

Tools

Every result is one MCP text blob: JSON {"ok": true, ...} or {"ok": false, "error": "...", "code": "..."}. Codes map UnauthorizedError / AuthError / RefusalError / NotFoundError / GrokBotError. Protobuf .raw is never serialized. Default session_id is "" (MAIN). Optional session_id on chat/live tools.

Reads have readOnlyHint. grokbot_delete_*, grokbot_put_secret, and grokbot_desktop with wake have destructiveHint. grokbot_send is mutating and idempotent if message_id is reused.

Roster

  • grokbot_whoami — email, user_id, SCHEMA_VERSION

  • grokbot_list_agents — id, agent_id, name, description, harness

  • grokbot_get_agent — list row plus todos, sessions, capabilities (agent_id required)

  • grokbot_create_agent — name required; harness box|temporal (default box); optional description, title. Only if the user asked.

  • grokbot_update_agent — agent_id; optional name, description, title, avatar_shape, avatar_color

  • grokbot_delete_agent — preview unless confirm=true (does not call delete on preview)

Chat — send vs wait

  • grokbot_send — fire-and-forget. Args: agent_id, text; optional session_id, message_id, rich_text, reply_to_id, attachments, fork. Returns message_id + delivery immediately (ACCEPTED_BOX, ACCEPTED_TEMPORAL, DUPLICATE, …). It does not wait for the assistant. Retry by sending the same message_id.

  • grokbot_wait — block until the next assistant-ish transcript entry after send, the agent stops (is_running false), or timeout (default 45s from GROKBOT_WAIT_TIMEOUT_SEC, clamp 1–120). Returns {timed_out, entries, is_running}. Honours MCP cancellation. Does not use MCP Tasks. Polls history if the watcher is down.

  • grokbot_status — GetGrokBotSendStatus for message_id

  • grokbot_history — default limit 40, max 100; optional before_seq. Entries use Forge entry_public shape: seq, entry_id, entry_kind, role, text, ts_ms, widget

  • grokbot_interrupt — returns had_active_run

  • grokbot_draft / grokbot_discard_draft / grokbot_react — wrap ChatSession. Draft needs exactly one of email or slack object.

Typical loop: send → wait (or history if you already know after_seq) → if quiet, pending.

Secrets

  • grokbot_list_secrets — names + descriptions only

  • grokbot_put_secret — agent_id, name, value; optional description. Result {ok, name} only. Value is never echoed in logs or the result. Only if the user asked.

  • grokbot_delete_secret — confirm=true required or preview-only

Widgets, handoff, desktop

  • grokbot_pending — widgets + HandoffRequested. Drops stale items older than 6 hours. Optional agent_id / session_id.

  • grokbot_resolve_widget — action: respond | dismiss | form | secret | approve. Requires agent_id + entry_id. Payload: value (respond/secret), values object (form), approved (approve).

  • grokbot_desktop — account sandbox (not per-agent; temporal has none). wake default false. wake=true requires confirm=true (may boot a hibernated box). Returns run_state, connect_url, and websocket header names (not token values). No in-process RFB/noVNC.

  • grokbot_end_handoff — client.end_handoff(agent_id, request_id); optional trigger (default DISMISSED)

Resources (JSON): grokbot://me, grokbot://agents, grokbot://agents/{id}/transcript, grokbot://agents/{id}/pending. Watcher persists Forge-shaped cursors in $GROKBOT_DATA_DIR/grokbot_cursors.json (default ./data, Docker /data). Stdio process death is fine.

Environment

Copy .env.example → .env. Empty values mean “unset / default”.

  • GROKBOT_TOKEN — Cursor JWT if no auth.json

  • GROKBOT_MCP_TOKEN — required for HTTP; Bearer shared secret

  • GROKBOT_MCP_HOST / GROKBOT_MCP_PORT / GROKBOT_MCP_PATH — HTTP bind (defaults 127.0.0.1, 8080, /mcp)

  • GROKBOT_MCP_RATE_PER_MIN — default 60

  • GROKBOT_DATA_DIR — default ./data (Docker /data)

  • GROKBOT_AUTH_FILE — override PKCE credential path (default ~/.config/grokbot-mcp/cursor-auth.json)

  • GROKBOT_WAIT_TIMEOUT_SEC — default 45, clamp 1..120

Security

  • HTTP will not listen without GROKBOT_MCP_TOKEN. Compare Bearer with hmac.compare_digest. /health is the only unauthenticated route.

  • Bind compose to localhost (127.0.0.1:8080). Treat the MCP token like a password.

  • Do not install this server on Grok Bots (inbound vs outbound).

  • Do not log JWT, MCP token, secret value, or desktop header values. Desktop returns header names only; put_secret / widget secret never echo the secret. login does not print the JWT unless --print-token.

  • delete_* without confirm=true is preview-only. Desktop wake=true needs confirm=true. Create/delete/put_secret/wake only when the user asked.

Tests

Offline (FakeClient, no network). Pytest asyncio_mode=auto, --cov=grokbot_mcp --cov-fail-under=70.

ruff check grokbot_mcp tests
pytest -q

Markers: live / live_mutate skip unless env is set. GROKBOT_LIVE_MUTATE=1 creates and sends on the production backend — throwaway agent only. GROKBOT_LIVE=1 is read-mostly (GetMe, list, watch, desktop wake=false).

CI (Forgejo)

Workflows under .forgejo/workflows/ (ci, sonarqube, owasp, container). No .github/workflows (GitHub is a push mirror). Jobs runs-on: docker with LAN checkout of this repo and grokbot-client. sonar.projectKey=kenzim_grokbot-mcp. OWASP DC JSON is not imported into Sonar.

Repository secrets:

  • SONAR_HOST_URL — SonarQube base URL

  • SONAR_TOKEN — analysis token

  • REGISTRY_USER plus REGISTRY_TOKEN or PACKAGE_WRITABLE_TOKEN — push git.stackken.com/kenzim/grokbot-mcp:{sha,latest,tag}

Out of scope: Forge inbound gateway, Electron keychain, usage/billing RPCs, MCP Tasks, in-process VNC, PyPI publish.

Available Tools

21 tools
grokbot_create_agentC

Create a Grok Bot agent. Only when the user asked. harness is box|temporal.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
titleNo
harnessNobox or temporalbox
descriptionNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and falls well short. It does not say whether names must be unique, whether creation is idempotent, what permissions are needed, or what happens on failure — all relevant for a mutation tool creating a persistent agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, with the action stated first. But the remaining two fragments are cryptic ('Only when the user asked') and partially redundant ('harness is box|temporal' duplicating the schema), so brevity here reflects under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter creation tool with no annotations and no output schema, the description is too thin. An agent gets no sense of side effects, required inputs beyond the schema's required field, or what a successful creation returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only harness is documented), so the description should compensate but does not. Its only parameter content, 'harness is box|temporal', merely restates the schema's own description and adds nothing about the undocumented name, title, and description parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a Grok Bot agent'), which is unambiguous against the create/update/delete/list sibling set. It stops short of explicitly differentiating itself from grokbot_update_agent or describing scope, but the core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Only when the user asked' is an explicit usage condition, functioning as a guardrail against proactive agent creation. However, it is a terse fragment that gives no guidance on prerequisites or when an alternative sibling (e.g., update) would be correct, leaving usage largely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_delete_agentA
Destructive

Delete an agent. Without confirm=true this is preview-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoMust be true to actually delete
agent_idYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, so the safety profile is known. The description adds real behavioral value by disclosing that without confirm=true the call is preview-only, i.e. a no-op dry run, which the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the destructive action stated first and the safety condition immediately after; nothing is wasted. Very brief, though, which limits how much it can convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter delete tool with no output schema, the description adequately covers the essential behavior (preview vs. actual delete) on top of annotations that already mark it destructive. Missing only usage context, which is a minor gap here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% – the 'confirm' field is already documented as 'Must be true to actually delete' in the schema. The description's preview-only sentence reinforces that meaning but adds no new syntax or semantics, and agent_id is undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete an agent') and clearly belongs to the delete family, distinguishable from create/update/get siblings. It does not explicitly name a sibling, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance or alternatives; it only describes what the call does and its preview behavior. An agent must infer the context of use from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_delete_secretB
Destructive

Delete a secret. Without confirm=true this is preview-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
confirmNoMust be true to actually delete
agent_idYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The destructiveHint=true annotation signals a mutating operation, but the description adds the critical nuance that the default behavior is a non-destructive preview, not an actual delete. That two-phase semantics is real behavioral context beyond the annotation, though nothing is said about reversibility or downstream effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and the important safety caveat immediately after. No filler, though the phrasing is terse enough that it borders on underspecified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers the confirm/preview mechanism but leaves the two required parameters and any error or edge-case behavior unexplained. It is minimally workable for an agent that already understands the secret-name convention.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: agent_id and name are undocumented in both schema and description, and the description adds no meaning for those required params. The confirm semantics it hints at are already fully specified in the schema, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a secret'), which cleanly separates it from siblings like grokbot_put_secret and grokbot_list_secrets. It does not explicitly name an alternative, but the verb+resource pair is unambiguous on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys the key usage rule: without confirm=true the call is preview-only, which is the main decision an agent must make. However, it offers no guidance on when to prefer this over alternatives or on any prerequisites beyond the confirm flag.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_desktopB
Destructive

Account-level sandbox VNC coordinates. wake=true requires confirm=true. Returns run_state, connect_url, and websocket header names (not token values).

ParametersJSON Schema
NameRequiredDescriptionDefault
wakeNoWake the sandbox. Requires confirm=true.
confirmNoRequired when wake=true

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The destructiveHint annotation already flags the safety profile, so the bar is lower. The description usefully adds the confirm requirement and names return fields, noting websocket headers are not token values, but it does not explain what is destroyed or what the destructive action actually does.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler, and it front-loads the resource domain before the wake/confirm rule and return details. Every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only two parameters and no output schema, the description partially compensates by naming return fields. However, it leaves the core purpose and when-to-use context underspecified for an agent choosing among many grokbot tools, so it is only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters are fully documented in the schema with the same wake/confirm dependency stated in the description. The description adds no syntax, format, or side-effect meaning beyond the schema, so this hits the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Account-level sandbox VNC coordinates' names a resource domain but uses no clear action verb, leaving the actual operation ambiguous. It is distinguishable from account-management siblings, but an agent cannot confidently tell whether this retrieves, creates, or wakes a sandbox.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives only the conditional prerequisite 'wake=true requires confirm=true' and no guidance on when to use this tool versus the many grokbot siblings. There is no context, exclusion, or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_discard_draftC

Discard a pending email or Slack draft for an entry.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
entry_idYes
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses only that the target must be a 'pending' draft; it says nothing about irreversibility, permission/ownership requirements, error behavior, or whether the entry or session is affected. For a destructive mutation tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and target come first. It is terse to the point of under-specification, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and two required parameters undocumented, the description leaves the agent guessing about what agent_id and entry_id must contain and what the tool returns. For a state-mutating tool this is not sufficient context to invoke confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (just session_id is documented; agent_id and entry_id are bare strings). The description's phrase 'for an entry' loosely gestures at entry_id but adds no format, type, or sourcing guidance to compensate for the undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Discard') and resource ('a pending email or Slack draft for an entry'), so the core action is unambiguous. It does not distinguish itself from the sibling grokbot_draft (create) or explain the 'entry' relationship, but the verb+resource pairing is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as grokbot_draft, nor any preconditions. The only implicit cue is the word 'pending', which hints the draft must still be open, but nothing tells the agent when discarding is appropriate or what happens if it is not pending.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_draftC

Send an email or Slack draft. Exactly one of email or slack is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoEmail draft {to, cc, from, subject, body}
slackNoSlack draft {target, body, workspace, thread}
agent_idYes
entry_idYes
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. For a tool whose name says 'draft' but whose text says 'send', the agent cannot tell whether a real message is delivered immediately, whether the draft remains pending, what permissions are needed, or what side effects occur — all material for a mutation with zero structured coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and immediately followed by the constraint; there is no filler. It is a little terse for a tool with two nested object shapes, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested objects, five parameters, no annotations, and no output schema, the description omits what happens on success, whether the draft is delivered or queued, and how agent_id/entry_id are consumed. It is not complete enough for an agent to invoke confidently without inspecting the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% and the mutual-exclusion rule between email and slack is real added meaning not expressed in the schema (which only marks agent_id/entry_id as required). But agent_id, entry_id, and session_id remain undocumented in both places, so the description only partly compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource pair is stated (send a draft), but 'draft' conflicts with 'send' — it is unclear whether this dispatches a queued draft or creates a pending one, especially with the sibling grokbot_discard_draft implying drafts are artifacts to be resolved later. No differentiation from the sibling grokbot_send is given, leaving the agent to guess which one to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Exactly one of email or slack is required' is a genuine invocation rule that constrains how to call it. However, there is no when-to-use guidance relative to grokbot_send or grokbot_discard_draft, and no prerequisites (e.g. whether an agent_id must be running) are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_end_handoffB

Clear a stuck desktop handoff (captcha / login) without waking the box.

ParametersJSON Schema
NameRequiredDescriptionDefault
triggerNoHand-back trigger; default DISMISSEDDISMISSED
agent_idYes
request_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one important trait: the handoff is cleared 'without waking the box,' i.e. non-disruptively. It does not say what happens to the agent/request state, whether the call is idempotent, what error occurs if the handoff is not actually stuck, or what is returned, so significant gaps remain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the core action and the key constraint are both stated immediately. Nothing needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small 3-parameter tool with no output schema, the description covers purpose plus one behavioral nuance, which is close to adequate. It still leaves the agent guessing about parameter meanings, failure behavior, and how it relates to sibling recovery tools, so it is minimum-viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only 'trigger' is documented in the schema) and the description adds no parameter meaning at all, not even for agent_id or request_id. Although those identifiers are largely self-evident, the description fails to compensate for the low schema coverage as required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Clear') and a specific resource ('stuck desktop handoff / captcha / login'), and adds a meaningful scope qualifier ('without waking the box'). An agent can understand the operation without opening the schema, though it stops short of explicitly contrasting this with near-neighbors like grokbot_resolve_widget or grokbot_interrupt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'stuck desktop handoff (captcha / login)' implies the triggering condition, so an agent can infer when it applies. However, there is no explicit when-to-use guidance and no routing to or away from the many related siblings (grokbot_desktop, grokbot_interrupt, grokbot_resolve_widget).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_get_agentB
Read-only

Get one agent plus todos, sessions, and runtime capabilities.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent id

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation already establishes this is a safe read. The description adds useful context by disclosing that the response bundles related todo, session, and capability data, but says nothing about authorization requirements, rate limits, or behavior for unknown/invalid agent ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that front-loads the core action and resource, then tacks on the payload contents. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description partially compensates by naming what is returned (todos, sessions, runtime capabilities), which is the main thing an agent needs to know. It stops short of describing error behavior or the shape of those sub-objects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (agent_id is documented), so the schema carries parameter meaning and the baseline is 3. The description adds no format or constraint detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (one agent) and enumerates the payload (todos, sessions, runtime capabilities), which distinguishes it from grokbot_list_agents and the mutating create/update/delete siblings. It does not name a sibling explicitly, keeping it just under a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus grokbot_list_agents or grokbot_status, and no prerequisites or exclusions are stated. The agent must infer usage from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_historyC
Read-only

Recent transcript entries (entry_public shape). Default limit 40, max 100.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
agent_idYes
before_seqNo
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description usefully adds a hard cap (max 100) that the schema does not express, plus the default. However, it says nothing about ordering, pagination via before_seq, or what a transcript entry contains beyond the shape name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse fragments with the return shape front-loaded and the limit constraint following. Nothing is wasted, though the telegraphic style trades clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should carry more of the return and pagination story, but it only names the entry_public shape. before_seq's role in cursor-style paging, agent_id semantics, and when to reach for this transcript are all missing for a 4-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (session_id alone is documented). The description compensates partially by clarifying the limit default and max 100, which the schema omits, but agent_id and especially before_seq remain entirely unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It identifies the resource (transcript entries) and notes the return shape (entry_public), but there is no explicit verb and no differentiation from siblings like grokbot_pending or grokbot_status that also surface agent activity. An agent can roughly infer 'read history' but not clearly where this fits among the 20 sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no conditions for choosing it over grokbot_pending or grokbot_status, and no mention of prerequisites. The word 'Recent' is the only hint of intended usage, which the agent must infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_interruptC

Interrupt the agent's current run. Returns had_active_run.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNouser
agent_idYes
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return value (had_active_run) but says nothing about whether the interrupt is graceful or abrupt, what happens to pending widgets/drafts, required permissions, or whether it's reversible. For a mutation-style control tool with zero annotation coverage this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two telegraphic sentences that front-load the action and include the return signal with zero filler. Appropriately sized for a single-action tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 3-parameter mutation tool with no annotations, no output schema, and only one undocumented-params return hint. Missing behavioral detail and parameter semantics that an agent needs to call it correctly, even given the simple surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: session_id is documented but agent_id and reason are not, and the description adds nothing about any of them. It does not explain the default 'user' value for reason or how target scope works. With low coverage and no compensating text, the gap is unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (interrupt) plus resource (the agent's current run), which is distinct from siblings like grokbot_send or grokbot_wait. It also states the return signal (had_active_run), so the agent knows what to expect. It stops short of explicitly differentiating from other lifecycle tools such as grokbot_end_handoff.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to interrupt versus alternatives (end_handoff, wait, resolve_widget) or when not to. No prerequisites or context about what state the agent must be in. The agent must infer all routing from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_list_agentsA
Read-only

List Grok Bot cloud agents (id, agent_id, name, description, harness).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation already declares this is a safe, non-mutating read. The description adds the returned field list (id, agent_id, name, description, harness), which is useful output-shape context, but says nothing about pagination, result limits, or filtering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the verb and resource front-loaded and zero filler. Every token is meaningful and nothing is duplicated from structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description usefully compensates by naming the returned fields, so an agent knows what comes back without inspecting a response. The only remaining gap is pagination or result-size behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify beyond the schema. The baseline of 4 applies for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List Grok Bot cloud agents') and even enumerates the fields returned, so an agent immediately knows what the tool produces. It does not explicitly differentiate itself from siblings like grokbot_get_agent, though the list/get contrast is strongly implied by naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus grokbot_get_agent or the other agent-management siblings. The reader must infer from the name alone that this is the enumeration variant; no conditions, prerequisites, or alternatives are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_list_secretsA
Read-only

List secret names and descriptions for an agent. Values are never returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint already declares this as a non-mutating read, so the bar is lowered, yet the description still adds the genuinely useful fact that secret values are never returned. It leaves out whether descriptions can be truncated or what permissions are needed, but the key security behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the security-relevant constraint (values are never returned) is placed prominently right after the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list with an annotation covering safety and no output schema, the description provides enough to call it correctly. Only minor gaps remain, such as permissions or result ordering/count.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter agent_id is undocumented in the schema, so the description carries the burden. It does not explain the parameter at all, but the name is largely self-explanatory for an agent-scoped list, keeping this at the minimum-viable level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (List) plus resource (secret names and descriptions) scoped to an agent. It clearly separates itself from grokbot_put_secret and grokbot_delete_secret by being the read/list counterpart, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives like grokbot_put_secret or grokbot_delete_secret. The intent is inferable from the verb but nothing is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_pendingB
Read-only

List pending widgets and HandoffRequested events. Drops items older than 6 hours.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNo
session_idNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=true, openWorldHint=true), and the description adds genuinely useful behavior beyond that: 'Drops items older than 6 hours' tells the agent results are bounded in time and may silently omit stale work. It stops short of a 5 because nothing is said about ordering, volume, or the shape of a returned item.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler, and the identification of what is listed comes before the retention caveat. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no output schema, the description covers what is returned at a high level plus the 6-hour retention boundary, which is the key operational fact. It remains incomplete on the two filter parameters and on what a 'widget' or 'HandoffRequested event' actually contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions agent_id or session_id, so it contributes no meaning for either filter. With two undocumented optional parameters, the agent cannot tell whether passing agent_id narrows results, changes scope, or is ignored — the description should have compensated and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and two concrete resources ('pending widgets and HandoffRequested events'), so an agent can tell it apart from sibling tools like grokbot_resolve_widget or grokbot_end_handoff. It does not explicitly name or contrast with the nearest alternatives (e.g. grokbot_status, grokbot_wait), which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no indication of how it relates to related siblings such as grokbot_status, grokbot_wait, or grokbot_resolve_widget. An agent must infer from the name alone that this is the polling/inspection call for outstanding work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_put_secretB
Destructive

Create or replace a secret. The value is never echoed in the result or logs.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
valueYes
agent_idYes
descriptionNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation destructiveHint=true already signals that this operation is destructive (replacing existing secrets). The description adds valuable context that the value is never echoed in results or logs, which is relevant for security-sensitive operations. However, it doesn't disclose other behaviors like permission requirements, idempotency, or what happens to the previous value on replacement beyond implying it's overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two short sentences that are front-loaded with the core action and a key behavioral note. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four parameters with zero schema descriptions and no output schema, the description is incomplete. It fails to explain required parameters, their purpose, or any side effects beyond the value not being echoed. For a secret-management tool with destructive implications, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning none of the four parameters are documented in the schema. The description mentions no parameters at all, leaving 'agent_id', 'name', 'value', and 'description' entirely unexplained. With low coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create or replace a secret.' It's clear what the tool does. However, it doesn't differentiate from sibling tools like grokbot_delete_secret or grokbot_list_secrets beyond the obvious action difference, which is somewhat self-evident from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool vs alternatives. While the action is clear, it doesn't specify prerequisites (e.g., requiring agent_id), when to use this over other secret-handling tools, or any conditional logic. No explicit when/when-not/alternatives are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_reactC

React to a transcript entry with an emoji.

ParametersJSON Schema
NameRequiredDescriptionDefault
emojiYes
agent_idYes
entry_idYes
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full disclosure burden. It does not say whether reactions are additive or replace existing ones, whether they can be removed, whether the acting agent must be in the session, or what happens if the same emoji is applied twice – meaningful gaps for a state-mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, stating the action and target immediately. It is arguably too sparse rather than too long, but the sentence itself is well-formed and waste-free.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutating tool with no annotations, no output schema, and 25% parameter coverage needs far more than one line to be safely invokable. The description leaves the emoji encoding, entry targeting, and result behavior entirely undefined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%: just session_id is documented ('Chat session id; default empty MAIN transcript'). The description adds nothing about emoji format (unicode vs shortcode), what entry_id references, or how agent_id is scoped, leaving three required parameters unexplained in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (react) plus resource (transcript entry) and the instrument (emoji), so the basic action is unmistakable. It stops short of differentiating itself from siblings like grokbot_send or grokbot_history, which also touch transcript content, so it does not reach 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no preconditions, and never names an alternative tool for related operations such as reading or sending to the transcript. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_resolve_widgetD

Resolve a widget: respond, dismiss, form, secret, or approve.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
valueNo
actionYes
valuesNo
agent_idYes
approvedNo
entry_idYes
request_idNo
session_idNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, openWorldHint=true, and idempotentHint=false, but the description adds zero behavioral context: no statement of side effects, permission requirements, whether actions are reversible, or what each action does to server state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is short and front-loaded, but it is under-specification rather than conciseness — it spends its few words restating the enum instead of conveying the one piece of information an agent actually needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with 9 parameters, 0% schema coverage, nested objects, no output schema, and openWorldHint=true demands far more than 11 words. The description leaves the required-field/action coupling, return behavior, and error handling completely undefined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, and the description only echoes the enum values already present in the schema. Parameters like value, values, approved, kind, request_id, and session_id are entirely unexplained in both places, leaving the agent no way to know which fields pair with which action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially the tool name stripped of its prefix ('Resolve a widget') plus a verbatim restatement of the schema's action enum. It never defines what a 'widget' is or what resolving one accomplishes, so an agent cannot distinguish its purpose from siblings like grokbot_pending beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite conditions, and no reference to any alternative tool. Nothing tells the agent when resolving a widget is appropriate versus calling grokbot_pending to discover one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_sendB
Idempotent

Send a user message (fire-and-forget). Returns message_id and delivery. Retry with the same message_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
forkNo
textYes
agent_idYes
rich_textNo
message_idNoReuse to retry an identical send
session_idNoChat session id; default empty MAIN transcript
attachmentsNo
reply_to_idNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation only declares idempotentHint=true, and the description meaningfully extends it: it discloses that the operation is fire-and-forget (no blocking response), names the return values (message_id, delivery), and explains the retry contract. Auth and attachment-handling behavior remain unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse sentences, with the core action and its async nature front-loaded and no filler. It may be slightly over-compressed for an 8-parameter tool, but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation tool with no annotations beyond idempotency and no output schema, the description leaves critical gaps: fork semantics, attachment formats, reply threading, and the text/rich_text relationship are all unexplained. It covers only the send action and retry path.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25%, so most of the 8 parameters are undocumented in both schema and description. The description clarifies the retry use of message_id but says nothing about agent_id, text vs rich_text, fork, attachments, or reply_to_id, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Send a user message') and adds a key behavioral qualifier ('fire-and-forget'). However, it does not differentiate from close siblings like grokbot_draft or grokbot_react, so an agent cannot fully disambiguate from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies an asynchronous sending context via 'fire-and-forget' and instructs on retry, but gives no explicit when-to-use or when-not-to-use guidance relative to alternatives such as grokbot_draft or grokbot_wait. There are many siblings that could plausibly overlap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_statusC
Read-only

GetGrokBotSendStatus for a previously sent message_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
message_idYes
session_idNoChat session id; default empty MAIN transcript

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true, the safety profile is already covered, and the description is consistent with it (a status query). It adds the precondition that the message must have been sent already, which is real context, but says nothing about what the returned status contains or how fresh it is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded and free of padding, which is structurally fine. Its brevity comes at the cost of under-specification rather than crispness, so it is adequate but not exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, two required params, and 33% schema coverage, the description carries a heavy burden it does not meet. It never explains what 'send status' actually returns or how the agent should use the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: only session_id is documented in the schema. The description nominally covers message_id ('previously sent message_id') but leaves agent_id completely unexplained and adds no format or meaning beyond the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (send status) and the anchor (a previously sent message_id), so the core operation is discernible. However, 'GetGrokBotSendStatus' is essentially the internal function name restated, and nothing distinguishes it from status-adjacent siblings like grokbot_wait, grokbot_pending, or grokbot_history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus grokbot_wait (wait for completion) or grokbot_pending (list outstanding items). The 'previously sent' precondition hints at usage but no alternative or exclusion is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_update_agentC

Update agent name, description, title, avatar_shape, or avatar_color.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
titleNo
agent_idYes
descriptionNo
avatar_colorNo
avatar_shapeNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the full behavioral burden, yet it omits critical mutation semantics: whether omitted fields are preserved (partial patch) or cleared, whether existing values are overwritten irreversibly, required permissions, and error behavior. 'Update' alone is not enough for a write tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that lists the updatable fields with zero filler or redundancy. It is efficient, though the extreme brevity leaves the definition under-specified rather than optimally scoped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no annotations, no output schema, and no parameter descriptions, this definition is too thin: it never explains partial-update behavior, permissions, or the required agent_id. An agent could call it but could easily misunderstand what happens to unlisted fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only restates schema property keys (name, description, title, avatar_shape, avatar_color) and omits the sole required parameter (agent_id) entirely. No format, length, or allowed-value guidance is added (e.g., valid avatar_shape values), so meaning beyond the schema is minimal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Update') and resource ('agent'), and enumerates the mutable fields, so an agent immediately knows this is the mutation counterpart to grokbot_get_agent / grokbot_create_agent. It stops short of explicitly naming those siblings or stating scope, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no mention of prerequisites (agent must exist), and no routing to alternatives such as grokbot_create_agent or grokbot_get_agent. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_waitB
Read-only

Block until the next assistant-ish transcript entry after send, the agent stops running, or timeout. Does not use MCP Tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
agent_idYes
after_seqNo
session_idNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint, so the description usefully adds that this call blocks and that it ends on a transcript entry, agent stop, or timeout. The note 'Does not use MCP Tasks' is a meaningful implementation detail an agent may need. It still does not say what is returned or how timeout interacts with the block.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, with the blocking behavior and termination conditions front-loaded. Minimal waste, though 'assistant-ish' is imprecise vocabulary that costs a little clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should explain the return value, which it does not. It adequately covers the blocking lifecycle but leaves parameters and results undocumented for a tool whose entire purpose is to suspend execution.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description carries the full burden and largely fails it. 'after send' weakly connects after_seq to a prior send, but timeout bounds (1-120s), session_id, and agent_id semantics are unexplained in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Block until...') and enumerates the three conditions that end the block, which tells an agent exactly what the call accomplishes. The 'assistant-ish transcript entry' phrasing is jargon and the relationship to grokbot_send is only implied, but the core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after send' hints that this is used following grokbot_send to wait for a response, and the termination conditions imply the intended call pattern. However, there is no explicit guidance to prefer this over polling grokbot_status or reading grokbot_history, so the alternative-selection logic is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grokbot_whoamiA
Read-only

Current Cursor account email, user_id, and grokbot SCHEMA_VERSION.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered by structured data. The description usefully discloses the three concrete values returned (email, user_id, SCHEMA_VERSION), which is real added context, but it says nothing about auth requirements, failure modes, or whether the identity is derived from a session or credential.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact clause with zero filler; the returned fields are listed up front and nothing is repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must enumerate the return values, and it does so completely for a zero-parameter read tool. Nothing essential to calling it correctly is missing, though a brief note on when the identity is useful would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema coverage is 100%, so there is nothing for the description to clarify. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact resource it returns — the current Cursor account email, user_id, and grokbot SCHEMA_VERSION — which is specific and clearly distinguishable from every sibling (all of which act on agents, secrets, or handoffs). It lacks an explicit verb like 'Return', but the noun phrase is unambiguous about what the tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool, no prerequisites, and no reference to alternatives. The identity-check use case is only inferable from the name 'whoami' and the returned fields, so the agent gets no explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv0.1.0
    • First observedgrokbot_create_agent
    • First observedgrokbot_delete_agent
    • First observedgrokbot_delete_secret
    • First observedgrokbot_desktop
    • First observedgrokbot_discard_draft
    • First observedgrokbot_draft
    • First observedgrokbot_end_handoff
    • First observedgrokbot_get_agent
    • First observedgrokbot_history
    • First observedgrokbot_interrupt
    • First observedgrokbot_list_agents
    • First observedgrokbot_list_secrets
    • First observedgrokbot_pending
    • First observedgrokbot_put_secret
    • First observedgrokbot_react
    • First observedgrokbot_resolve_widget
    • First observedgrokbot_send
    • First observedgrokbot_status
    • First observedgrokbot_update_agent
    • First observedgrokbot_wait
    • First observedgrokbot_whoami

TDQS

C2.7/5.0

Scored across 21 tools

Disambiguation4/5

Most tools target distinct resources or actions, and the agent and secret CRUD groups are clearly separated. Some overlap exists among send/status/wait/draft and between pending/history, but descriptions provide enough context to choose correctly.

Naming Consistency3/5

All tool names share the grokbot_ prefix and snake_case, which is consistent. However, the set mixes verb_noun patterns (list_agents, create_agent, resolve_widget) with bare nouns/verbs (status, history, wait, desktop, whoami), making the convention less predictable.

Tool Count3/5

21 tools is on the heavy side for a single server, even though the capabilities span agents, messaging, widgets, secrets, desktop, and handoff. Each area is distinct, but some niche operations could likely be consolidated.

Completeness4/5

Core CRUD for agents and secrets, messaging, widget handling, desktop access, and handoff are all covered. Minor gaps remain, such as todo/session management, draft listing/approval, and transcript search, but the main workflows are supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables ChatGPT Desktop and Codex to delegate substantial work to a locally installed Grok Build agent. Supports consultations, background builder/tester jobs, cancellation, session discovery, and transcript export.
    7
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to manage Grok Bots through MCP: create, list, search, and delete bots, send and receive messages, read conversation transcripts, search message history, and check usage.
    11
    2
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP clients like Codex and Claude Code to start or continue conversations with Grok Build, track progress, read completed answers, and cancel running turns using an existing Grok login.
    1
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP-compatible agents like ChatGPT and Claude to securely access a Grok Bot workstation's folder-scoped files, shell, and Git tools over HTTPS+OAuth or local stdio, and to orchestrate specialized Grok Bot agents via a message bridge.
    1
    MIT