Grok Gadgets Gateway
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Grok Gadgets Gatewaylist my gadgets and turn the simulator LED green"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Grok Gadgets gateway
An MCP gateway and configurable software simulator for gadgets targeting Grok. It discovers device capabilities, routes commands, and reports state and events.
Experimental alpha. Local MCP simulation and software transports are tested. Native Grok invocation receipts, mobile clients, physical C124 operation, and a reviewed authenticated route from a cloud Bot to local devices remain open.
flowchart LR
C["Local MCP client"] --> G["Gateway"]
G --> S["Software C124 simulator"]
G --> D["Linux / ESP32 device application"]
B["Grok Bot: invocation evidence pending"] -.-> GSolid paths describe local software interfaces. An execution report is a device acknowledgement; physical effects require a separate observation. The simulator has no physical effects and makes no Grok API calls.
Choose a first step
Try without hardware: run the installed MCP demonstration below.
Customize: use the simulator guide and a separate
my-light.json.Connect your application: read local operation and the versioned protocol.
Contribute: use CONTRIBUTING, support, and GitHub Issues.
Related MCP server: MCP-Edge
Run the installed demonstration
Requirements: uv, Python 3.11, and the gateway wheel. Native Apple Silicon Python 3.11.15 is the tested baseline. Installation may need network access; the simulator needs no account, API key, hardware, or open device listener.
Build the wheel from this source checkout; package releases are not yet published. Use
uv sync --locked and uv build. Place
grok_gadgets_gateway-0.1.0a1-py3-none-any.whl in an otherwise empty working folder,
open a terminal there, and run:
uv venv --python 3.11 --seed .venv
.venv/bin/python -m pip install ./grok_gadgets_gateway-0.1.0a1-py3-none-any.whl
.venv/bin/python -m grok_gadgets_gateway.demoThe official local MCP client launches a subprocess gateway and asserts discovery, green RGB, state readback, safe retry, invalid-command rejection, simulated button edges, offline failure, and reconnect. The final JSON includes:
{"led_status": "executed", "simulated": true, "physical_verified": false, "grok_verified": false}That is a subset of the report, not a native Grok receipt. The demo deliberately enables test controls in its child simulator only.
An ordinary local MCP client launches:
.venv/bin/grok-gadgets-gateway --simulatorUse the absolute executable path in the client's configuration. The client supervises
the process, which speaks MCP on stdio and waits for requests; it is not an interactive
terminal. Ordinary tools are gadgets_list_devices, gadgets_get_state,
gadgets_command, gadgets_command_status, gadgets_read_events, and
gadgets_diagnostics. --test-controls is a separate explicit opt-in.
Compatibility and evidence
Package 0.1.0a1 and protocol 0.1.0 are separate version identifiers.
Declared Python >=3.11 support is not evidence for every interpreter or platform.
Path | Evidence | Remaining limit |
macOS arm64, CPython 3.11.15 | Source checks and fresh installed-wheel MCP demo | Independent human reproduction |
Linux aarch64, CPython 3.11.17 container | Gateway domain/TCP used in recorded Linux SDK acceptance | Standalone gateway MCP/platform matrix |
Simulator | Configured identity, RGB, delay, offline/reconnect | No physical device |
Device TCP / USB framing | Authenticated loopback and software pseudo-terminal checks | Physical cable, board, and OS permissions |
Windows / Intel Mac | Not verified | Clean installation and runtime tests |
Grok / mobile / C124 | Native invocation/mobile/physical evidence pending | Supported client route and observed hardware acceptance |
Launch verification records the documented quick start. Simulator evidence and historical alpha checks retain their actual scopes.
Architecture and boundaries
This repository owns the gateway, simulator, and canonical schemas. Device libraries belong in the Linux SDK and ESP32 SDK. Shared architecture, roadmap, and policies live in the hub. The demo needs no sibling checkout.
Device TCP binds to loopback only and requires per-device credentials outside Git. It is separate from MCP stdio. A cloud Bot cannot execute a path on your computer. Remote HTTPS/OAuth connectivity is not implemented by these transports. Read architecture, security, and release preparation.
Troubleshooting and support
Symptom | Next step |
Server waits silently | Launch through an MCP client, or run the demonstration instead. |
No simulated device | Include |
Config rejected | Use strict v1 JSON and the documented bounds. |
Dependency installation fails | Use the selected Python 3.11 environment above; another system interpreter/architecture is not the verified baseline. |
Device unavailable | Inspect state; reconnect explicitly in a test session or diagnose the agent. |
Unconfirmed / timed out | Inspect state and recover safely; never invent a new retry ID for an uncertain physical action. |
Use SUPPORT, SECURITY, and CODE_OF_CONDUCT. Do not post credentials, household state, private event bodies, or raw account captures in issues.
Original code is Apache-2.0; retain NOTICE and installed dependency licenses. This independent project is exclusively for Grok and is not affiliated with xAI.
History note
Pre-publication commit dates were reconstructed across 29 September–5 October 2026 at the owner’s request. Verification records retain their actual execution dates. See the history and privacy record.
Available Tools
6 toolsgadgets_commandB
Request a device action with a stable retry ID; distinguish accepted from executed.
| Name | Required | Description | Default |
|---|---|---|---|
| arguments | Yes | ||
| device_id | Yes | ||
| capability | Yes | ||
| command_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two meaningful traits: the call is asynchronous (accepted ≠ executed) and command_id acts as a stable retry/idempotency key. It omits auth requirements, whether the request can fail silently, and any rate or capability constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the async/retry semantics are packed in tightly. It is arguably too compressed, since the clauses are telegraphic rather than explanatory, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-required-parameter mutation tool with a free-form nested arguments object, no annotations, and no output schema, the description is far too thin. It never explains what a capability is, how arguments are shaped, or what the accepted response contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 required parameters, and the description only illuminates command_id ('stable retry ID'). device_id, capability, and especially the free-form nested 'arguments' object (additionalProperties: true) are left with no guidance on format or capability-dependent content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Request a device action' with a retry identifier, which is clearly distinct from the read-oriented siblings (gadgets_get_state, gadgets_read_events). It does not, however, explicitly contrast itself with gadgets_command_status, which covers the closest adjacent concern.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'distinguish accepted from executed' hints that execution must be verified elsewhere (presumably gadgets_command_status) and that command_id enables safe retries. No explicit when-to-use, prerequisites, or named alternative is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gadgets_command_statusC
Read execution acknowledgement; timed_out/unconfirmed require human recovery.
| Name | Required | Description | Default |
|---|---|---|---|
| command_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely useful behavioral trait: timed_out and unconfirmed states require human recovery rather than automated retry. However, it omits whether the call is read-only, whether polling is expected or the command may still be in flight, and what the other possible states are.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with the read operation front-loaded and the recovery caveat second; no filler. It is slightly too terse and jargon-dependent to be fully self-explanatory, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented command_id, the description leaves key gaps: it does not say where command_id comes from, what fields the acknowledgement contains, or what the normal (non-error) success outcome looks like. The recovery note is the only substantive disclosure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter command_id has 0% schema description coverage, and the description adds nothing about it — not the expected format, nor that it is the identifier returned by gadgets_command. Since there is one undocumented required parameter, the description should have compensated and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Read' plus the resource 'execution acknowledgement' communicates that this retrieves the result/status of a command, but 'execution acknowledgement' is jargon that only loosely maps to 'check the status of a previously submitted command'. It hints at a relationship with gadgets_command but never states it, so the agent must infer the distinction from siblings like gadgets_get_state and gadgets_read_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no mention of when this should be called relative to gadgets_command, and no alternative named. The clause about timed_out/unconfirmed describes outcome handling, not selection criteria for this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gadgets_diagnosticsA
Get an allowlisted support report without tokens, state, arguments or raw errors.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully discloses the output profile: the report is allowlisted and excludes tokens, state, arguments, and raw errors, which tells the agent the result is safe to surface and intentionally lossy. It omits access requirements, whether it is strictly read-only, and any rate/size behavior, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every clause (allowlisted, and the four excluded data classes) carries distinct information the agent cannot get from the empty schema or absent annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 0-param tool with no output schema, the description supplies the key missing context: what kind of report comes back and that secrets/raw internals are stripped. It does not indicate the report's structure or sections, but the essential safety-relevant framing is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics for the description to compensate for. Baseline for a 0-param tool applies; no misinterpretation is possible.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get an allowlisted support report.' It further characterizes the artifact (a redacted support report), which implicitly distinguishes it from siblings like gadgets_get_state and gadgets_read_events. It does not name a sibling or contrast scope explicitly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context, no prerequisites, and no routing to alternatives among the five sibling gadgets_* tools. An agent can infer this is a diagnostic/support-export call from the name, but nothing in the description says when to reach for it versus gadgets_get_state or gadgets_read_events.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gadgets_get_stateB
Read reported state and freshness; physical effect is not verified.
| Name | Required | Description | Default |
|---|---|---|---|
| device_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose one genuinely valuable and non-obvious trait — that the value is *reported* state and 'physical effect is not verified' — warning the agent not to treat it as confirmed actuation. However, it omits permissions, error behavior, and how 'freshness' is expressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the core read semantic comes first and the caveat second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description adequately signals the return shape (state plus freshness) and the key reliability caveat. It could say more about error conditions or required device state, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single device_id parameter, but the parameter is self-describing by name and type. The description adds no format, scope, or ID-source guidance beyond the schema, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Read reported state and freshness,' which is concrete about what is retrieved from a device. It implicitly separates itself from siblings like gadgets_read_events or gadgets_diagnostics by naming the state/freshness pair, but never explicitly names an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus gadgets_command_status, gadgets_read_events, or gadgets_diagnostics. The only usage signal is the implied 'read state,' leaving the agent to guess which of several read-style siblings to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gadgets_list_devicesB
Discover devices, capability schemas, availability and simulation labels.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It hints at returned content (capability schemas, availability, simulation labels) but never states that the operation is non-mutating, whether device discovery is scoped or filtered, or any auth/rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is appropriately sized, though the compressed noun list of return contents reads slightly like a spec fragment rather than guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and no parameters, the description is the only source of information, and it partially compensates by listing what is returned. It still omits what an agent should do with the results and how this discovery step relates to the other gadgets_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline case where no parameter explanation is needed. The description legitimately adds no parameter detail because none exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ("Discover") and resource ("devices") and even enumerates what is discovered (capability schemas, availability, simulation labels). It is clear what the tool does, but it never distinguishes itself from siblings like gadgets_get_state or gadgets_diagnostics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus the five sibling tools, nor any stated prerequisite or ordering (e.g., call this first to obtain device IDs before gadgets_command). Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gadgets_read_eventsC
Read bounded ordered event history; null cursor recovers explicit history loss.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| device_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden, and it does disclose two non-obvious traits: results are bounded and ordered, and a null cursor signals explicit history loss. It still omits pagination semantics (how to advance the cursor), permission requirements, and what the returned events contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core purpose front-loaded and no filler. It is arguably too terse for the domain, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cursor-paginated read tool with three undocumented parameters, no annotations, and no output schema, the definition leaves too much unspecified. An agent cannot determine how to page, what device_id does, or what happens when the history-loss path triggers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so all three parameters (limit, cursor, device_id) are undocumented in the schema. The description partially compensates by explaining the null-cursor case and implying bounding via 'limit', but device_id and cursor advancement are never addressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Read ... event history') with two qualifiers ('bounded', 'ordered'), which separates it reasonably well from siblings like gadgets_get_state and gadgets_diagnostics. It is not a tautology, but it never says what an 'event' is or how this differs concretely from reading state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives among the five sibling tools. The only usage-adjacent statement is the cursor=null behavior, which is a mechanism rather than a selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
gadgets_command - First observed
gadgets_command_status - First observed
gadgets_diagnostics - First observed
gadgets_get_state - First observed
gadgets_list_devices - First observed
gadgets_read_events
TDQS
Scored across 6 tools
Each tool targets a distinct stage of the device interaction lifecycle: discovery, state read, command dispatch, command acknowledgement, event history, and diagnostics. The descriptions explicitly separate reported state from execution acknowledgement and command request from command status.
All names use a consistent snake_case pattern with the gadgets_ prefix and a predictable verb/noun or action-noun structure. The slight noun form in gadgets_diagnostics does not break the overall convention.
Six tools is well-scoped for a device gateway: it covers discovery, inspection, control, execution tracking, history, and support without redundant or missing categories. Every tool has a clear role.
The set covers the core lifecycle from discovery through command acknowledgement and event history, with diagnostics for support. A minor gap is the absence of an explicit command cancellation or recovery operation, though status descriptions note human recovery for timed-out/unconfirmed cases.
Maintenance
Related MCP Connectors
An authenticated remote MCP server for user-owned devices and one-shot capability invocation.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Remote MCP server for product discovery catalog and retrieving product details.
Guarded MCP server for agent-readable business truth, provenance, readiness, and discovery.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol (MCP) server for seamless integration with peripheral devices connected to your computer. Control, monitor, and manage hardware devices through a unified API.5MIT
- AlicenseNot gradedqualityCmaintenanceEnables cloud LLM agents to discover and invoke physical hardware on edge and IoT devices through standard MCP tools, bridging constrained device channels like UART, BLE, and Wi-Fi.3MIT
- AlicenseNot gradedqualityAmaintenanceProvides a local MCP server for Bluetooth Low Energy automation, enabling scanning, GATT inspection, characteristic reads and writes, and evidence capture, diff, and replay with safety guards.10MIT
- AlicenseAqualityCmaintenanceExposes digital twin device data capabilities as MCP tools, allowing MCP clients to query device status, read real-time metrics, fetch time-series data, and check alerts. Includes a simulated PLC driver with a clean interface for connecting real devices via OPC-UA, Modbus, or gateway APIs.6MIT