Skip to main content
Glama

mcp-spatial-perception

CI Publish PyPI version Python versions License: MIT Ruff

Give your AI agent eyes on the physical world.

An MCP server that exposes real-time spatial perception queries from DePIN edge-vision nodes to any LLM that speaks the Model Context Protocol — Claude Desktop, Cursor, Claude Code, or a custom agent framework.

PyPI page


The problem

LLM agents can read files, browse the web, and call APIs — but they're blind to the physical world. They can't answer questions like:

  • "Is there foot traffic at location X right now?"

  • "What is the live visual state of camera node #402?"

  • "Is the loading dock clear before I dispatch the truck?"

There's no standard way for an agent to ask those questions.

Related MCP server: Cheqd MCP Toolkit

The solution

mcp-spatial-perception is a small, dependency-free MCP server that sits between your agent and a DePIN network of edge-vision nodes. The agent calls a tool, the server queries the network, and structured JSON describing the scene comes back — detections, bounding boxes, confidence scores, environmental state, geolocation.

┌─────────────┐   MCP / JSON-RPC   ┌───────────────────────┐   DePIN query   ┌──────────────┐
│  LLM agent  │ ─────────────────► │ mcp-spatial-perception│ ──────────────► │ edge nodes   │
│ (Claude,    │ ◄───────────────── │  (this repo)          │ ◄────────────── │ (cameras,    │
│  Cursor…)   │   structured JSON  └───────────────────────┘   telemetry     │  sensors)    │
└─────────────┘                                                               └──────────────┘

Install

pip install mcp-spatial-perception

Requires Python 3.10+. No runtime dependencies — pure standard library.

Use with Claude Desktop

Add this to your claude_desktop_config.json:

{
  "mcpServers": {
    "spatial-perception": {
      "command": "mcp-spatial-perception"
    }
  }
}

Restart Claude Desktop. The query_spatial_feed tool will appear in the tool picker.

Use with Cursor / Claude Code / any MCP client

The server speaks JSON-RPC 2.0 over stdio. Any MCP-compatible client can launch it via the mcp-spatial-perception console script:

mcp-spatial-perception

Or directly, if you have the source checked out:

python mcp_spatial.py

Use programmatically

from mcp_spatial import handle_mcp_request

response = handle_mcp_request(
    {
        "jsonrpc": "2.0",
        "id": 1,
        "method": "tools/call",
        "params": {"name": "query_spatial_feed", "arguments": {"node_id": "node_402"}},
    }
)

print(response["result"]["content"][0]["text"])

The query_spatial_feed tool

Input

Field

Type

Required

Description

node_id

string

yes

The DePIN node ID to query.

Output — structured JSON describing the node's current visual state:

{
  "node_id": "node_402",
  "timestamp": 1770000000,
  "location": { "lat": 9.0765, "lon": 7.3986 },
  "detections": [
    {
      "object": "delivery_truck",
      "confidence": 0.94,
      "bounding_box": [120, 80, 450, 300]
    },
    {
      "object": "person",
      "confidence": 0.88,
      "bounding_box": [50, 60, 110, 200]
    }
  ],
  "environmental": { "light_level": "daylight", "obscured": false }
}

Field

Meaning

node_id

Echo of the queried node.

timestamp

Unix seconds when the frame was captured.

location

Latitude / longitude of the node.

detections

Objects found in the frame, with bounding boxes.

environmental

Lighting, occlusion, and other scene conditions.


FAQ

Why an MCP server and not just a REST API? Because MCP is what Claude Desktop, Cursor, and Claude Code speak natively. A REST API would require each client to write a custom integration. An MCP server is a drop-in tool for every MCP-aware agent.

Why is the feed mocked right now? To keep the protocol surface testable and stable while the DePIN adapter is built. The mock_node_feed function is a single, well-isolated seam — swapping it for a real network client doesn't touch the MCP logic.

Does this run the vision model? No. Edge nodes do the frame extraction and inference on-device; this server relays the structured result. That's the point of DePIN — the compute is at the edge, not in your agent's process.


Status

⚠️ Alpha. The MCP protocol surface is stable, but the node feed is currently mocked. A real DePIN adapter is on the roadmap.

Roadmap

  • Spec-compliant MCP server (initialize, tools/list, tools/call)

  • Published to PyPI with Trusted Publishing

  • CI: lint (ruff), type-check (mypy strict), test (pytest) on Python 3.10–3.12

  • Replace mock feed with a real DePIN adapter

  • Add query_by_location(lat, lon) tool

  • Add subscribe_to_node(node_id) push notifications

  • Frame snapshot retrieval

  • Auth / signed node requests


Development

git clone https://github.com/jamie643/mcp-spatial-perception.git
cd mcp-spatial-perception
pip install -e ".[dev]"

ruff check .          # lint
ruff format --check . # format check
mypy mcp_spatial.py   # type-check (strict)
pytest                # tests + coverage

Contributing

Issues and pull requests are welcome. For substantial changes, please open an issue first to discuss what you'd like to change.

License

MIT © 2026 Jamie643

Available Tools

1 tool
query_spatial_feedC

Fetch real-time edge-parsed visual detection data from a physical DePIN node feed.

ParametersJSON Schema
NameRequiredDescriptionDefault
node_idYesThe DePIN node ID to query.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Fetch' and 'real-time' imply a read against a live source, but it says nothing about rate limits, auth requirements, whether results are a snapshot or stream, freshness windows, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is appropriately sized for a one-parameter tool, though the terseness comes at the cost of missing detail rather than from genuine compression.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must explain the return shape, but it never says what fields a detection record contains, what the freshness or time-window semantics of 'real-time' are, or what an empty result means. An agent can call it but cannot anticipate the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter (node_id) is already documented in the schema as 'The DePIN node ID to query.' The description adds no meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch') and resource ('real-time edge-parsed visual detection data from a physical DePIN node feed'), which is enough to distinguish it from nothing since there are no siblings. The resource noun is jargon-heavy ('edge-parsed', 'DePIN node feed') and never grounds what a detection datum actually is, but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no alternatives named (none exist as siblings). The single word 'real-time' hints at a freshness use case but the agent is left to infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.4
    • First observedquery_spatial_feed

TDQS

B3/5.0

Scored across 1 tool

Disambiguation5/5

With only a single tool in the set, there is no possibility of overlap or misselection. The purpose—fetching real-time edge-parsed detections from a DePIN node—is unambiguous.

Naming Consistency4/5

The lone name query_spatial_feed follows a clean verb_noun snake_case convention that is easy to read. A single tool cannot demonstrate a consistent pattern across the set, so this is scored slightly below full marks.

Tool Count2/5

One tool is too thin for a domain like spatial perception, which naturally involves node discovery, multi-region queries, and time-window filtering. A single feed reader leaves the server under-scoped.

Completeness2/5

The surface only supports fetching live detections; there is no way to list available nodes, filter by region/object class, or query historical data. These gaps will force agents into dead ends for anything beyond a raw live read.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers