Skip to main content
Glama
mctlhq

Newton MCP Gateway

by mctlhq

Newton MCP Gateway

An experimental Model Context Protocol bridge for Archetype AI Newton, the foundation model for physical-world sensor data.

This project explores two complementary integration patterns:

  1. Newton as an MCP capability — make physical-world intelligence available to any MCP-compatible agent (Claude, ChatGPT, enterprise copilots, custom hosts) through a small, auditable gateway over Newton's publicly documented /query API.

  2. MCP as an action boundary for Newton — connect physical-world understanding to external digital and physical capabilities through policy-controlled, auditable and verified actions.

Status: experimental, independent project. The MCP server and mock backend are functional. The real Newton adapter is implemented against public documentation but has not yet been validated against a live account; it activates only with an authorized ATAI_API_KEY. This project is not affiliated with or endorsed by Archetype AI.

Why

Newton turns raw sensor streams into physical understanding. MCP is becoming the common interoperability boundary for agent capabilities. Connecting the two creates value in both directions:

        DIGITAL AGENTS                                 PHYSICAL WORLD
  Claude / ChatGPT / enterprise agents          sensors, machines, environments
                 │                                           ▲
                 │ MCP                                       │ actuators (via MCP servers)
                 ▼                                           │
        newton-mcp-gateway  ──────►  Newton  ──────►  MCP Action Runtime
        (Direction B: implemented)   physical         (Direction A: proposal)
                                     understanding    policy · approval · verify

The deeper idea is not MCP itself. It is that Physical AI needs an open, safe, auditable and closed-loop boundary between understanding the physical world and changing it — and MCP may be a good foundation for that boundary.

Related MCP server: Jeltz

Quickstart (no credentials needed)

git clone https://github.com/mctlhq/newton-mcp-gateway
cd newton-mcp-gateway
uv sync
NEWTON_BACKEND=mock uv run newton-mcp        # stdio MCP server

Claude Desktop / Claude Code configuration:

{
  "mcpServers": {
    "newton": {
      "command": "uv",
      "args": ["--directory", "/path/to/newton-mcp-gateway", "run", "newton-mcp"],
      "env": { "NEWTON_BACKEND": "mock" }
    }
  }
}

With real Newton access (variable names follow Archetype's docs):

NEWTON_BACKEND=api ATAI_API_KEY=... ATAI_API_ENDPOINT=https://api.u1.archetypeai.app/v0.5 uv run newton-mcp

Tools

Deliberately few. Each maps to a documented Direct Query pattern.

Tool

Newton model family

What it does

newton_query

Newton C (text / image / video reasoning)

Natural-language question grounded in inline text/JSON events or uploaded file_ids. Use system_prompt to force structured JSON.

newton_embed_timeseries

Omega encoder

Channel-first sensor window → one 768-dim embedding per channel.

Planned (see issues): image analysis via the Files API, running Newton Agent bundles (osm, anomaly-discovery, rare-event-detection, task-verification) and paging their results.

Mock results are labelled backend: "mock" and text outputs start with [mock]. They are never presented as real Newton output.

Direction A: the Physical Action Contract (proposal)

A successful MCP tool call is not a successful physical action. set_temperature(23) can return 200 OK while the AC is offline, the wrong zone was changed, or the room simply does not cool. Physical actions need policy, approval and outcome verification on top of ordinary tool calling.

This project proposes a tool-independent Physical Action Contract (schemas/, v0.1):

{
  "goal": "reduce_room_temperature",
  "reason": "The occupied kitchen reached 29.4 C",
  "confidence": 0.96,
  "target": { "type": "environment", "location": "kitchen" },
  "constraints": { "desired_temperature_c": 23, "minimum_temperature_c": 20, "maximum_temperature_c": 25 },
  "risk": "low",
  "verification": { "condition": "temperature_c <= 24", "timeout_seconds": 600 }
}

Newton says what should happen. The MCP environment knows how. A small action runtime sits in between:

contract → capability match (MCP tool discovery) → policy (auto / confirm / deny)
        → human approval if required → execute → observe → verify → succeed / retry / escalate

The contract can be produced by a Newton /query with a strict JSON system prompt — a documented usage pattern — or by a Newton Agent's output. The policy engine in newton_mcp/action/ is a deterministic first cut. The runtime, capability resolver and verifier are the next phases.

The first real actuator testbed is a smart-home MCP server (lights, HVAC, speaker), chosen because it is real hardware with benign, reversible actions. The interface itself is domain-independent: Home Assistant, BMS, OPC-UA, ROS or enterprise workflow MCP servers plug in the same way.

What is confirmed vs. proposed

Source

/query request/response shape, model families, env variable names

Archetype public docs

Agents API (blueprints → bundles → runs, results/events)

Archetype public docs

This gateway's MCP tool mapping

this project

Physical Action Contract, policy engine, action runtime

this project's experimental proposal

Nothing here reverse-engineers private endpoints or circumvents access controls.

Development

uv sync --group dev
uv run pytest

Layout: src/newton_mcp/newton/ (backend protocol, mock, real adapter) · src/newton_mcp/action/ (contract, policy) · src/newton_mcp/server.py (MCP server) · schemas/ · examples/ · docs/.

License

MIT

Available Tools

2 tools
newton_embed_timeseriesA
Read-only

Encode a sensor window with the Newton Omega encoder. Input is channel-first: outer list = channels, inner lists = samples. Returns one 768-dim embedding per channel. Leave normalize=false unless cross-window amplitude is irrelevant.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
channelsYes
normalizeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint, openWorldHint), so the bar is lower; the description still adds real behavioral context by specifying the input layout semantics ('channel-first') and the output cardinality (one 768-dim embedding per channel), which the annotations cannot convey. It does not disclose the model default behavior, but that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with purpose, then input shape, then the normalize caveat. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the return value needn't be re-explained, and the description covers the non-obvious input structure and the normalize tradeoff. The undocumented 'model' parameter is the only gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it does for the critical 'channels' parameter ('outer list = channels, inner lists = samples') and gives a decision rule for 'normalize'. Only the 'model' parameter is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Encode a sensor window'), names the concrete encoder ('Newton Omega'), and its scope is narrow enough to distinguish it from the sibling newton_query without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this tool over newton_query, and no prerequisites described. The only conditional ('Leave normalize=false unless...') is a parameter-level instruction, not a when-to-use rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

newton_queryB
Read-only

Ask Newton (text-reasoning model) a natural-language question about physical-world data. Ground it with inline text/JSON events or previously uploaded file_ids. Use system_prompt to force structured JSON output.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
queryYes
file_idsNo
json_eventsNo
text_eventsNo
system_promptNo
max_new_tokensNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so safety and open-domain behavior are covered. The description adds genuinely useful behavioral context (it is a reasoning model that can be grounded with events/files and steered via system_prompt), but says nothing about latency, cost, token limits, or failure modes of an LLM call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and then the grounding/steering options. Every sentence contributes, with only minor duplication of the 'model' concept.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and annotations cover the read-only profile. Still, with 7 parameters at 0% coverage, the omission of model and max_new_tokens and of event-array formats leaves an agent guessing on nontrivial inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description carries the burden. It explains query, file_ids, json_events, text_events, and system_prompt, but leaves model and max_new_tokens entirely undocumented and gives no format details for the event arrays.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Ask) and resource (Newton, a text-reasoning model) scoped to physical-world data, which is clearly distinct from the sibling newton_embed_timeseries. It stops short of explicitly naming the sibling to differentiate, so it lands just below the top band.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete invocation guidance — ground the query with inline text/JSON events or previously uploaded file_ids, and use system_prompt to force structured JSON. However, it never states when to prefer this over newton_embed_timeseries or any prerequisites for the grounding inputs, leaving the choose-a-tool decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observednewton_embed_timeseries
    • First observednewton_query

TDQS

A3.7/5.0

Scored across 2 tools

Disambiguation5/5

The two tools target clearly distinct operations: natural-language querying/reasoning versus timeseries embedding. There is no overlap in purpose or inputs, so an agent can easily select the correct tool.

Naming Consistency4/5

Both tools use the newton_ prefix and snake_case, which is consistent. However, newton_query lacks the explicit object noun that newton_embed_timeseries includes, a minor deviation from a strict verb_noun pattern.

Tool Count3/5

With only two tools, the surface feels thin for a gateway, falling into the borderline range. Each tool is distinct and earns its place, but the server could likely benefit from additional support tools.

Completeness3/5

The query tool references previously uploaded file_ids, yet there is no tool to upload, list, or manage files, creating a notable gap. Core query and embedding operations are present, but the grounding workflow is incomplete without file handling.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables LLMs to connect to n8n workflows, Blender 3D, Notion, and filesystem tools through a single MCP hub with hot-pluggable bridges. Features a pedagogical transparency layer for educational purposes and supports multiple LLM providers including Claude, OpenAI, Gemini, and Ollama.
    4
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Connects AI assistants to physical devices via MCP, enabling LLMs to reason about sensor data, perform cross-device correlation, and interpret anomalies.
    4
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    MCP server for chatting with physical-world data from robotics, drones, automotive, and IoT sources using natural language. It generates auditable SQL queries over Apache Arrow/DuckDB to let you analyze, summarize, and build data pipelines.
    25
    397
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP clients to discover and create agents, durably submit messages with verifiable replies, and search source-linked local memory. Exposes eight authenticated tools with owner-approved OAuth, transcript-backed verification, and explicit control over history import and embedding.
    MIT