Oneiros MCP Server
The Oneiros MCP Server exposes a JEPA-style latent world model as agentic tools, enabling planning and navigation in a 2D point-mass environment via learned latent space.
encode_observation: Convert a 6D observation vector[x, y, vx, vy, goal_x, goal_y]into a compact latent embedding using the trained JEPA encoder.predict_rollout: Given a starting latent and a sequence of 2D acceleration actions, roll the learned latent dynamics forward step-by-step to imagine future trajectories — without touching the real environment.plan_to_goal: Use latent-space Model Predictive Control (MPC via the Cross-Entropy Method) to search over action sequences, score them by predicted latent distance to the goal, and return the best next action. Call repeatedly each step for receding-horizon planning.reset_env: Initialize the 2D point-mass environment to a deterministic starting state (with an optional seed), returning the initial observation and goal position.step_env: Apply a 2D acceleration action to advance the environment one timestep, returning the new observation, reward (negative distance to goal), a done flag, and current distance to goal.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Oneiros MCP Serverplan a path to the goal at (5,3)"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Oneiros
Status: complete research artifact — results verified. v0.2 adds image-based perception as a demonstrated result (the four pixel gates below), closing v0.1's documented next step; the V-JEPA 2 scale-up remains future work (needs GPU + pretrained weights).
A JEPA-style latent world model that an agent calls as a tool — over MCP — to plan.
Oneiros is a small, CPU-only research artifact at the intersection of agentic systems and world models. It trains a Joint Embedding Predictive Architecture (JEPA) world model on a 2D point-mass environment, then exposes that model as a set of Model Context Protocol tools. An agent plans by calling the learned predictive world model as a tool — encoding observations into latents, rolling latent dynamics forward, and running model-predictive control entirely in latent space — rather than the usual pattern of an LLM calling hand-written functions.
Why predict in latent space
A JEPA predicts the future in a learned latent space, not in pixels or
tokens. Given an observation o_t and an action a_t, it learns an encoder
f and a predictor g such that g(f(o_t), a_t) matches f(o_{t+1}) — there
is no decoder and no pixel-reconstruction loss. This matters because
reconstruction wastes capacity modeling perceptually salient but
control-irrelevant detail (texture, lighting, background), while a latent
predictor is free to discard everything that does not help it anticipate the
future. The cost is a well-known failure mode: latent prediction can collapse to
a constant (every observation maps to the same point, making prediction
trivially perfect). Oneiros defeats collapse with an EMA target encoder,
stop-gradients, and a VICReg-style variance + covariance penalty, and then uses
the resulting latent dynamics for planning.
Related MCP server: MuJoCo MCP Server
Architecture
flowchart LR
subgraph Env["Point-mass environment (numpy)"]
O["obs o_t"]
ON["obs o_t+1"]
end
subgraph WM["JEPA world model (torch, CPU)"]
F["encoder f"]
FT["EMA target f_target<br/>(stop-grad)"]
G["predictor g"]
O --> F --> Z["latent z_t"]
Z --> G
A["action a_t"] --> G
G --> ZH["z_hat_t+1"]
ON --> FT --> ZT["z_t+1 (target)"]
ZH -. "MSE + VICReg<br/>variance/covariance" .-> ZT
end
subgraph MCP["MCP server (agentic interface)"]
T1["encode_observation"]
T2["predict_rollout"]
T3["plan_to_goal"]
T4["reset_env / step_env"]
end
subgraph Agent["Agent loop"]
P["latent MPC planner<br/>(CEM over g)"]
end
WM --> MCP
MCP <--> Agent
P -->|"first action"| EnvThe agent never sees the environment's dynamics. It calls plan_to_goal, which
encodes the current and goal observations, searches action sequences by rolling
the predictor g forward H steps in latent space (cross-entropy method),
scores each candidate by predicted-latent distance to the goal, and returns the
first action. The planner replans every step (receding-horizon MPC).
Verified results
Numbers below are from an actual run on this machine (CPU only, seed 0). Train
with python -m oneiros.train and reproduce the diagnostics with
python -m oneiros.demo_agent.
Honesty gate | Metric | Result |
(a) Predictor beats no-op baseline | next-latent MSE vs identity baseline | 0.0185 vs 0.1076 (ratio 0.17 — ~5.8x better) |
(b) Latent not collapsed | per-dim latent std (mean / min) | 1.04 / 1.00 (threshold 0.1) |
(c) MPC beats random | goal-reaching success over 20 seeds | MPC 95% vs random 15-20% |
Training takes about 21 seconds for 4000 steps. The checkpoint
(oneiros/checkpoint.pt, ~240 KB) is committed so the demo, MCP server, and
planning tests run without retraining.
The pixel gates (v0.2): image-based perception, demonstrated
v0.1 documented why the image encoder could not beat the identity baseline. The diagnosis had two parts, and each got a principled fix rather than a knob-twiddle:
Consecutive frames were nearly identical (the blob moves ~a pixel per step), so "predict no change" was already an excellent predictor. Fix: the swift environment preset (
PointMassConfig.swift()) — largerdtand acceleration so the agent moves several pixels per frame, plus speed-proportional drag so the dynamics are genuinely non-linear.A single frame hides velocity — the dynamics are second-order, so no single-frame predictor can recover the next state, and the convolutional encoder exploited this by temporal smoothing (mapping consecutive frames to nearly identical latents; the dataset-wide variance penalty does not forbid it — more training made it worse, 0.89 → 0.94 MSE ratio). Fixes: two-frame stacking (velocity becomes observable from pixels, the same reason pixel world models from DQN to V-JEPA consume clips, not stills) and a delta-variance penalty (VICReg-style hinge on the std of
z_{t+1} - z_t) that forbids the temporal collapse outright.
Results from the committed checkpoint_image.pt (~735 KB), heldout data,
enforced as tests in tests/test_gates_image.py:
Pixel gate | Metric | Result |
(d) Image predictor beats no-op | heldout next-latent MSE ratio vs identity | 0.11 (vector model: 0.17) |
(e) Latent not collapsed, incl. temporally | per-dim std / one-step delta MSE | 1.10 / 1.08 (pre-fix delta was 0.012) |
(f) Pixel MPC beats random | goal-reach rate + median steps over 20 seeds | 20/20, median 9.5 steps vs 27.5 random |
(g) Imagination useful at horizon | compounded H-step rollout MSE ratio | 0.45 at H=4, 0.86 at H=12 |
One honest negative, measured and deliberately not gated: a privileged
linear-dynamics MPC reading the true 4D state still reaches the goal about
twice as fast as the pixel planner (median ~5 vs ~9.5 steps). Planning from
pixels has not caught planning from privileged state, and gate (g) shows
open-loop imagination degrading by horizon 12 — which is exactly why the
planner replans every step. The baselines live in oneiros/baselines.py;
reproduce with the image-training command under Run.

The agent drives the point-mass to the goal (green star) using only the world-model planning interface.
Latent prediction error | Planning success |
|
|
What this is — and isn't
This is a genuine, end-to-end demonstration that (1) a non-trivial latent dynamics model can be learned without collapse and without reconstruction, and (2) planning in that latent space solves a control task far better than chance, all behind an agentic tool interface.
It is not at scale. The environments are toy 2D point-masses, the models
are tiny (a few hundred KB). The default committed model uses a
vector-state observation on the original environment; the committed
image model (v0.2, checkpoint_image.pt) perceives stacked 32x32 frames
on the swift environment and clears its own four gates above — v0.1's
"image-based perception is the documented next step" is now a demonstrated
result, with the diagnosis (frame similarity + hidden velocity + temporal
smoothing) and fixes documented rather than hand-waved. What remains future
work is the scale-up: a frozen pretrained perception encoder such as
V-JEPA 2 with a learned latent dynamics head on
top, exactly the recipe this toy mirrors — that step needs a GPU and
pretrained weights, which this CPU-only artifact deliberately does not
assume.
Install
Requires Python 3.12. A project-local virtual environment is recommended.
python -m venv .venv
# Windows: .venv\Scripts\activate | Unix: source .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install -e ".[dev]"torch is installed from the CPU wheel index — no GPU is needed or used.
Run
# Train the JEPA world model (writes oneiros/checkpoint.pt). ~21s on CPU.
python -m oneiros.train --obs-mode vector --steps 4000
# Train the image world model on the swift environment (~3 min on CPU).
# This is the exact recipe of the committed checkpoint: two-frame stacking,
# the delta-variance penalty, and an explicit --checkpoint so the vector
# model's default path is not clobbered.
python -m oneiros.train --obs-mode image --env swift --steps 6000 \
--frame-stack 2 --delta-var-weight 25 --checkpoint oneiros/checkpoint_image.pt
# Run the scripted agent: drive to the goal via latent MPC, write the GIF +
# diagnostic plots to assets/.
python -m oneiros.demo_agent --seed 0 --k 20
# Tests, including the three honesty gates (uses the committed checkpoint).
pytest -q
# Lint.
ruff check oneiros testsMCP server
The world model is exposed over MCP as a stdio server:
# vector model (default)
python -m oneiros.mcp_server
# the committed image model: stacked-frame observations, swift environment
python -m oneiros.mcp_server --obs-mode image
# any other checkpoint
python -m oneiros.mcp_server --checkpoint /path/to/checkpoint.pt--obs-mode selects which committed world model the server exposes. The
environment tools serve observations in the loaded model's format — 6D
vectors, or the stacked rendered frames the image model trains on — and
reset_env returns a ready-made goal_observation for the planning tools.
The served environment always uses the configuration the checkpoint was
trained on (the swift preset for the image model).
Tools:
Tool | Purpose |
| observation -> latent |
| roll latent dynamics |
| latent-space MPC; returns the next action toward a goal |
| full best action sequence + the imagined latent path |
| loaded model's obs mode, dims, and honesty-gate metrics |
| drive the point-mass environment; a |
To register the server with Claude Desktop, add this to
claude_desktop_config.json (use absolute paths for your checkout):
{
"mcpServers": {
"oneiros": {
"command": "C:/path/to/Oneiros/.venv/Scripts/python.exe",
"args": ["-m", "oneiros.mcp_server"],
"cwd": "C:/path/to/Oneiros"
}
}
}On Unix the command is .venv/bin/python. Append "--obs-mode", "image"
to args to serve the committed image world model instead of the vector one.
Repository layout
oneiros/
env.py # deterministic 2D point-mass environment (numpy)
model.py # JEPA encoder + predictor + VICReg regularizers
data.py # random-policy rollout replay buffer
train.py # JEPA training loop, evaluation, checkpoint I/O
planner.py # latent-space MPC (CEM / random shooting)
baselines.py # pixel episode runners + privileged linear-MPC baseline
mcp_server.py # MCP tools exposing the world model
demo_agent.py # scripted agent + diagnostics (GIF, plots)
checkpoint.pt # committed vector model (~240 KB)
checkpoint_image.pt # committed image model, swift env (~735 KB)
tests/ # determinism, shapes, and the three honesty gates
assets/ # generated GIF and diagnostic figuresSee SYNERGY.md for how the same latent-dynamics idea connects to regime-aware modeling in time series.
License
MIT — see LICENSE.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Alicense-qualityDmaintenanceA domain-agnostic MCP server for autonomous experimentation, generalizing Karpathy's autoresearch pattern into a reusable server that any AI agent can drive, pointed at any domain defined by a JSON configuration.Last updated3Apache 2.0
- Alicense-qualityDmaintenanceExposes MuJoCo physics simulation to AI assistants via 65 MCP tools, enabling natural language control of robotics simulation, trajectory optimization, contact analysis, and video export.Last updated8MIT
- Alicense-qualityDmaintenanceEnables AI agents to interactively explore PDDL planning problems by exposing a PDDL engine as MCP tools for initialization, action execution, state inspection, and goal checking.Last updated4Apache 2.0
- AlicenseCqualityAmaintenanceEnables AI agents to autonomously develop and test Godot 4 games through an MCP-based feedback loop, providing tools for authoring, running, observing, playtesting, and verifying game projects.Last updated203662MIT
Related MCP Connectors
Agent Replay Debugger MCP — record every agent step + deterministic replay. Step-debugger for
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to control Unreal E…
Shared long-term memory vault for AI agents with 20 MCP tools.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/mal0ware/Oneiros'
If you have feedback or need assistance with the MCP directory API, please join our Discord server

