reflex
Routes requests to Obsidian as the selected MCP backend, enabling decisions such as choosing Obsidian for note search and retrieval.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@reflexDecide which MCP server should handle: 'what did I bookmark about cilium?'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
reflex
A System-1 decision layer for homelab and self-hosted platforms: give it a state and typed questions, it returns calibrated routing decisions in a single forward pass — no generation, no parsing, no hallucination.
Built on convaiinnovations/laya
(System-1 decision models, Apache-2.0), packaged as a production service:
YAML policies in, MCP/HTTP routing decisions out.
flowchart LR
subgraph Consumers
GW[agentgateway]
TH[toolhive / vmcp]
AG[agents / hermes]
end
subgraph reflex
API["/decide (FastAPI)"]
MCP["MCP server (Streamable HTTP)"]
POL["policies/*.yaml → typed question schemas"]
RTR["ONNX runtime (English base checkpoint)"]
end
GW --> API
TH --> MCP
AG --> MCP
API --> POL --> RTR
MCP --> POLWhat it routes
Policy | Decision | Example |
| which MCP server / backend should handle a request | karakeep vs hindsight vs netbox vs obsidian |
| which tool inside a chosen server, plus relevance score |
|
| which agent profile should take the message | ops vs chat vs default |
All three are choice-type questions with calibrated probabilities — schemas
live in policies/*.yaml, so new routes ship without retraining. When base
accuracy isn't enough, the production fine-tune path (shadow traffic →
domain dataset → fine-tune → repackage) produces your own checkpoint, the same
way laya-typed-decisions derives from laya.
Related MCP server: jev-mcp
What it is not
Not an embedding service — no vectors; use TEI for embeddings.
Not a reranker —
scoreprimitives can rank candidates, but a dedicated cross-encoder reranker is the right tool for large candidate sets.Not an LLM — it cannot chat, summarize, or generate anything.
Not a guardrail — never sits in a deny path; deny-path scanning belongs to fail-closed sidecars (e.g. mcp-guardrails).
Not enforceable pre-fine-tune — the bundled base checkpoint is advisory/shadow-only. Live measurements (2026-09-25, see docs/EVALUATION.md): correct top-1 on clear routes, but borderline confidence is far below enforcement grade and CPU latency is ~1 s/decision. Do not gate traffic on it until a fine-tuned, temperature-refit checkpoint is promoted.
Quickstart
pip install -r requirements.txt
uvicorn app.server:app --host 0.0.0.0 --port 9000curl -s localhost:9000/decide/mcp -d '{"state": {"text": "what did I bookmark about cilium last week"}}'
# {"policy":"mcp","answer":{"choice":"karakeep","probability":0.91},...}Docker / Kubernetes: see Dockerfile and deploy/k8s.
Runtime: ONNX, no torch
The release image runs onnxruntime against the bundled English base
checkpoint — the root of tozp/laya-onnx
(model.onnx + tokenizer.json + rl_agent_config.json; the
laya-typed-decisions artifact is a separately tuned checkpoint and is
deliberately NOT bundled: shadow data must come from the base we intend to
fine-tune) — no torch, no transformers, no CUDA libs. That takes the image
from ~5.4 GB (default PyPI
torch wheel pulls the full nvidia stack) to ~2 GB fp32, or ~800 MB
with --build-arg MODEL_FILE=model_int8.onnx (measured RSS: fp32 ≈ 2.4 GiB
after warm-up → size pods at ~3 Gi limit; int8 ≈ 0.75 GiB → ~1 Gi limit).
Inference is one forward pass per call; encode/decode mirrors laya's
contract exactly (temperature clamping included), so decisions are
interchangeable with the torch path.
Measured (production nodes: Intel 13900H 6P+8E/96 GB; i3-N305 = dev box, conservative lower bound; details in docs/EVALUATION.md):
runtime | 13900H (cluster) | N305 (dev) | warm RSS | image |
torch 0.3.0 (previous) | 0.3–0.6 s (measured live, 2-CPU limit) | — | ~1.5 GiB | 5.36 GB |
ONNX fp32 (default) | ~1.0 s (measured live 2026-09-25, 2-CPU limit) | 9 s (measured) | 2.4 GiB → 3 Gi limit | ~2 GB |
ONNX int8 (build-arg) | ~1–1.7 s (inferred) | 5 s (measured) | 0.75 GiB → 1 Gi limit | ~800 MB |
ONNX fp32 decisions are bit-identical to the torch deployment (4/4 cases, probabilities and confidence to 4 decimals). Current latency is bounded by the export's static 512-token graph, not the hardware — CPU is adequate for advisory/shadow traffic; inline high-QPS enforcement (~33 ms on the model card's T4) wants an accelerator (iGPU/OpenVINO, Apple-silicon host, or cloud burst).
# default: tozp/laya-onnx fp32
docker build -t reflex .
# lean variant, or your own fine-tuned export (see finetune/):
docker build --build-arg MODEL_FILE=model_int8.onnx -t reflex .
docker build --build-arg MODEL_REPO=<your-hf-repo> -t reflex .Decision tracing
Set OTEL_EXPORTER_OTLP_ENDPOINT (standard OTEL env vars) and every policy
decision emits one span with gen_ai.* attributes — input, choice,
probability, checkpoint — rendered natively as generations in langfuse or
any OTLP backend. The SDK is inert when the endpoint is unset. Raw-schema
calls (/decide, reflex_decide) are not traced: no policy label, varying
schema, no dataset value.
Shadow deployments MUST set OTEL_TRACES_SAMPLER=always_on. The SDK default
(parentbased_always_on) drops spans under unsampled parents, and MCP
front-proxies sample aggressively (toolhive: 5%) — that silently shrinks the
shadow dataset to a biased sliver.
This is the shadow-mode log source for step 1 below: spans out, export from your backend, distill, fine-tune.
Fine-tuning is a production feature
reflex treats fine-tuning as an operational stage, not research:
Shadow mode — log every decision + the route actually taken.
Distill — build the domain dataset from traffic logs (teacher = your production LLM or rule labels).
Fine-tune — upstream notebook (laya_finetune_typed_decisions) against the dataset; refit temperature on your data (required — base checkpoints ship over-confident).
Repackage & promote — versioned checkpoint into the image, shadow again, then enforce.
See ARCHITECTURE.md for the design deep dive and finetune/ for the pipeline contract.
License
Apache-2.0 (see LICENSE). Model weights: Apache-2.0 (Convai Innovations).
This server cannot be deployed
Maintenance
Related MCP Connectors
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
MCP server for building and testing AI agents with multi-model experimentation and insights.
AI routing, memory, guardrails, and governance. Routes across Claude, GPT, Gemini.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables cross-lingual semantic intent routing and multilingual agent dispatch through a deterministic, zero-dependency MCP server, allowing AI agents to parse and route user queries across languages and delegate to appropriate agents with structured JSON output.8-
- AlicenseAqualityCmaintenanceEnables agents to make calibrated decisions via six MCP tools for classification, relevance ranking, claim verification, action gating, next-step control, and model listing, using Jev's System One model without generating text.622MIT
- AlicenseAqualityCmaintenanceProvides agents with fast, typed, calibrated decision tools for classification, scoring, yes/no checks, and gating risky tool calls.51,141 npm2MIT
- AlicenseAqualityBmaintenanceEnables MCP clients to run fast, local, non-autoregressive System-1 decisions on text, including classification, triage, routing, and safety gating, via six typed tools with probabilistic outputs.208Apache 2.0