Laya MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Laya MCP ServerRead laya_guide, then use laya_decide to classify 'I was charged twice' as refund or billing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Local CPU decisions and MCP-native route coordination. One container and one resident model. MCP, REST, Swagger and a test console. No GPU or paid API.
By htrnguyen · Tiếng Việt
START
Requires Docker Compose v2+. Start the pure MCP service without global client registration:
docker compose up -d --buildFirst startup downloads dependencies/model weights. The container remains loopback-only with its existing CPU and memory limits. Setup starts the local service only and does not edit MCP client configuration.
Open | URL |
Test console | |
Swagger | |
MCP | |
Health |
The primary tool endpoint is http://127.0.0.1:30765/mcp.
Related MCP server: oh-my-laya
USE
Read laya_guide, then use laya_decide to classify “I was charged twice and need a refund.” Include caller/task and show the tool result.
Tool | Purpose |
| Usage/schema/limits, no inference |
| Loaded model, device, revision |
| 1–3 typed questions and normalized results |
| Compatible single-classification tool |
| Advisory routes over caller-supplied candidates; no command execution |
| Search supplied metadata and retrieve route traces |
| Unverified MCP caller self-report |
| Read/update local route policy |
Choice selects among 2–10 labels. Noul estimates yes/no probability. Score rates ordered descriptions and is experimental. Noul/score always require review. Short input and one question are the best starting point for CPU latency.
REST: POST /v1/decide uses state, questions, caller, task and returns validated decisions,
review flags, usage, timing and trace ID. Raw /v1/systemone remains available.
Examples are in the console and Swagger. Agents can read /llms.txt, /guide.md, /skill.md,
/rules.md, /integrations.json, /openapi.json, or MCP resource laya://guide.
CONFIGURE & OPERATE
Edit .env, then rerun setup. Defaults are in .env.example.
Main settings: LAYA_PORT=30765, LAYA_THREADS=4, LAYA_CPUS=4, LAYA_MEMORY=4g.
sudo docker compose ps
sudo docker compose logs -f --tail 50
sudo docker compose stop
sudo docker compose start
sudo docker compose down # keeps model filesLogs go only to Docker stdout. Default LAYA_LOG_FORMAT=compact shows result, caller/task,
latency and trace. Use json for bounded input/output details, then rerun setup.
Docker rotates at 10 MiB × 2 files; recreating a container may discard its old logs.
Successful health polls are hidden. Nested MCP/REST timings must not be added together.
Keep data/models: it avoids downloading weights again. Setup caches, optional reports and private
config backups are ignored by Git. No persistent application log folder is needed;
setup removes old request log files only after successful deployment.
restart: always needs an existing container; after compose down, run setup again.
LIMITS & CHECKS
This is a trusted local-only service, without authentication. Do not expose it directly to LAN/Internet. Logs in JSON mode may contain private text. Structured credential keys are redacted; free-form secrets are not. Model confidence is not proof or permission to act. Use tests/code for exact checks, and the main agent for writing, planning or multi-step reasoning. Invalid/truncated normalized responses are rejected.
uv run --no-project --with-requirements tests/requirements.txt python -m unittest discover -s tests -v
sudo docker compose exec -T laya python /app/scripts/verify_mcp.py
python3 scripts/benchmark.py --runs 100An earlier six-case repeated CPU smoke test measured p50 103 ms, p95 116 ms on Ryzen 5 5600H (4 threads, 100 warm calls). Not an accuracy guarantee or a benchmark of the multi-question API. Full container rebuild/reboot and browser interaction must be verified on the target host.
SOURCES & LICENSE
Laya / Convai Innovations — upstream runtime; pinned commit, docs, benchmarks.
Laya weights — the pinned runtime uses its multilingual subfolder; see multilingual limitations and SOURCE.json.
/healthreports the actual loaded revision.Swagger UI — bundled 5.33.1; asset provenance.
Docker Compose / uv.
TypeSafe / Jev and Jev, Clearly Explained — Avi Chawla — architectural reading only.
Integration author: htrnguyen. Apache-2.0 · NOTICE · CHANGELOG. Original Laya/Swagger licenses remain with their code. Model weights are not committed. This project is not affiliated with TypeSafe and does not claim equivalent quality to Jev. Direct dependencies are pinned; base-image/transitive dependencies are not fully locked.
This server cannot be deployed
Maintenance
Related MCP Connectors
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
MCP server for building and testing AI agents with multi-model experimentation and insights.
HiveCompute MCP Server — decentralized inference router for AI agents
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to offload quick judgment calls like shipping readiness, file triage, and claim verification to a fast local MCP server with calibrated confidence and safe fallbacks.6MIT
- AlicenseAqualityBmaintenanceEnables local Laya-MLX typed decisions (classification, scoring, risk routing, and yes/no) for coding agents like Codex, Claude Code, DeepSeek Harness, and Pi via MCP, running entirely on Apple Silicon Macs.13MIT
- AlicenseNot gradedqualityAmaintenanceEnables local typed decision-making over MCP Streamable HTTP, running Laya and Von models to turn arbitrary state into choice, score, and yes/no answers with probabilities.3Do What The F*ck You Want To Public
- AlicenseAqualityBmaintenanceProvides fast, offline typed decision tools (classify, score, check, triage) to agents via the Model Context Protocol, enabling safe local inference across 100+ languages with calibrated confidence and no text generation.878 PyPIMIT