omp-laya-judge
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@omp-laya-judgeClassify this feedback as positive, neutral, or negative: 'The update is great!'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
omp-laya-judge
Local System-1 judge for oh-my-pi,
powered by laya. Typed decisions
(choice/bool/score) at mean 402ms (p50 238ms, p95 1108ms) on CPU,
0 LLM tokens burned, nothing leaves the machine.

The decision layer, mid-turn. Every card is a real answer from the local
sidecar - the notes gate picking project notes, the edit gate judging a file
change, the bash gate reading a destructive command, the recovery gate
choosing what to do about a failed tool - drawn by the same row builder the
TUI uses (hooks/lib/record.ts) and regenerated with
python benchmark/live_feed_gif.py:

Measured head-to-head
Same 12 classification questions, real model via omp -p vs this plugin
on CPU, no GPU:
LLM | laya-judge | |
latency / question | 18.4s (16–25s sampled) | mean 402ms (p50 238ms, p95 1108ms) |
tokens / question | 728 | 0 |
accuracy (12-case bench) | n/a (reference) | 10/12 |
Reproduce everything below from committed artifacts:
python benchmark/run.py # -> benchmark/results.json (schema 2)
python benchmark/calibrate.py # -> benchmark/calibration.json (gate metrics)
python benchmark/chart.py # -> assets/benchmark.svg
python benchmark/before_after.py # -> assets/before-after.gifRelated MCP server: agent-fastpath
The 0.6 gate, honestly
Policy: answers with confidence ≥ 0.6 are auto-accepted, below escalates
to the LLM. What that bought on the 12-case bench
(benchmark/calibration.json, computed — not asserted):
auto-accept 9/12, 3 escalated to the LLM
2 misses total: both caught by the gate, 0 escaped
false-accept 0% (0 of the 9 auto-accepts were wrong)
Both misses are score questions — the severity head lands on the right
neighbourhood (2.4 for a 3) but not the exact bucket, and both sit under the
gate, so they escalate instead of being trusted.
Arithmetic no longer reaches the model at all. The parity misses that used to
escape at confidence 0.80/0.84 are answered by exact arithmetic in
server/core.py:resolve_arithmetic, which claims a yes/no question only when
the state carries exactly one distinct number and the instruction names an
operation on it (parity, primality, divisibility, comparison). Everything
ambiguous is still laya's.
Checkpoint A/B/C + random
(benchmark/compare.json, reproduced by benchmark/compare_models.py).
These are raw checkpoint scores — the model answering alone, with no
arithmetic resolver — which is why the english row reads 8 while the
headline above reads 10: the two extra cases are parity, answered exactly by
server/core.py rather than by the checkpoint.
EN 12-case | TR 4-case | mean conf | mean ms | |
english | 8 | 2 | 0.68 | 210 |
typed-decisions | 8 | 3 | 0.54 | 198 |
multilingual | 6 | 2 | 0.76 | 111 |
random (coin flip) | 6 | 0 | — | 0 |
TR sample is small (4); the honest read is laya >> chance (2–3 vs 0), not a checkpoint coronation.
Which judge answers
The sidecar speaks POST /v1/systemone, so the model behind it is a setting,
not a fork. Measured on the same 56 cases with one grading function
(benchmark/compare_backends.py; routing 12, triage 20, severity 11, the real
mixed JSON call 6, Turkish 7):
backend | correct | median | p95 | route | triage | severity | mixed | Turkish | ECE |
laya/english (default) | 33/56 | 242ms | 491ms | 11/12 | 16/20 | 3/11 | 3/6 | 0/7 | 0.443 |
von 1.2 | 29/56 | 269ms | 486ms | 9/12 | 15/20 | 4/11 | 1/6 | 0/7 | 0.348 |
laya/multilingual | 26/56 | 88ms | 173ms | 9/12 | 13/20 | 3/11 | 1/6 | 0/7 | 0.480 |
Switch backends by environment; the hook never changes your models for you.
LAYA_BACKEND=systemone LAYA_UPSTREAM_URL=http://127.0.0.1:8000 ... # von, or jev
LAYA_BACKEND=systemone LAYA_UPSTREAM_URL=https://api.typesafe.ai/v1/systemone \
LAYA_UPSTREAM_KEY_ENV=TYPESAFE_API_KEY ... # jev, your key
LAYA_MODEL=multilingual ... # faster, less sureLAYA_UPSTREAM_KEY_ENV names the variable holding the key, so the key itself
never lands in a config file. /info reports which backend is actually
answering, and arithmetic stays local on every backend because it is exact
and free everywhere.
Two honest limits from that table: Turkish is 0/7 on all three - including
laya's multilingual checkpoint - so the English checkpoints are not usable
for non-English text, and the layer should not gate on it. And laya's
answer_confidence is the field to threshold, not the confidence it also
returns; the gates use the calibrated one, which raised local auto-accept
from 7/12 to 9/12 with false-accept still at 0%.
Demos

Grounded quiz (python demo/quiz.py): 6/8, mean 282ms, 0 tokens —
both misses came in under 0.2 confidence, exactly the cases the escalate
rule covers.

Snake (python demo/snake.py): score 12, 300 moves, alive, 12 shield
interventions, ~0.9s/move on CPU, 0 tokens. Watch every decision in the
browser: open web/snake.html (canvas replay with live probability bars,
play/pause/speed/scrub); quiz at web-quiz.html. Recipe (from
laya-mlx): a deterministic planner
describes each direction over a tiny state (Safe route: yes. Food reachable: …), a safety shield executes the best SAFE move. Short parallel criteria are
the whole trick — greedy end-to-end choice without the planner scores 0 (see
demo/snake.py header for the measured failure modes). Honest calibration:
per-move confidences on game states run at noise level (0.005-0.05; the quiz
gets 0.2-1.0 on text) — navigation is planner + shield with laya ranking, the
same division of labor as laya-mlx. The maze variant (demo/maze.py) is a
documented negative: corridors trap the noise walk (1 crumb/200 steps).

Tetris (python demo/tetris.py): score 5720, 43 lines, 120 pieces, alive,
~1s/placement on CPU, 0 tokens. Per piece the planner enumerates every
legal landing, shortlists 6 by classic features (lines, holes, height,
bumpiness), laya picks one in a batched call, and a shield keeps the stack
out of the top-4 danger zone. Replay in browser: web/tetris.html (colored
board, candidate cards with live probabilities, scrub).
Languages
The server routes per request: Latin script → english, everything else →
multilingual (one extra ~0.7GB download, then cached). Turkish works today
(e.g. fatura/departman routing correct) but with lower confidence than
English — the same 0.6 gate applies. To pin a checkpoint, set LAYA_MODEL
in .mcp.json's env; the manifest ships unpinned because pinning
disables routing.
Install (step by step)
Requirements: Python 3.10+, pip, ~3GB disk (checkpoints), oh-my-pi.
pip install -r server/requirements.txt # laya==0.3.20, mcp<2, torch, transformers>=4.48,<5
omp plugin marketplace add F0Rextasy/omp-marketplace
omp plugin install laya-judge@forextasy # or: omp plugin link ./omp-laya-judge
omp plugin list # laya-judge should show ● enabledVerify inside omp:
Call the tool mcp__laya-judge__judge_info and reply with its exact JSON output.Expected: {"model": "laya/auto", "device": "cpu", "loaded": true, ...}.
First start imports torch and loads the checkpoint (~10s warm, longer on a
cold machine while both checkpoints download — the manifest ships
timeout: 600000 for exactly that), then ~0.16s per judgment.
Troubleshooting (all hit during development, all fixed in this repo):
transformers< 4.48 cannot load ModernBERT → pintransformers>=4.48,<5.mcp>=2renamed FastMCP → pinmcp<2.On Windows, torch init order matters → the sidecar binds first, loads the router on a background thread, and keeps answering
judge_info(loaded:false, thenstartup_errorinstead of hanging if it fails). Measured: ~20s toloaded:trueon CPU.python.EXE/py.EXEuppercase spawn failures → use the shipped.cmdwrapper.
Use
The plugin exposes mcp__laya-judge__judge(state, questions) plus the
laya-judge skill (triage → pre-filter → escalate). Question shapes mirror
oh-my-pi's eval judge():
choice:criteria: {label: rubric}→{choice, probabilities, confidence}bool: yes/no statement →{bool: P(true)}(layanoulhead is[false, true])score:criteria: [lowest … highest]→{score, legend, probabilities, confidence}
Every reply carries model (laya/english or laya/multilingual, derived
from laya's own routing block), per-call latency_ms, and laya's usage/
routing block so provenance survives in the transcript.
Tests
python -m unittest discover -s . -t . -p "test_*.py" # fast suite (CI)
LAYA_SLOW_TESTS=1 python -m unittest discover -s . -t . -p "test_*.py" # + model/stdio/routing/concurrencyThe fast suite pins the mapping layer, the HTTP sidecar contract, repo
hygiene (requirements/license/versions/manifest), and — via
tests/test_readme_lint.py + tests/test_calibration.py — that every
number on this page still matches the committed JSON. The slow suite needs
cached checkpoints: accuracy floor, MCP stdio handshake against a spawned
server, non-Latin routing, Turkish diacritics, and parallel judgments.
laya as the harness's judge (System One)
oh-my-pi resolves typed judgments through a judge role chain. A candidate
whose API is a judgment API is a native System One backend — the same slot
TypeSafe's jev and OpenRouter's Decisions API occupy. laya already speaks
that wire protocol (state + noul/choice/score → typed answers), so
the integration is not a shim: the sidecar answers the harness's own
decisions.
# one-time, user environment (Windows shown)
setx TYPESAFE_BASE_URL http://127.0.0.1:3777
setx TYPESAFE_API_KEY laya-local
# then, in ~/.omp/agent/config.yml:
# modelRoles:
# judge: typesafe/jev-latestGET /v1/models feeds oh-my-pi's discovery, so the sidecar's checkpoints
appear as native judge models:
$ omp models --kind=judge
typesafe (3)
├── english # discovered from the sidecar
├── jev-latest # bundled seed, baseUrl redirected to the sidecar
└── multilingual # discovered from the sidecarWith that set, laya answers the harness's decisions itself — the main model
never calls a judge tool. Verified live: the find cascade sent
systemone: 58 question(s) straight from oh-my-pi to the sidecar.
Where laya wins, measured. 1 question answers in 0.3–0.8s with zero API tokens; the same judgment through a chat model costs seconds and tokens.
Where it loses, measured. laya is ~100ms per question on CPU, so wide
batches lose to one remote call: 58 questions take 8.6s locally. The sidecar
chunks requests at 8 and merges them to stay inside the harness's 10s
judgment budget, but the find cascade issues many such batches per search,
which is why this setup ships find.enabled: off. Re-enable it on faster
hardware, or keep the cloud judge for that one path.
Because a ChainJudge timeout throws instead of falling through, the sidecar
refuses pathological requests (over LAYA_MAX_TOTAL_QUESTIONS, default 256)
with a non-transient 422 instead of hanging on the model lock.
You see the choices. The sidecar keeps a ring buffer of every System One
answer (GET /v1/decisions?since=N) and the turn-end card pulls what the
harness picked, so laya's own decisions read like any other call:
⚡ laya ▸ 2 decisions this turn — level=xhigh · stopped=0.23 · avg 254msWithout it the native judge role was silent — the harness calls laya directly, so no tool result ever reached the transcript.
The decision layer (harness-side)
The plugin no longer waits to be called. hooks/pre/laya-decide.ts asks the
sidecar directly at the points a keel-style selector would, and every gate is
fail-open: a timeout, a dead sidecar, or a malformed reply means no
decision, never a block.
# | Event | Decision | Acts at |
D1 |
| health-check, respawn the sidecar (throttled 60s) | always |
D2 |
| is this command destructive? | warn ≥0.6, block ≥0.85 |
D3 |
| mechanical work → | ≥0.6 |
D4 |
| which recovery path? + escalate effort? | ≥0.6 |
D5 |
| is the task actually finished? | P<0.6 → one continuation |
D6 |
| which context categories to preserve | ≥0.6 |
D7 |
| which project notes are relevant | ≥0.6 |
D8 |
| run focused checks before the full suite | ≥0.6, once per turn |
D2 and D8 only ask after a prefilter match (destructive/privileged/network/ kill/force-push patterns; bare full-suite runners), so ordinary commands cost zero. The sidecar self-exits after 10 idle minutes; the next prompt respawns it.
Measured gate behaviour (this is the honest part)
Probed against the shipped laya/english checkpoint on CPU, 16-command risk
set and the recovery/routing/note gates:
D2 carries real signal. Destructive commands score 0.43–0.78 (mean 0.64); ordinary read commands 0.12–0.34 (mean 0.22). Of four question framings tested, the long statement used here separated best.
D2 cannot rank deletions. Routine cleanup (
rm -rf build0.68) and catastrophe (rm -rf /0.75) overlap almost completely. No threshold separates them: 0.6 flags 5/8 routine commands, 0.8 catches nothing. That is why the block bar stays at 0.85 — the shipped checkpoint never reaches it, and the live behaviour is a caution, not a stop. Enriching the state with harness-extracted features made it worse (measured), so it is not done.D3, D4, D5, D8 are noise on this checkpoint — confidence 0.01–0.30, and D3 inverts ("find why auth drops sessions" routes to
sonic). They stay wired at the 0.6 gate, which they do not clear, so they cost one cheap round-trip and change nothing today. They become live if a stronger checkpoint is selected — re-measure before trusting them, exactly asLAYA_MODELpinning changes the calibration.
In short: the wiring is real and measured, and on the shipped checkpoint the only gate that acts is the destructive-command caution. Treat the rest as instrumented infrastructure, not as behaviour you should rely on.
Layout
.mcp.json— stdio server declaration (unpinned,timeout: 600000)server/bridge.py— thin FastMCP stdio bridge; no model importsserver/core.py— pure mapping (no torch), batch + unpack layerserver/sidecar.py— sole model owner and HTTP sidecar (POST /judge,GET /info)server/laya-judge.cmd— Windows launcherskills/laya-judge/SKILL.md— when to use / when to escalaterules/laya-auto.md— auto-routing rule (arithmetic excluded)commands/laya.md+hooks/pre/laya*.ts—/layastatus command, startup banner, live decision feedhooks/pre/laya-decide.ts+hooks/lib/record.ts— the harness-side decision layer (8 gates) and the shared decision queue behind the turn-end cardbenchmark/— reproducible accuracy + latency + gate proofdemo/— quiz/snake/tetris + maze (all seeded,*-stats.json), rendered through the shareddemo/style.pylooktests/— fast + slow suites;.github/workflows/test.ymlruns fastassets/—before-after.gif,benchmark.svg,quiz.gif,snake.gif,tetris.gif
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Decision-only prompt routing and firewall checks for local/cloud routing, PII and jailbreak risk.
- DatagoatOAuthio.datagoat
Governed decision engine: yes/no, score, choice and rank answers about cases, from past outcomes.
Deterministic contextual decision arbitration and action routing for autonomous software. Takes current state, context, or intent plus caller-supplied candidate actions, state transitions, routes, refusals, escalations, tools, or models and returns a deterministic ordered candidate field. Also provides persistent machine representations for memory, retrieval, indexing, and downstream coherence measurement.
A paid remote MCP for ZeroLang, built to return verdicts, receipts, usage logs, and audit-ready JSON
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables typed decisions (yes/no, choice, score) with transparent preflight checks, honest confidence reporting, and health monitoring.510 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to offload quick judgment calls like shipping readiness, file triage, and claim verification to a fast local MCP server with calibrated confidence and safe fallbacks.3MIT
- AlicenseNot gradedqualityCmaintenanceEnables MiMo Desktop to run a typed-decision engine through local stdio MCP, returning choice/score/noul predictions with calibrated confidence for auto-execution, LLM review, or escalation—without generating text.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to perform typed decisions (choice, score, noul) over text, emails, tickets, or JSON documents, with automatic language routing and source ranking across 100+ languages via a single forward pass.Apache 2.0