Skip to main content
Glama

omp-laya-judge

Local System-1 judge for oh-my-pi, powered by laya. Typed decisions (choice/bool/score) at mean 402ms (p50 238ms, p95 1108ms) on CPU, 0 LLM tokens burned, nothing leaves the machine.

before/after

The decision layer, mid-turn. Every card is a real answer from the local sidecar - the notes gate picking project notes, the edit gate judging a file change, the bash gate reading a destructive command, the recovery gate choosing what to do about a failed tool - drawn by the same row builder the TUI uses (hooks/lib/record.ts) and regenerated with python benchmark/live_feed_gif.py:

live feed

Measured head-to-head

Same 12 classification questions, real model via omp -p vs this plugin on CPU, no GPU:

LLM judge()

laya-judge

latency / question

18.4s (16–25s sampled)

mean 402ms (p50 238ms, p95 1108ms)

tokens / question

728

0

accuracy (12-case bench)

n/a (reference)

10/12

Reproduce everything below from committed artifacts:

python benchmark/run.py           # -> benchmark/results.json (schema 2)
python benchmark/calibrate.py     # -> benchmark/calibration.json (gate metrics)
python benchmark/chart.py         # -> assets/benchmark.svg
python benchmark/before_after.py  # -> assets/before-after.gif

Related MCP server: agent-fastpath

The 0.6 gate, honestly

Policy: answers with confidence ≥ 0.6 are auto-accepted, below escalates to the LLM. What that bought on the 12-case bench (benchmark/calibration.json, computed — not asserted):

  • auto-accept 9/12, 3 escalated to the LLM

  • 2 misses total: both caught by the gate, 0 escaped

  • false-accept 0% (0 of the 9 auto-accepts were wrong)

Both misses are score questions — the severity head lands on the right neighbourhood (2.4 for a 3) but not the exact bucket, and both sit under the gate, so they escalate instead of being trusted.

Arithmetic no longer reaches the model at all. The parity misses that used to escape at confidence 0.80/0.84 are answered by exact arithmetic in server/core.py:resolve_arithmetic, which claims a yes/no question only when the state carries exactly one distinct number and the instruction names an operation on it (parity, primality, divisibility, comparison). Everything ambiguous is still laya's.

Checkpoint A/B/C + random

(benchmark/compare.json, reproduced by benchmark/compare_models.py). These are raw checkpoint scores — the model answering alone, with no arithmetic resolver — which is why the english row reads 8 while the headline above reads 10: the two extra cases are parity, answered exactly by server/core.py rather than by the checkpoint.

EN 12-case

TR 4-case

mean conf

mean ms

english

8

2

0.68

210

typed-decisions

8

3

0.54

198

multilingual

6

2

0.76

111

random (coin flip)

6

0

—

0

TR sample is small (4); the honest read is laya >> chance (2–3 vs 0), not a checkpoint coronation.

benchmark

Which judge answers

The sidecar speaks POST /v1/systemone, so the model behind it is a setting, not a fork. Measured on the same 56 cases with one grading function (benchmark/compare_backends.py; routing 12, triage 20, severity 11, the real mixed JSON call 6, Turkish 7):

backend

correct

median

p95

route

triage

severity

mixed

Turkish

ECE

laya/english (default)

33/56

242ms

491ms

11/12

16/20

3/11

3/6

0/7

0.443

von 1.2

29/56

269ms

486ms

9/12

15/20

4/11

1/6

0/7

0.348

laya/multilingual

26/56

88ms

173ms

9/12

13/20

3/11

1/6

0/7

0.480

Switch backends by environment; the hook never changes your models for you.

LAYA_BACKEND=systemone LAYA_UPSTREAM_URL=http://127.0.0.1:8000 ...   # von, or jev
LAYA_BACKEND=systemone LAYA_UPSTREAM_URL=https://api.typesafe.ai/v1/systemone \
             LAYA_UPSTREAM_KEY_ENV=TYPESAFE_API_KEY ...             # jev, your key
LAYA_MODEL=multilingual ...                                        # faster, less sure

LAYA_UPSTREAM_KEY_ENV names the variable holding the key, so the key itself never lands in a config file. /info reports which backend is actually answering, and arithmetic stays local on every backend because it is exact and free everywhere.

Two honest limits from that table: Turkish is 0/7 on all three - including laya's multilingual checkpoint - so the English checkpoints are not usable for non-English text, and the layer should not gate on it. And laya's answer_confidence is the field to threshold, not the confidence it also returns; the gates use the calibrated one, which raised local auto-accept from 7/12 to 9/12 with false-accept still at 0%.

Demos

quiz game

Grounded quiz (python demo/quiz.py): 6/8, mean 282ms, 0 tokens — both misses came in under 0.2 confidence, exactly the cases the escalate rule covers.

snake

Snake (python demo/snake.py): score 12, 300 moves, alive, 12 shield interventions, ~0.9s/move on CPU, 0 tokens. Watch every decision in the browser: open web/snake.html (canvas replay with live probability bars, play/pause/speed/scrub); quiz at web-quiz.html. Recipe (from laya-mlx): a deterministic planner describes each direction over a tiny state (Safe route: yes. Food reachable: …), a safety shield executes the best SAFE move. Short parallel criteria are the whole trick — greedy end-to-end choice without the planner scores 0 (see demo/snake.py header for the measured failure modes). Honest calibration: per-move confidences on game states run at noise level (0.005-0.05; the quiz gets 0.2-1.0 on text) — navigation is planner + shield with laya ranking, the same division of labor as laya-mlx. The maze variant (demo/maze.py) is a documented negative: corridors trap the noise walk (1 crumb/200 steps).

tetris

Tetris (python demo/tetris.py): score 5720, 43 lines, 120 pieces, alive, ~1s/placement on CPU, 0 tokens. Per piece the planner enumerates every legal landing, shortlists 6 by classic features (lines, holes, height, bumpiness), laya picks one in a batched call, and a shield keeps the stack out of the top-4 danger zone. Replay in browser: web/tetris.html (colored board, candidate cards with live probabilities, scrub).

Languages

The server routes per request: Latin script → english, everything else → multilingual (one extra ~0.7GB download, then cached). Turkish works today (e.g. fatura/departman routing correct) but with lower confidence than English — the same 0.6 gate applies. To pin a checkpoint, set LAYA_MODEL in .mcp.json's env; the manifest ships unpinned because pinning disables routing.

Install (step by step)

Requirements: Python 3.10+, pip, ~3GB disk (checkpoints), oh-my-pi.

pip install -r server/requirements.txt   # laya==0.3.20, mcp<2, torch, transformers>=4.48,<5
omp plugin marketplace add F0Rextasy/omp-marketplace
omp plugin install laya-judge@forextasy  # or: omp plugin link ./omp-laya-judge
omp plugin list                          # laya-judge should show ● enabled

Verify inside omp:

Call the tool mcp__laya-judge__judge_info and reply with its exact JSON output.

Expected: {"model": "laya/auto", "device": "cpu", "loaded": true, ...}. First start imports torch and loads the checkpoint (~10s warm, longer on a cold machine while both checkpoints download — the manifest ships timeout: 600000 for exactly that), then ~0.16s per judgment.

Troubleshooting (all hit during development, all fixed in this repo):

  • transformers < 4.48 cannot load ModernBERT → pin transformers>=4.48,<5.

  • mcp>=2 renamed FastMCP → pin mcp<2.

  • On Windows, torch init order matters → the sidecar binds first, loads the router on a background thread, and keeps answering judge_info (loaded:false, then startup_error instead of hanging if it fails). Measured: ~20s to loaded:true on CPU.

  • python.EXE/py.EXE uppercase spawn failures → use the shipped .cmd wrapper.

Use

The plugin exposes mcp__laya-judge__judge(state, questions) plus the laya-judge skill (triage → pre-filter → escalate). Question shapes mirror oh-my-pi's eval judge():

  • choice: criteria: {label: rubric} → {choice, probabilities, confidence}

  • bool: yes/no statement → {bool: P(true)} (laya noul head is [false, true])

  • score: criteria: [lowest … highest] → {score, legend, probabilities, confidence}

Every reply carries model (laya/english or laya/multilingual, derived from laya's own routing block), per-call latency_ms, and laya's usage/ routing block so provenance survives in the transcript.

Tests

python -m unittest discover -s . -t . -p "test_*.py"             # fast suite (CI)
LAYA_SLOW_TESTS=1 python -m unittest discover -s . -t . -p "test_*.py"   # + model/stdio/routing/concurrency

The fast suite pins the mapping layer, the HTTP sidecar contract, repo hygiene (requirements/license/versions/manifest), and — via tests/test_readme_lint.py + tests/test_calibration.py — that every number on this page still matches the committed JSON. The slow suite needs cached checkpoints: accuracy floor, MCP stdio handshake against a spawned server, non-Latin routing, Turkish diacritics, and parallel judgments.

laya as the harness's judge (System One)

oh-my-pi resolves typed judgments through a judge role chain. A candidate whose API is a judgment API is a native System One backend — the same slot TypeSafe's jev and OpenRouter's Decisions API occupy. laya already speaks that wire protocol (state + noul/choice/score → typed answers), so the integration is not a shim: the sidecar answers the harness's own decisions.

# one-time, user environment (Windows shown)
setx TYPESAFE_BASE_URL http://127.0.0.1:3777
setx TYPESAFE_API_KEY laya-local
# then, in ~/.omp/agent/config.yml:
#   modelRoles:
#     judge: typesafe/jev-latest

GET /v1/models feeds oh-my-pi's discovery, so the sidecar's checkpoints appear as native judge models:

$ omp models --kind=judge
typesafe (3)
├── english        # discovered from the sidecar
├── jev-latest     # bundled seed, baseUrl redirected to the sidecar
└── multilingual   # discovered from the sidecar

With that set, laya answers the harness's decisions itself — the main model never calls a judge tool. Verified live: the find cascade sent systemone: 58 question(s) straight from oh-my-pi to the sidecar.

Where laya wins, measured. 1 question answers in 0.3–0.8s with zero API tokens; the same judgment through a chat model costs seconds and tokens.

Where it loses, measured. laya is ~100ms per question on CPU, so wide batches lose to one remote call: 58 questions take 8.6s locally. The sidecar chunks requests at 8 and merges them to stay inside the harness's 10s judgment budget, but the find cascade issues many such batches per search, which is why this setup ships find.enabled: off. Re-enable it on faster hardware, or keep the cloud judge for that one path.

Because a ChainJudge timeout throws instead of falling through, the sidecar refuses pathological requests (over LAYA_MAX_TOTAL_QUESTIONS, default 256) with a non-transient 422 instead of hanging on the model lock.

You see the choices. The sidecar keeps a ring buffer of every System One answer (GET /v1/decisions?since=N) and the turn-end card pulls what the harness picked, so laya's own decisions read like any other call:

⚡ laya ▸ 2 decisions this turn — level=xhigh · stopped=0.23 · avg 254ms

Without it the native judge role was silent — the harness calls laya directly, so no tool result ever reached the transcript.

The decision layer (harness-side)

The plugin no longer waits to be called. hooks/pre/laya-decide.ts asks the sidecar directly at the points a keel-style selector would, and every gate is fail-open: a timeout, a dead sidecar, or a malformed reply means no decision, never a block.

#

Event

Decision

Acts at

D1

session_start / before_agent_start

health-check, respawn the sidecar (throttled 60s)

always

D2

tool_call (bash)

is this command destructive?

warn ≥0.6, block ≥0.85

D3

tool_call (task)

mechanical work → sonic?

≥0.6

D4

tool_result (error)

which recovery path? + escalate effort?

≥0.6

D5

session_stop

is the task actually finished?

P<0.6 → one continuation

D6

session.compacting

which context categories to preserve

≥0.6

D7

before_agent_start

which project notes are relevant

≥0.6

D8

tool_call (bash)

run focused checks before the full suite

≥0.6, once per turn

D2 and D8 only ask after a prefilter match (destructive/privileged/network/ kill/force-push patterns; bare full-suite runners), so ordinary commands cost zero. The sidecar self-exits after 10 idle minutes; the next prompt respawns it.

Measured gate behaviour (this is the honest part)

Probed against the shipped laya/english checkpoint on CPU, 16-command risk set and the recovery/routing/note gates:

  • D2 carries real signal. Destructive commands score 0.43–0.78 (mean 0.64); ordinary read commands 0.12–0.34 (mean 0.22). Of four question framings tested, the long statement used here separated best.

  • D2 cannot rank deletions. Routine cleanup (rm -rf build 0.68) and catastrophe (rm -rf / 0.75) overlap almost completely. No threshold separates them: 0.6 flags 5/8 routine commands, 0.8 catches nothing. That is why the block bar stays at 0.85 — the shipped checkpoint never reaches it, and the live behaviour is a caution, not a stop. Enriching the state with harness-extracted features made it worse (measured), so it is not done.

  • D3, D4, D5, D8 are noise on this checkpoint — confidence 0.01–0.30, and D3 inverts ("find why auth drops sessions" routes to sonic). They stay wired at the 0.6 gate, which they do not clear, so they cost one cheap round-trip and change nothing today. They become live if a stronger checkpoint is selected — re-measure before trusting them, exactly as LAYA_MODEL pinning changes the calibration.

In short: the wiring is real and measured, and on the shipped checkpoint the only gate that acts is the destructive-command caution. Treat the rest as instrumented infrastructure, not as behaviour you should rely on.

Layout

  • .mcp.json — stdio server declaration (unpinned, timeout: 600000)

  • server/bridge.py — thin FastMCP stdio bridge; no model imports

  • server/core.py — pure mapping (no torch), batch + unpack layer

  • server/sidecar.py — sole model owner and HTTP sidecar (POST /judge, GET /info)

  • server/laya-judge.cmd — Windows launcher

  • skills/laya-judge/SKILL.md — when to use / when to escalate

  • rules/laya-auto.md — auto-routing rule (arithmetic excluded)

  • commands/laya.md + hooks/pre/laya*.ts — /laya status command, startup banner, live decision feed

  • hooks/pre/laya-decide.ts + hooks/lib/record.ts — the harness-side decision layer (8 gates) and the shared decision queue behind the turn-end card

  • benchmark/ — reproducible accuracy + latency + gate proof

  • demo/ — quiz/snake/tetris + maze (all seeded, *-stats.json), rendered through the shared demo/style.py look

  • tests/ — fast + slow suites; .github/workflows/test.yml runs fast

  • assets/ — before-after.gif, benchmark.svg, quiz.gif, snake.gif, tetris.gif

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables typed decisions (yes/no, choice, score) with transparent preflight checks, honest confidence reporting, and health monitoring.
    510 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to offload quick judgment calls like shipping readiness, file triage, and claim verification to a fast local MCP server with calibrated confidence and safe fallbacks.
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MiMo Desktop to run a typed-decision engine through local stdio MCP, returning choice/score/noul predictions with calibrated confidence for auto-execution, LLM review, or escalation—without generating text.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to perform typed decisions (choice, score, noul) over text, emails, tickets, or JSON documents, with automatic language routing and source ranking across 100+ languages via a single forward pass.
    Apache 2.0