Skip to main content
Glama

laya-mcp

把 Laya System-1 决策模型封装成一个工程化的 MCP Server(FastAPI 宿主)

一次前向传播回答一组「带类型的问题」,返回校准概率 + 可直接分支执行的决策契约, 让任意支持 MCP 的 Agent(Claude Code / Codex / Cursor / 自研 Agent)都能在毫秒级完成自动决策。

Python FastAPI MCP Laya


目录


Related MCP server: laya-mcp

这是什么

laya-mcp 把 Laya(一个 System-1 决策模型族) 包装成一个常驻服务,同时对外暴露两种接口:

接口

地址

面向谁

MCP Streamable HTTP

http://127.0.0.1:8077/mcp

任意 MCP 客户端 / Agent

REST API

http://127.0.0.1:8077/v1/...

脚本、看板、健康检查、非 MCP 调用方

OpenAPI 文档

http://127.0.0.1:8077/docs

人

它不是「又一个 LLM 接口」。Laya 是非生成式的:一次 forward pass 回答一组结构化问题, 输出的是校准过的概率分布,而不是一段文本。这使得它可以在 10~50ms 内做出「该走哪条路」的判断, 成本约为调用一次完整 LLM 的千分之一。

与传统做法对比

laya-test.py(见 examples/quickstart_router.py)演示了最小用法: 加载 Router → 构造问题 → router.predict() → 用阈值手写 if。

这个工程把它产品化了:

关注点

laya-test.py

laya-mcp

模型驻留

每次进程启动重新加载

常驻单例 + 启动预热(warmup)

阈值策略

每个调用方自己写 if conf >= 0.8

决策契约,由服务统一输出 grade

并发

单线程阻塞

线程池 + 事件循环不阻塞 + 结果缓存

接入方式

只能进程内 Python 调用

MCP / REST / CLI 三种入口共用同一套逻辑

可观测性

print

结构化日志 + 指标(计数 / 延迟分位) + /healthz

错误处理

直接抛异常

统一错误码 + HTTP 状态映射

配置

硬编码

LAYAMCP_* 环境变量 / .env

Agent 可发现性

无

工具自描述 + resources 知识库 + prompt 模板


核心概念:决策契约(Contract)

这是整个项目最重要的设计。 模型给出的是概率,但 Agent 需要的是策略: 「我现在可以直接执行吗?还是必须找人确认?」

原始概率无法直接回答这个问题——两个策略 0.79 / 0.78 时,最高分看起来很高,但它不是一个决策。 因此每次调用都会额外返回一个 contract 块,把概率翻译成 4 种 grade,Agent 直接分支即可:

grade

含义

Agent 应该做什么

✅ auto_execute

置信度越过阈值,且与次优有明显差距

直接执行,不要再花 token 重新推理

⚠️ escalate

高置信但存在近距平局,或问题本身语义上就该由人拍板

不要仅凭模型答案行动,走完整 LLM 分析或请人确认

🛑 ask_human

低置信 + 高风险场景

必须转人工

🔁 reconsider

低置信 + 低风险

可以重试:换更精确的 criteria、增加候选、或交给 LLM

grade 由三个信号共同决定(而不是单一置信度)

信号

作用

为什么单独看它不行

confidence

校准后的置信度

候选数量多时 softmax 会被抬高,必须用 Laya 的温度校准值

margin

top1 − top2

0.79 / 0.78 的两条相邻策略,不管最高分多高都不构成决策

needs_review

模型自带的「需人工复核」头

对语义上就要求人拍板的问题,高置信也依然是 escalate

契约结构示例

{
  "grade": "auto_execute",
  "autonomous": true,
  "requires_human": false,
  "threshold": 0.8,
  "margin_threshold": 0.05,
  "decisions": {
    "selected_strategy": {
      "value": "redis_cache",
      "confidence": 0.9231,
      "margin": 0.41,
      "grade": "auto_execute",
      "reason": "confidence 0.92 clears 0.80 and beats the runner-up by 0.41",
      "requires_human": false
    }
  },
  "grade_counts": { "auto_execute": 1, "escalate": 0, "ask_human": 0, "reconsider": 0 },
  "reasons": [],
  "summary": "all 1 answers clear the auto-execute bar; proceed",
  "routing": { "model": "english", "reason": "English Latin text" }
}

聚合规则:一个被卡住的问题会卡住整个请求。 整体 grade 取所有问题中最「保守」的那个 (优先级 ask_human > escalate > reconsider > auto_execute), 因为对 Agent 而言「部分可执行」比「整体不可执行」更危险。

Agent 典型分支伪代码

result = mcp.call("decide", state=diff, preset="code_change_risk")
grade = result["contract"]["grade"]

if grade == "auto_execute":
    merge_pull_request()                       # 直接干,零推理成本
elif grade == "escalate":
    run_full_llm_review(diff)                  # 花 token 深挖
elif grade == "ask_human":
    notify_reviewer(result["contract"]["reasons"])
else:  # reconsider
    result = mcp.call("decide", state=diff, preset="code_change_risk",
                      threshold="critical")    # 换阈值 / 换问题重试

为什么值得用

  • 大幅降低 Agent 决策成本:把「要不要做 / 走哪条路」交给 System-1 模型,避免为此唤起完整 LLM 的 CoT。

  • 不引入幻觉风险:非生成式,输出是封闭候选集上的概率,不会「编造」答案。

  • 策略集中在服务端:阈值、margin、风险等级都是配置,不用在每个 Agent 里重复实现。

  • 对 Agent 友好:工具自带详细 description 与 structuredContent,另有 laya://guide 等资源做知识注入。

  • 一套逻辑三种入口:MCP / REST / CLI 共用 laya_mcp.core,不会出现行为漂移。

  • 生产可用:健康检查、指标、结构化日志、超时与体积上限、线程安全、优雅降级(模型加载失败时 /healthz 返回 degraded 而非崩溃)。


架构

                        ┌─────────────────────────────┐
   MCP Agent  ────────► │  /mcp  StreamableHTTP       │
 (Claude Code /         │  (9 tools / 4 resources /   │
  Codex / Cursor)       │   2 prompts)                │
                        ├─────────────────────────────┤
   脚本 / 看板 ────────► │  /v1/...  REST (FastAPI)    │
                        └──────────────┬──────────────┘
                                       │  共用同一套实现
                        ┌──────────────▼──────────────┐
                        │        laya_mcp.core        │
                        │  settings / schema / preset │
                        │  engine(常驻+池+缓存)         │
                        │  contract(决策契约) / metrics │
                        └──────────────┬──────────────┘
                                       │
                        ┌──────────────▼──────────────┐
                        │  laya.Router (System-1)     │
                        │  english / multilingual /   │
                        │  typed-decisions            │
                        └─────────────────────────────┘

目录结构

laya-mcp/
├── src/laya_mcp/
│   ├── core/                 # 与传输方式无关的领域层
│   │   ├── settings.py       #   LAYAMCP_* 配置、阈值等级
│   │   ├── schema.py         #   公开问题 schema ↔ Laya 原生格式
│   │   ├── presets.py        #   5 个内置问题集
│   │   ├── contract.py       #   ★ 概率 → grade 的决策契约
│   │   ├── engine.py         #   单例 Router + 线程池 + LRU 缓存
│   │   ├── errors.py         #   统一错误分类与 HTTP 映射
│   │   └── metrics.py        #   计数 / 延迟分位
│   ├── api/                  # FastAPI 层
│   │   ├── app.py            #   create_app(),挂载 /mcp 与 /v1
│   │   ├── routes.py         #   REST 端点(不含决策逻辑)
│   │   ├── schemas.py        #   请求 / 响应模型
│   │   └── dependencies.py   #   DI:engine / settings / metrics
│   ├── mcp_server/           # MCP 层
│   │   ├── server.py         #   工具 / 资源 / Prompt 注册
│   │   └── questions.py      #   问题解析(inline / preset / 快捷参数)
│   ├── cli.py                # 命令行入口
│   └── logging_config.py     # 结构化日志
├── examples/quickstart_router.py   # 原始 laya-test.py 用法
├── tests/
└── pyproject.toml

关键实现说明

  • 模型只加载一次:DecisionEngine 进程级单例。启动时预热(LAYAMCP_WARMUP=true), 之后每次请求都是热态,毫秒级返回。

  • 不阻塞事件循环:推理跑在有界线程池里(LAYAMCP_WORKER_THREADS), torch 内部会释放 GIL,因此 CPU 上并发也有收益。

  • 结果缓存:对相同 (state, questions) 做进程内 LRU 缓存(默认 256 条), Agent 重试同样的问题几乎零成本;响应中 cached=true 可辨别。

  • MCP 挂载方式(容易踩坑):不能用 streamable_http_app() 再 mount 到 FastAPI—— Starlette 不会为子应用执行 lifespan,session manager 起不来,/mcp 会全部失败。 正确做法是在本应用的 lifespan 里进入 session manager:

    manager = StreamableHTTPSessionManager(app=server._lowlevel_server, ...)
    app.mount("/mcp", StreamableHTTPASGIApp(manager))
    # lifespan: async with manager.run(): yield

    参见 src/laya_mcp/api/app.py 的模块注释。

  • MCP SDK 版本:基于 mcp 2.x(from mcp.server.mcpserver import MCPServer)。 与 1.x 的 mcp.server.fastmcp.FastMCP 不兼容,升级时注意导入路径变化。


快速开始

环境要求

  • Python 3.13+

  • uv(推荐)

  • 磁盘约 3 GB(三个 checkpoint 合计约 2.2 GB)

  • CPU 即可运行;有 CUDA / Apple MPS 会自动加速

  • 首次运行需要联网从 HuggingFace 下载权重,之后可完全离线

安装

git clone <your-repo-url> laya-mcp
cd laya-mcp
uv sync                       # 创建 .venv 并安装依赖

自检

uv run laya-mcp doctor

输出示例:

laya-mcp 0.2.0
python: 3.13.x
settings: models=['english', 'multilingual'] device=auto port=8077
laya: 0.3.5 -> .../laya/__init__.py
torch: 2.x.x cuda=False
HF 缓存: 2.20 GB,包含 laya 检查点: True
fastapi: ok
uvicorn: ok
mcp: ok
pydantic_settings: ok
结论: 环境可用

启动服务

uv run laya-mcp serve
INFO  laya_mcp.engine  loading laya router  models=['english','multilingual'] device=cpu
INFO  laya_mcp.engine  laya router ready  loaded=['english','multilingual']
INFO  laya_mcp.api     engine ready in 67.8s
INFO  mcp.server...    StreamableHTTP session manager started
Uvicorn running on http://127.0.0.1:8077

⏳ 首次启动约 60~90 秒(在 CPU 上加载两个 checkpoint)。之后所有请求都是热态, 单次决策通常 10~50ms。

第一条请求

curl -s http://127.0.0.1:8077/v1/decide \
  -H 'content-type: application/json' \
  -d '{
    "state": "Fix a typo in README, one file, no interface change",
    "preset": "code_change_risk"
  }' | python -m json.tool

或者完全不启服务,直接一次性决策:

uv run laya-mcp ask \
  --state "Fix a typo in README, one file, no interface change" \
  --preset code_change_risk

接入 MCP Agent

服务启动后,任何支持 MCP 的客户端都可以直接接入。关键信息只有一条:

MCP Streamable HTTP 端点:http://127.0.0.1:8077/mcp

Claude Code

claude mcp add --transport http laya http://127.0.0.1:8077/mcp

Codex / Cursor / 通用 mcp.json

{
  "mcpServers": {
    "laya": {
      "type": "streamable-http",
      "url": "http://127.0.0.1:8077/mcp"
    }
  }
}

只给单个 Agent 用(stdio,无 HTTP)

不想起 HTTP 服务时,可以用 stdio 传输,由客户端自己拉起进程:

{
  "mcpServers": {
    "laya": {
      "command": "uv",
      "args": ["run", "--directory", "/abs/path/to/laya-mcp", "laya-mcp", "serve", "--transport", "stdio"]
    }
  }
}

stdio 模式下进程由客户端按需拉起,建议启动后先让 Agent 调一次 preload_warmup, 否则第一次决策会等待模型加载。HTTP 模式则无需关心——服务启动时已预热。

让 Agent「自动决策」的推荐接入方式

  1. 在系统提示里注入用法。服务已经在 MCP instructions 里写了使用指南, 另外可以让 Agent 读取 laya://guide 资源获取完整的问题编写规范。

  2. 强制 Agent 先查契约。约定:任何有副作用的动作前,先调用 decide / should_escalate, 并只在 contract.grade == "auto_execute" 时直接执行。

  3. 按风险等级传 threshold。删库、发钱、改线上配置这类动作,传 "high" 或 "critical"。

  4. 用 preset 起步,再按需内联问题。preset 与内联 questions 会合并到同一次前向传播, 边际成本几乎为零。


MCP 工具参考

共 9 个工具,全部返回结构化结果(文本 + structuredContent 字段同名)。

工具

用途

是否加载模型

典型耗时

decide

主力工具:对 state 回答一组带类型的问题,返回答案 + 契约

是

10~50ms

decide_yes_no

最省成本:单个是非问题 + 契约

是

~10ms

decide_shortlist

候选标签极多(几十~上百)时,先 embedding 粗筛 top-k 再决策

是

20~80ms

route_request

只做路由:判断语言 / 脚本 / 工作流,返回该用哪个 checkpoint 及原因

否

<1ms

list_presets

列出全部内置问题集及其 question id

否

<1ms

describe_questions

回显并校验一组问题(id、kind、候选、等级),不推理

否

<1ms

should_escalate

一站式:给定动作,直接回答「是否需要人工介入」

是

10~50ms

server_status

就绪状态、已加载 checkpoint、缓存条数、延迟分位

否

<1ms

preload_warmup

强制加载 checkpoint(仅 stdio 或关闭预热时需要)

是(慢)

30~90s

调用顺序建议(先便宜后昂贵)

route_request  <  list_presets / describe_questions  <  decide_yes_no  <  decide  <  decide_shortlist
   (<1ms)                  (<1ms)                      (~10ms)         (10~50ms)      (20~80ms)

decide 参数速查

参数

类型

说明

state

string | object | array

必填。模型真正读到的上下文:代码 / diff / 工单 / trace / 工具返回值。放原始材料,不要放摘要,摘要会掉准确率

questions

object | array

内联问题集,见 Question Schema

preset

string

内置问题集名;可与 questions 合并

threshold

number | string

auto_execute 所需置信度。数字 ∈ [0,1],或等级 trivial/low/normal/high/critical

model

string

强制 checkpoint:english / multilingual / typed-decisions

lang

string

提示 state 的语言(如 zh / en),跳过语言检测

task

string

强制使用 typed-decisions 的四个工作流之一

use_cache

bool

是否复用相同请求的缓存结果,默认 true

返回值结构

{
  "contract": { /* 见上文「决策契约」 */ },
  "grade": "auto_execute",          // 整体 grade,Agent 只需看这个
  "autonomous": true,               // 是否可完全自主执行
  "requires_human": false,
  "summary": "all 5 answers clear the auto-execute bar; proceed",
  "answers": {
    "blast_radius": {
      "kind": "choice",
      "value": "single_file",
      "confidence": 0.9642,
      "margin": 0.8713,
      "grade": "auto_execute",
      "reason": "confidence 0.96 clears 0.80 and beats the runner-up by 0.87",
      "requires_human": false,
      "probabilities": { "single_file": 0.9642, "module": 0.0929, "...": 0.0 },
      "raw": { /* Laya 原始返回,未加工 */ }
    }
  },
  "routing": { "model": "english", "reason": "English Latin text", "repo": "convaiinnovations/laya" },
  "latency_ms": 31.4,
  "cached": false,
  "usage": {}
}

Resources 与 Prompts

除了工具,服务还暴露 MCP Resources(可被 Agent 当知识库读取)与 Prompts(可复用模板)。

Resources

URI

内容

laya://guide

问题编写指南、kind 说明、grade 语义(强烈建议 Agent 读取)

laya://schema

decide / describe_questions 接受的问题 JSON Schema

laya://presets

全部内置问题集(JSON)

laya://status

实时就绪状态与延迟,等同 server_status

Prompts

名称

用途

decide_before_acting

把「打算执行的动作」转成一组问题集的填空模板

triage_message

对一条用户消息做分诊(意图 / 紧急度 / 情绪 / 流失风险)


内置问题集(Preset)

Preset 是问题的骨架,不是答案。它们覆盖了 Agent 最常遇到的五类场景:

名称

问题数

适用场景

包含的 question id

code_change_risk

5

代码改动 前置风险评审:爆炸半径、可回滚性、数据影响、测试覆盖、是否需人工

blast_radius, reversible, data_impact, test_coverage, needs_human_review

support_triage

4

客服工单分诊:意图、紧急度、情绪、流失风险

intent, is_urgent, frustration, churn_risk

input_guard

5

不可信输入的护栏:越狱、提示注入、敏感数据、危害程度、话题

jailbreak, prompt_injection, sensitive_data, harm_severity, topic

routing

4

模型/算力路由:难度、领域、是否需要工具、敏感性

difficulty, domain, needs_tools, is_sensitive

typed_decisions_agent_trace

5

Agent trace 可观测性(与 typed-decisions checkpoint 的工作流对齐)

action, needs_review, outcome, risk, urgency

# 查看全部 preset
uv run laya-mcp presets
# 查看单个 preset 的完整问题定义
curl -s http://127.0.0.1:8077/v1/presets/code_change_risk | python -m json.tool

preset + 自定义问题可以合并,一次前向传播全部回答:

{
  "state": "<diff 或工单>",
  "preset": "code_change_risk",
  "questions": {
    "target_release": {
      "kind": "choice",
      "instructions": "这个改动更适合进哪个版本?",
      "criteria": {
        "hotfix": "紧急修复,必须立刻发",
        "next_minor": "下个小版本即可",
        "backlog": "不急,排入待办"
      }
    }
  }
}

问题(Question)Schema

Laya 原生词汇是 choice / score / noul(且使用 t / ins / crit 短键)。 本工程定义了一套稳定、自描述的公开 schema,调用方不需要了解底层格式:

{
  "question_id": {
    "kind": "choice",                  // choice | score | yes_no
    "instructions": "选哪个方案?",       // 给模型的指令
    "criteria": {                      // 见下方「三种 kind」
      "redis_cache": "适合读多写少、容忍秒级延迟",
      "db_index": "适合查询条件固定、数据量中等"
    },
    "weight": 1.0,                     // 可选,建议性权重
    "description": "..."               // 可选,仅在罗列时展示
  }
}

三种 kind

kind

对应 Laya 类型

criteria 格式

说明

choice

choice

{标签: 何时选它}

N 选 1。每个候选的适用条件要写具体、可区分

score

score

["等级0", "等级1", ...]

有序等级列表,返回等级下标

yes_no

noul

{"true": "应该成立时的陈述"}

是非问题。陈述句要正向表述(true 表示「可以继续」)

编写要点(直接影响准确率)

  1. 一个问题只问一件事。不要问「这个改动安全吗」,而是拆成爆炸半径 / 可回滚 / 数据影响。

  2. 候选之间必须互斥且穷尽。留一个 other 兜底往往比硬选更好。

  3. criteria 写判定条件,不写形容词。✅「涉及共享 schema,会被其他团队消费」 ❌「影响很大」。

  4. state 放原始材料,代码、diff、完整工单原文,不要放你自己的总结。

  5. 问题数量 2~6 个最佳(上限由 LAYAMCP_MAX_QUESTIONS 控制,默认 32), 每个问题的边际成本极低——多问一个几乎不增加耗时。

  6. 不要用 Laya 做生成任务。它只做分类/打分,不是写作模型。

校验问题集

写好后先让服务回显校验,避免把无效问题发到模型:

curl -s http://127.0.0.1:8077/v1/questions/describe \
  -H 'content-type: application/json' \
  -d '{"preset": "input_guard"}' | python -m json.tool

REST API

Base path:/v1。完整交互式文档见 /docs。

决策

方法

路径

说明

POST

/v1/decide

回答一组带类型的问题(等同 MCP decide)

POST

/v1/decide/yes-no

单个是非问题

POST

/v1/decide/shortlist

大规模候选:embedding 粗筛 + 决策

POST

/v1/route

只做路由,不推理(微秒级)

目录 / 运维

方法

路径

说明

GET

/v1/presets

全部内置问题集(含问题定义)

GET

/v1/presets/{name}

单个问题集(规范化后的最终形态)

POST

/v1/questions/describe

校验并回显一组问题

GET

/v1/models

可用 checkpoint 及加载状态

GET

/v1/healthz

存活检查(模型加载失败返回 200 + degraded)

GET

/v1/status

引擎状态、阈值、缓存、延迟分位

GET

/v1/metrics

原始计数与延迟数据(供 Prometheus 等抓取)

GET

/v1/version

服务版本与 laya 版本

示例

# 1) 是非问题:这个动作能直接执行吗?
curl -s http://127.0.0.1:8077/v1/decide/yes-no \
  -H 'content-type: application/json' \
  -d '{
    "state": {"action": "DROP TABLE users", "env": "production"},
    "statement": "this action is safe to run without human review",
    "threshold": "critical"
  }' | python -m json.tool

# 2) 中文工单分诊(自动路由到 multilingual checkpoint)
curl -s http://127.0.0.1:8077/v1/decide \
  -H 'content-type: application/json' \
  -d '{"state": "客户反馈发票被重复扣款,语气很急", "preset": "support_triage"}' \
  | python -m json.tool

# 3) 只路由,不推理
curl -s http://127.0.0.1:8077/v1/route \
  -H 'content-type: application/json' \
  -d '{"state": "给我翻译一下这段中文"}' | python -m json.tool
# {"model":"multilingual","reason":"non-Latin script (han, 100% of letters); ..."}

统一错误格式

所有错误都是同一形状,并带 request_id 便于排查:

{
  "error": "invalid_question",
  "message": "at most 32 questions per call",
  "details": { "count": 40, "limit": 32 },
  "request_id": "0f3a9c1b2d4e5f60"
}

命令行 CLI

uv run laya-mcp --help

命令

说明

laya-mcp serve

启动服务(默认 FastAPI + MCP Streamable HTTP)

laya-mcp serve --transport stdio

仅提供 MCP,走标准输入输出,适合本地 Agent

laya-mcp ask --state ... --choices ...

一次性决策,不启服务(等价于 laya-test.py 的用法)

laya-mcp presets

列出内置问题集

laya-mcp tools

列出 MCP 工具(不加载模型)

laya-mcp doctor

检查环境、依赖与 checkpoint 缓存

serve 常用参数

uv run laya-mcp serve \
  --host 0.0.0.0 \
  --port 8077 \
  --models english,multilingual,typed-decisions \
  --device auto \
  --log-level INFO

参数

说明

--transport

streamable-http(默认,含 REST)/ stdio(仅 MCP)/ sse

--host / --port

监听地址与端口(默认 127.0.0.1:8077)

--models

常驻 checkpoint,逗号分隔,或 all

--device

auto / cpu / cuda / mps

--reload

开发热重载(仅 HTTP)

--no-warmup

启动时不加载模型(首次请求会变慢)

ask 示例

# N 选 1(--choices 是快捷写法)
uv run laya-mcp ask \
  --state "高并发下数据库查询变慢,Query 全表扫描,QPS 5000/s" \
  --instructions "根据性能瓶颈选择最合适的优化方案" \
  --choices "redis_cache,db_index,read_write_split,search_engine"

# 读取文件作为 state(@ 前缀)
uv run laya-mcp ask --state @./diff.patch --preset code_change_risk

# 输出完整 JSON(含契约)
uv run laya-mcp ask --state "..." --preset input_guard --json

# 是非问题
uv run laya-mcp ask --state "<发布计划>" \
  --statement "this can ship without human approval"

ask 的可读输出示例:

5/5 answers fail the auto-execute bar (...): ask a human before acting

blast_radius: single_file  (choice, confidence 0.64, margin 0.82, ask_human)
needs_human_review: True   (yes_no, confidence 0.60, margin 0.21, ask_human)

checkpoint: english - English Latin text

配置项(环境变量)

所有配置项前缀为 LAYAMCP_,也可写在项目根目录的 .env 里 (见 .env.example)。

服务

变量

默认值

说明

LAYAMCP_HOST

127.0.0.1

监听地址

LAYAMCP_PORT

8077

监听端口

LAYAMCP_TRANSPORT

streamable-http

streamable-http / stdio / sse

LAYAMCP_MCP_PATH

/mcp

MCP 挂载路径

LAYAMCP_API_PREFIX

/v1

REST 前缀

LAYAMCP_DOCS_ENABLED

true

是否开启 /docs

推理

变量

默认值

说明

LAYAMCP_MODELS

english,multilingual

常驻 checkpoint,逗号分隔或 all

LAYAMCP_DEVICE

auto

auto / cpu / cuda / mps(auto 优先 CUDA → MPS → CPU)

LAYAMCP_DEFAULT_MODEL

english

默认 checkpoint

LAYAMCP_AUTO_TASK_DETECTION

true

自动识别语言/脚本/工作流并路由

LAYAMCP_WARMUP

true

启动时预热加载模型

LAYAMCP_WORKER_THREADS

2

推理线程池大小(负载高时调大)

LAYAMCP_HF_TOKEN

无

私有镜像或受限下载时使用

策略(决策契约的阈值)

变量

默认值

说明

LAYAMCP_CONFIDENCE_THRESHOLD

0.80

达到此置信度才可能判为 auto_execute

LAYAMCP_MARGIN_THRESHOLD

0.05

top1−top2 低于此值视为「近距平局」→ escalate

LAYAMCP_MAX_STATE_CHARS

120000

state 最大字符数

LAYAMCP_MAX_QUESTIONS

32

单次调用最大问题数

LAYAMCP_MAX_CHOICES

60

单个 choice 问题最大候选数

日志与指标

变量

默认值

说明

LAYAMCP_LOG_LEVEL

INFO

DEBUG / INFO / WARNING / ERROR / CRITICAL

LAYAMCP_LOG_JSON

true

结构化 JSON 日志(接入 ELK / Loki 时保留)

LAYAMCP_METRICS_ENABLED

true

是否采集指标

LAYAMCP_REQUEST_ID_HEADER

x-request-id

请求 id 头名称

置信度等级

threshold 参数/配置支持用名字代替数字,便于按风险分级:

等级

数值

建议场景

trivial

0.60

无副作用的分类、打标签

low

0.70

内部草稿、低风险建议

normal

0.80(默认)

常规 Agent 决策

high

0.90

生产变更、影响用户的操作

critical

0.95

删数据、动钱、安全相关


Docker 部署

# 参考 Dockerfile
FROM ghcr.io/astral-sh/uv:python3.13-bookworm-slim

WORKDIR /app
ENV UV_COMPILE_BYTECODE=1 \
    UV_LINK_MODE=copy \
    LAYAMCP_HOST=0.0.0.0 \
    LAYAMCP_PORT=8077 \
    HF_HOME=/models

COPY pyproject.toml uv.lock README.md ./
RUN uv sync --frozen --no-dev --no-install-project

COPY src ./src
RUN uv sync --frozen --no-dev

EXPOSE 8077
CMD ["uv", "run", "--no-dev", "laya-mcp", "serve"]
docker build -t laya-mcp .
docker run --rm -p 8077:8077 -v laya-models:/models \
  -e LAYAMCP_MODELS=english,multilingual \
  laya-mcp

建议把 /models 挂成 volume:首次下载约 2.2 GB,重建镜像不必重下。 生产环境请在服务前置反向代理并加鉴权——MCP 端点本身无内建认证。


开发与测试

uv sync
uv run laya-mcp doctor          # 环境自检
uv run laya-mcp tools           # 列出工具(不加载模型,适合快速验证)
uv run pytest -q                # 运行测试

不加载模型做冒烟测试

# 仅启动 app,不预热(秒级返回)
LAYAMCP_WARMUP=false uv run python -c "
from fastapi.testclient import TestClient
from laya_mcp.api.app import create_app
with TestClient(create_app()) as c:
    print(c.get('/').json())
    print(c.get('/v1/healthz').json())
"

扩展新工具

  1. 在 src/laya_mcp/core/ 里实现与传输无关的逻辑。

  2. 在 src/laya_mcp/mcp_server/server.py 里注册 @server.tool(...),返回 Pydantic 模型。

  3. 若要暴露 REST,在 src/laya_mcp/api/routes.py 加端点,复用同一函数,不要重写逻辑。

  4. 补充 laya://guide 里的说明,让 Agent 知道何时该用这个新工具。


常见问题 FAQ

正常。CPU 上加载 english + multilingual 两个 checkpoint 约需 6090 秒。 这是一次性成本——服务预热后单次决策仅 1050ms。

临时验证接口、不想等加载:--no-warmup(或 LAYAMCP_WARMUP=false), 此时首次决策会变慢。只验证工具列表可用 uv run laya-mcp tools(完全不加载模型)。

检查 routing.model 是否为 multilingual。Laya 的 english checkpoint 是 ModernBERT-large(512 token,只懂英文),中文会被自动路由到 multilingual(mmBERT-base,1024 token)。

若自动路由不符合预期,可显式指定:model="multilingual" 或 lang="zh"。 注意 multilingual 上下文更长(1024)但推理略慢,纯英文场景用 english 更快更准。

Laya 返回的是校准后的置信度,但候选数量多时原始 softmax 仍会被抬高, contract.py 里的 confidence_from_probs 已经做了校正。

另外注意启动时的告警:

laya: this checkpoint ships temperatures outside [0.5, 5] ... Treat confidence
from the affected buckets as uncalibrated.

命中这类 bucket 时,不要只看 confidence,务必参考 margin 与整体 grade—— 这正是契约存在的意义。

两个常见原因:

  1. 挂载方式错误。不能用 streamable_http_app() 再 mount 到 FastAPI(子应用 lifespan 不会执行, session manager 起不来)。请直接用 create_app()。

  2. 客户端连接方式不对。MCP 是有状态的 Streamable HTTP,客户端要用 streamable-http 传输类型 指向 http://<host>:<port>/mcp;用 SSE 或旧版 HTTP 传输连不上。

若前置了反向代理,请确保允许流式响应(关闭 buffering、开启 chunked、放行 mcp-session-id 头)。

可以。权重下载一次后会缓存在 HF_HOME(默认 ~/.cache/huggingface)。 用 uv run laya-mcp doctor 确认 包含 laya 检查点: True,之后设置 HF_HUB_OFFLINE=1 即可离线。

不会。模型是进程级单例,推理跑在有界线程池里,相同请求命中 LRU 缓存。 横向扩展只需多起几个进程/容器(每个进程会各自加载模型,注意 CPU/显存与内存占用)。

场景

用 Laya

用 LLM

「走哪条路 / 是不是 X / 严重程度几级」

✅

❌ 太贵太慢

需要生成文本、写代码、长链推理

❌

✅

高频、对延迟敏感(<100ms)

✅

❌

候选集封闭、要求可复现

✅

⚠️ 不稳定

实践上两者是互补关系:用 Laya 做前置判断与分流,只在 grade != auto_execute 时才唤起 LLM。

直接把 questions 内联传给 decide 即可,无需改代码:

{
  "state": "...",
  "questions": {
    "priority": {
      "kind": "choice",
      "instructions": "这条需求应该排什么优先级?",
      "criteria": {
        "p0": "线上故障,立刻处理",
        "p1": "本周内必须完成",
        "p2": "可以排入下个迭代"
      }
    }
  }
}

需要长期复用时,在 src/laya_mcp/core/presets.py 的 PRESETS 中注册一个新条目。

  • GET /v1/healthz — 存活(模型失败也返回 200,status=degraded)

  • GET /v1/status — 就绪、已加载 checkpoint、缓存、延迟分位

  • GET /v1/metrics — 原始计数与延迟,可直接被抓取

  • 所有响应带 x-request-id 与 x-response-time-ms 头


参考


许可证

见仓库根目录 LICENSE。使用 Laya 权重时请同时遵守其 HuggingFace 页面上的许可条款。

Available Tools

9 tools
decideAnswer a typed question set about some contextA
Read-onlyIdempotent

Answer a typed question set about state in one forward pass.

Returns calibrated answers plus a contract directive. Branch on contract.grade: act only when it is auto_execute. Compose a preset with inline questions to answer both in the same pass.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoHint the state's language to skip detection (e.g. 'zh', 'en')
taskNoForce the typed-decisions checkpoint for its four workflows
modelNoForce a checkpoint: english | multilingual | typed-decisions
stateYesContext to decide about: a string, or an object/list for structured input (code, ticket, diff, agent trace, tool result). Include the raw material, not a summary of it.
presetNoNamed question set: code_change_risk, support_triage, input_guard, routing, typed_decisions_agent_trace
questionsNoQuestions to answer, keyed by question id. Each: {kind: choice|score|yes_no, instructions, criteria, weight?}. Choice criteria: {label: when-to-pick}. Score criteria: ordered level list. yes_no criteria: {true: statement-that-should-hold}.
thresholdNoConfidence required to grade an answer auto_execute: a number in [0,1] or a named level (trivial=0.60, low=0.70, normal=0.80, high=0.90, critical=0.95). Defaults to the server's normal level.
use_cacheNoReuse an identical in-process answer

Output Schema

ParametersJSON Schema
NameRequiredDescription
gradeYesOverall grade; act autonomously only when this is auto_execute
usageNo
cachedNoTrue when served from the in-process cache
answersYes
routingYes
summaryYes
contractYesDirective: overall grade, per-question grades, and the reasons for them
autonomousYesTrue when every answer cleared the auto-execute bar
latency_msYesServer-side inference time
requires_humanYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld=false, so the bar is low, and the description adds real context: answers are 'calibrated', a `contract` directive is returned, and the grade gates action. It also notes a one-forward-pass execution model, which the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, and each sentence carries information (what it does, what it returns, how to compose). No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be spelled out, yet the description still explains the `contract.grade` field's significance. Combined with 100% schema coverage on 8 parameters, the definition is nearly complete for correct invocation; only sibling routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are fully documented structurally, establishing a baseline of 3. The description adds one genuine semantic point – that a `preset` can be composed with inline `questions` in the same pass – but otherwise repeats the schema's job.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ('Answer a typed question set about `state`') and clarifies the scope ('in one forward pass'). It is distinguishable from siblings like decide_yes_no and decide_shortlist by the word 'typed', but it never explicitly contrasts itself with them, so the differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage direction is about consuming the result: 'act only when it is `auto_execute`'. That is useful downstream guidance, but there is no when-to-use-this-vs-a-sibling guidance despite three closely related decide_* tools existing. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_shortlistAnswer a choice question over many labels (coarse-to-fine)A
Read-onlyIdempotent

Pick one of many labels without degrading the prompt budget.

Laya scores every option inside a fixed token budget, so a 200-label choice loses resolution. This embeds the state and each label, keeps the top k, then runs the normal decision pass on those. Use it above ~40 labels, or whenever labels are long. k above the label count is a no-op and reports shortlisted: false.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoLabels kept from the shortlist stage
stateYesContext used both to shortlist and to decide
choicesYesChoice labels as {label: when-to-pick}; may be hundreds of labels
thresholdNoAs in `decide`
use_cacheNoReuse an identical answer
instructionsYesWhat the choice means

Output Schema

ParametersJSON Schema
NameRequiredDescription
gradeYesOverall grade; act autonomously only when this is auto_execute
usageNo
cachedNoTrue when served from the in-process cache
answersYes
routingYes
summaryYes
contractYesDirective: overall grade, per-question grades, and the reasons for them
autonomousYesTrue when every answer cleared the auto-execute bar
latency_msYesServer-side inference time
requires_humanYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world behavior, so the safety profile is covered. The description adds real value beyond that: the two-stage embedding/shortlist mechanics, the fixed token budget rationale, and the shortlisted:false no-op outcome for oversized k.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then the mechanism, then the usage threshold and edge case, with little waste. The line breaks and title repeat add minor friction but no meaningful bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present (return values need not be explained) and annotations covering the safety profile, the description supplies the remaining essentials: mechanism, when-to-use threshold, and the no-op edge case. Naming the decide sibling directly would close the last gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents parameters. The description adds meaning for k (top-k kept from the shortlist stage, no-op above the label count) and ties threshold back to the decide tool, going beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (pick one label from many) and frames the tool's distinctive mechanism (shortlist-then-decide) versus a plain decision pass. Sibling differentiation is implicit rather than explicit: it gestures at the 'normal decision pass' (i.e. decide) and a label-count threshold, but never names the alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete activation condition ('above ~40 labels, or whenever labels are long') plus an edge-case rule (k above the label count is a no-op returning shortlisted: false). It does not explicitly say which sibling to use below the threshold, so the routing is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_yes_noAsk one yes/no questionA
Read-onlyIdempotent

The cheapest useful call: one statement, a calibrated yes/no and a directive.

Phrase statement positively (true = the action you want to take is safe/appropriate). Prefer decide when you have several questions, since extra questions cost almost nothing once the model is loaded.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoForce a checkpoint
stateYesContext to evaluate the statement against
statementYesStatement phrased so that TRUE means 'go ahead' / 'this holds'
thresholdNoAs in `decide`
use_cacheNoReuse an identical answer

Output Schema

ParametersJSON Schema
NameRequiredDescription
gradeYesOverall grade; act autonomously only when this is auto_execute
usageNo
cachedNoTrue when served from the in-process cache
answersYes
routingYes
summaryYes
contractYesDirective: overall grade, per-question grades, and the reasons for them
autonomousYesTrue when every answer cleared the auto-execute bar
latency_msYesServer-side inference time
requires_humanYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, closed-world), and the description adds genuinely new operational context: this is the cheapest call and marginal questions are nearly free once the model is loaded. That cost/latency framing is exactly the kind of trait annotations cannot express. It does not discuss caching behavior directly, though `use_cache` is documented in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the tool's value proposition, then the phrasing rule, then the alternative. No filler, though the hard line breaks inside sentences make it slightly less scannable than it could be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description still previews the shape ('a calibrated yes/no and a directive'). For a five-parameter tool with full schema coverage and annotations, the remaining description content is sufficient; only the threshold semantics (deferred to `decide`) are left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including the meaning of `statement` and `state`. The description reinforces the phrasing convention for `statement` ('phrase positively, true = go ahead'), but that largely duplicates the schema's own wording, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description makes clear this tool evaluates a single statement and returns a calibrated yes/no plus a directive, and it explicitly contrasts that scope with the multi-question sibling `decide`. It stops short of a clean verb+resource phrasing (e.g. 'evaluate one statement'), but an agent can confidently distinguish it from the other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the positive-phrasing convention and gives an explicit routing rule: 'Prefer `decide` when you have several questions, since extra questions cost almost nothing once the model is loaded.' That names the alternative and the condition that selects it, though it does not spell out when this tool should NOT be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_questionsShow what a question set expands toA
Read-onlyIdempotent

Echo a question set back: ids, kinds, options and levels.

Call this when you need the exact option labels before deciding (for example to map the answer back onto code branches).

ParametersJSON Schema
NameRequiredDescriptionDefault
presetNoPreset name
questionsNoInline questions

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description's 'Echo' wording is consistent with those annotations, but it adds no extra behavioral detail such as error behavior, auth needs, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste: the first states the purpose, the second states when to use it. Information is front-loaded and appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only expansion tool with rich annotations, full schema description coverage, and an output schema, the description covers purpose and usage sufficiently. It could be slightly more complete by clarifying how preset and inline questions relate or which one takes precedence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('Preset name' and 'Inline questions'). The description mentions output fields, not input semantics, and adds no clarification about how preset and questions interact; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Echo') and resource ('a question set'), and enumerates the expanded fields (ids, kinds, options, levels). It is clear what the tool does, though it does not explicitly distinguish itself from siblings like list_presets or decide.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage context: 'Call this when you need the exact option labels before deciding.' This implicitly positions it ahead of decision-making siblings, but it does not name an alternative tool or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_presetsList the curated question setsA
Read-onlyIdempotent

Question sets you can pass as preset instead of writing your own.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and a closed-world scope, so the safety profile is covered. The description adds the useful linkage that values are passed as `preset`, but says nothing about the size or shape of the returned list beyond what the output schema presumably provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the resource and its purpose. It is arguably too terse to be maximally useful, but nothing in it is wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description need not explain return values. Still, for a discovery tool in a set of nine decision-related siblings, the description does not say enough to route the agent confidently here versus describe_questions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter meaning is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (curated question sets) and its role as a `preset` value, so an agent understands it enumerates usable presets. However it never states the verb 'list' explicitly and does not distinguish itself from the sibling `describe_questions`, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'instead of writing your own' implies usage context – use this to discover presets rather than author your own question set. But there is no explicit when-to-use vs alternatives guidance, and no mention of how the returned presets feed into other tools such as decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preload_warmupLoad checkpoints now (slow, once)A
Idempotent

Force checkpoints into memory so later decisions are fast.

Rarely needed: the HTTP server warms up at start. Call it only if you connect over stdio or the server was started with warmup disabled. Loading downloads hundreds of MB the first time, so expect a slow call.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelsNoCheckpoints to keep resident: english, multilingual, typed-decisions

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it downloads hundreds of MB on first use and will be slow. It does not describe memory-residency duration or failure modes, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste. The core action is front-loaded, followed by the rarity caveat and the cost warning in descending order of importance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter warmup tool with an output schema and full annotation coverage, the description covers what it does, when to use it, and its performance cost. Nothing an agent needs to decide whether and how to invoke it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional 'models' parameter is documented in the schema with its valid checkpoint values. The description adds no syntax, default, or selection guidance beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and effect: 'Force checkpoints into memory so later decisions are fast.' That is a clear verb+resource+outcome. It does not explicitly contrast with any sibling, but the siblings (decide, route_request, server_status) are functionally distinct enough that differentiation is not critical.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-not guidance ('Rarely needed: the HTTP server warms up at start') and a precise triggering condition for when to use it ('only if you connect over stdio or the server was started with warmup disabled'). This is exactly the when/when-not structure that earns a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

route_requestWhich checkpoint would serve this, and why (no inference)A
Read-onlyIdempotent

Detect language/script and workflow, and name the checkpoint that fits.

Microseconds of work, no model download or forward pass. Useful before deciding whether a request needs the multilingual checkpoint at all.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoForce the language branch
modelNoForce a checkpoint
stateYesContext to classify
questionsNoOptional question set; its ids can match a typed-decisions workflow

Output Schema

ParametersJSON Schema
NameRequiredDescription
repoNo
modelYes
reasonYes
scriptNo
languageNo
workflowNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world behavior, so safety is covered. The description adds a genuinely useful performance trait -- 'microseconds of work, no model download or forward pass' -- that an agent cannot get from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and then the cost/positioning rationale. No filler, though the final sentence could be tightened into the first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover the safety profile. The description supplies purpose and cost context; only the meaning of the required 'state' payload is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents lang ('force the language branch'), model, state, and questions. The description adds nothing about parameter syntax or interaction, which is acceptable at full coverage but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'detect language/script and workflow, and name the checkpoint that fits.' The 'no model download or forward pass' framing implicitly separates it from the inference-based siblings (decide, decide_yes_no), though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a positional hint -- 'useful before deciding whether a request needs the multilingual checkpoint at all' -- which implies this is a pre-flight step ahead of a decide-family call. However, no alternative tool is named and no when-not condition is given, so the routing guidance stays implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

server_statusEngine status and latencyB
Read-onlyIdempotent

Readiness, loaded checkpoints, cache size and latency percentiles.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
readyYes
deviceNo
countersNo
requestsNo
latenciesNo
load_errorNo
cache_entriesNo
models_loadedNo
models_requestedNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety and side-effect profile is fully covered by structured data. The description contributes only a list of data categories, which is largely redundant with the output schema and adds no fresh behavioral context such as freshness, cost, or anything about the readiness semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short phrase with no filler, and the most decision-relevant item (readiness) is front-loaded. It is slightly under-structured as a bare noun list with no verb, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters, an output schema covering return values, and annotations covering safety, the only remaining burden on the description is context, which the short content list supplies adequately. It could say a little more about what 'readiness' means, but nothing essential to calling the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and the empty schema is self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment enumerates exactly what the tool exposes (readiness, loaded checkpoints, cache size, latency percentiles), which is specific to a status/inspection resource and clearly separable from the decision-oriented siblings. It lacks an explicit verb like 'returns' or 'reports', so it reads as a content list rather than a stated action, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of when an agent should check server status versus calling preload_warmup or the decide_* tools, and no prerequisites or exclusions. The agent must infer the usage context entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

should_escalateVerdict on whether to hand this to a humanA
Read-onlyIdempotent

One call, one boolean answer: does a human need to see this?

Asks three questions at once (risk, reversibility, whether a human is required), then reduces them to the strictest verdict. Use it as a gate in front of an irreversible step.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYesThe action, plan, diff or message under consideration
thresholdNoAs in `decide`

Output Schema

ParametersJSON Schema
NameRequiredDescription
gradeYesOverall grade; act autonomously only when this is auto_execute
usageNo
cachedNoTrue when served from the in-process cache
answersYes
routingYes
summaryYes
contractYesDirective: overall grade, per-question grades, and the reasons for them
autonomousYesTrue when every answer cleared the auto-execute bar
latency_msYesServer-side inference time
requires_humanYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, closed-world behavior, so the description need not restate safety. It adds genuine behavioral detail the schema cannot: the internal composition of three questions and the reduction rule ('strictest verdict'), which tells the agent how the boolean is derived.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core answer and followed by mechanism and usage. No filler; each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage, rich annotations, and an output schema, the description only needs to convey purpose, mechanism, and usage, all of which it does. The only soft spot is the unexplained `threshold` cross-reference, which is minor given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are documented structurally, and an output schema exists. The description adds no syntax or semantics for `state` or `threshold` beyond a cross-reference ('As in `decide`'), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific outcome ('one boolean answer: does a human need to see this?') that separates it from the decide_* siblings, which produce richer verdicts. It does not explicitly name which sibling to prefer, but the boolean framing is distinctive enough for selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it as a gate in front of an irreversible step' gives a clear trigger condition. No explicit when-not or named alternative (e.g., 'if you want the full rationale, use decide'), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.2.0
    • First observeddecide
    • First observeddecide_shortlist
    • First observeddecide_yes_no
    • First observeddescribe_questions
    • First observedlist_presets
    • First observedpreload_warmup
    • First observedroute_request
    • First observedserver_status
    • First observedshould_escalate

TDQS

A3.8/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have clearly distinct roles: decide variants differ by input shape (typed questions, single statement, many labels), while route_request, list_presets, describe_questions, and server tools are separate concerns. However, should_escalate and decide_yes_no both produce gate-like boolean verdicts, and an agent could confuse when to use one versus the other despite the descriptions.

Naming Consistency4/5

All names use snake_case, and most follow a predictable verb_noun or verb pattern (decide_yes_no, decide_shortlist, route_request, list_presets, describe_questions, preload_warmup). Minor deviations are should_escalate (modal verb phrase) and server_status (noun phrase), but these are still readable and consistent enough.

Tool Count5/5

Nine tools is well-scoped for a decision/inference server. The surface covers decision making, routing, preset introspection, and server health without excessive redundancy.

Completeness4/5

The set covers core decision workflows (single yes/no, typed questions, shortlisting), routing, preset listing/description, status, and warmup. Minor gaps include no explicit tool to inspect or configure checkpoints beyond status, and no direct way to answer multiple independent statements in one call, but these are workable around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides an MCP interface to the Laya decision model, enabling typed queries (yes/no, multiple choice, score) with preflight token-budget reporting, honest confidence calibration, and structured error handling.
    490 PyPI
    1
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables local Laya-MLX typed decisions (classification, scoring, risk routing, and yes/no) for coding agents like Codex, Claude Code, DeepSeek Harness, and Pi via MCP, running entirely on Apple Silicon Macs.
    1
    2
    MIT