laya-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@laya-mcpassess code change risk and tell me if I can auto-merge"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
laya-mcp
把 Laya System-1 决策模型封装成一个工程化的 MCP Server(FastAPI 宿主)
一次前向传播回答一组「带类型的问题」,返回校准概率 + 可直接分支执行的决策契约, 让任意支持 MCP 的 Agent(Claude Code / Codex / Cursor / 自研 Agent)都能在毫秒级完成自动决策。
目录
Related MCP server: laya-mcp
这是什么
laya-mcp 把 Laya(一个 System-1 决策模型族)
包装成一个常驻服务,同时对外暴露两种接口:
接口 | 地址 | 面向谁 |
MCP Streamable HTTP |
| 任意 MCP 客户端 / Agent |
REST API |
| 脚本、看板、健康检查、非 MCP 调用方 |
OpenAPI 文档 |
| 人 |
它不是「又一个 LLM 接口」。Laya 是非生成式的:一次 forward pass 回答一组结构化问题, 输出的是校准过的概率分布,而不是一段文本。这使得它可以在 10~50ms 内做出「该走哪条路」的判断, 成本约为调用一次完整 LLM 的千分之一。
与传统做法对比
laya-test.py(见 examples/quickstart_router.py)演示了最小用法:
加载 Router → 构造问题 → router.predict() → 用阈值手写 if。
这个工程把它产品化了:
关注点 |
|
|
模型驻留 | 每次进程启动重新加载 | 常驻单例 + 启动预热(warmup) |
阈值策略 | 每个调用方自己写 | 决策契约,由服务统一输出 |
并发 | 单线程阻塞 | 线程池 + 事件循环不阻塞 + 结果缓存 |
接入方式 | 只能进程内 Python 调用 | MCP / REST / CLI 三种入口共用同一套逻辑 |
可观测性 |
| 结构化日志 + 指标(计数 / 延迟分位) + |
错误处理 | 直接抛异常 | 统一错误码 + HTTP 状态映射 |
配置 | 硬编码 |
|
Agent 可发现性 | 无 | 工具自描述 + resources 知识库 + prompt 模板 |
核心概念:决策契约(Contract)
这是整个项目最重要的设计。 模型给出的是概率,但 Agent 需要的是策略: 「我现在可以直接执行吗?还是必须找人确认?」
原始概率无法直接回答这个问题——两个策略 0.79 / 0.78 时,最高分看起来很高,但它不是一个决策。
因此每次调用都会额外返回一个 contract 块,把概率翻译成 4 种 grade,Agent 直接分支即可:
| 含义 | Agent 应该做什么 |
✅ | 置信度越过阈值,且与次优有明显差距 | 直接执行,不要再花 token 重新推理 |
⚠️ | 高置信但存在近距平局,或问题本身语义上就该由人拍板 | 不要仅凭模型答案行动,走完整 LLM 分析或请人确认 |
🛑 | 低置信 + 高风险场景 | 必须转人工 |
🔁 | 低置信 + 低风险 | 可以重试:换更精确的 criteria、增加候选、或交给 LLM |
grade 由三个信号共同决定(而不是单一置信度)
信号 | 作用 | 为什么单独看它不行 |
| 校准后的置信度 | 候选数量多时 softmax 会被抬高,必须用 Laya 的温度校准值 |
| top1 − top2 | 0.79 / 0.78 的两条相邻策略,不管最高分多高都不构成决策 |
| 模型自带的「需人工复核」头 | 对语义上就要求人拍板的问题,高置信也依然是 escalate |
契约结构示例
{
"grade": "auto_execute",
"autonomous": true,
"requires_human": false,
"threshold": 0.8,
"margin_threshold": 0.05,
"decisions": {
"selected_strategy": {
"value": "redis_cache",
"confidence": 0.9231,
"margin": 0.41,
"grade": "auto_execute",
"reason": "confidence 0.92 clears 0.80 and beats the runner-up by 0.41",
"requires_human": false
}
},
"grade_counts": { "auto_execute": 1, "escalate": 0, "ask_human": 0, "reconsider": 0 },
"reasons": [],
"summary": "all 1 answers clear the auto-execute bar; proceed",
"routing": { "model": "english", "reason": "English Latin text" }
}聚合规则:一个被卡住的问题会卡住整个请求。 整体
grade取所有问题中最「保守」的那个 (优先级ask_human>escalate>reconsider>auto_execute), 因为对 Agent 而言「部分可执行」比「整体不可执行」更危险。
Agent 典型分支伪代码
result = mcp.call("decide", state=diff, preset="code_change_risk")
grade = result["contract"]["grade"]
if grade == "auto_execute":
merge_pull_request() # 直接干,零推理成本
elif grade == "escalate":
run_full_llm_review(diff) # 花 token 深挖
elif grade == "ask_human":
notify_reviewer(result["contract"]["reasons"])
else: # reconsider
result = mcp.call("decide", state=diff, preset="code_change_risk",
threshold="critical") # 换阈值 / 换问题重试为什么值得用
大幅降低 Agent 决策成本:把「要不要做 / 走哪条路」交给 System-1 模型,避免为此唤起完整 LLM 的 CoT。
不引入幻觉风险:非生成式,输出是封闭候选集上的概率,不会「编造」答案。
策略集中在服务端:阈值、margin、风险等级都是配置,不用在每个 Agent 里重复实现。
对 Agent 友好:工具自带详细 description 与
structuredContent,另有laya://guide等资源做知识注入。一套逻辑三种入口:MCP / REST / CLI 共用
laya_mcp.core,不会出现行为漂移。生产可用:健康检查、指标、结构化日志、超时与体积上限、线程安全、优雅降级(模型加载失败时
/healthz返回degraded而非崩溃)。
架构
┌─────────────────────────────┐
MCP Agent ────────► │ /mcp StreamableHTTP │
(Claude Code / │ (9 tools / 4 resources / │
Codex / Cursor) │ 2 prompts) │
├─────────────────────────────┤
脚本 / 看板 ────────► │ /v1/... REST (FastAPI) │
└──────────────┬──────────────┘
│ 共用同一套实现
┌──────────────▼──────────────┐
│ laya_mcp.core │
│ settings / schema / preset │
│ engine(常驻+池+缓存) │
│ contract(决策契约) / metrics │
└──────────────┬──────────────┘
│
┌──────────────▼──────────────┐
│ laya.Router (System-1) │
│ english / multilingual / │
│ typed-decisions │
└─────────────────────────────┘目录结构
laya-mcp/
├── src/laya_mcp/
│ ├── core/ # 与传输方式无关的领域层
│ │ ├── settings.py # LAYAMCP_* 配置、阈值等级
│ │ ├── schema.py # 公开问题 schema ↔ Laya 原生格式
│ │ ├── presets.py # 5 个内置问题集
│ │ ├── contract.py # ★ 概率 → grade 的决策契约
│ │ ├── engine.py # 单例 Router + 线程池 + LRU 缓存
│ │ ├── errors.py # 统一错误分类与 HTTP 映射
│ │ └── metrics.py # 计数 / 延迟分位
│ ├── api/ # FastAPI 层
│ │ ├── app.py # create_app(),挂载 /mcp 与 /v1
│ │ ├── routes.py # REST 端点(不含决策逻辑)
│ │ ├── schemas.py # 请求 / 响应模型
│ │ └── dependencies.py # DI:engine / settings / metrics
│ ├── mcp_server/ # MCP 层
│ │ ├── server.py # 工具 / 资源 / Prompt 注册
│ │ └── questions.py # 问题解析(inline / preset / 快捷参数)
│ ├── cli.py # 命令行入口
│ └── logging_config.py # 结构化日志
├── examples/quickstart_router.py # 原始 laya-test.py 用法
├── tests/
└── pyproject.toml关键实现说明
模型只加载一次:
DecisionEngine进程级单例。启动时预热(LAYAMCP_WARMUP=true), 之后每次请求都是热态,毫秒级返回。不阻塞事件循环:推理跑在有界线程池里(
LAYAMCP_WORKER_THREADS), torch 内部会释放 GIL,因此 CPU 上并发也有收益。结果缓存:对相同
(state, questions)做进程内 LRU 缓存(默认 256 条), Agent 重试同样的问题几乎零成本;响应中cached=true可辨别。MCP 挂载方式(容易踩坑):不能用
streamable_http_app()再mount到 FastAPI—— Starlette 不会为子应用执行 lifespan,session manager 起不来,/mcp会全部失败。 正确做法是在本应用的 lifespan 里进入 session manager:manager = StreamableHTTPSessionManager(app=server._lowlevel_server, ...) app.mount("/mcp", StreamableHTTPASGIApp(manager)) # lifespan: async with manager.run(): yield参见
src/laya_mcp/api/app.py的模块注释。MCP SDK 版本:基于 mcp 2.x(
from mcp.server.mcpserver import MCPServer)。 与 1.x 的mcp.server.fastmcp.FastMCP不兼容,升级时注意导入路径变化。
快速开始
环境要求
Python 3.13+
uv(推荐)
磁盘约 3 GB(三个 checkpoint 合计约 2.2 GB)
CPU 即可运行;有 CUDA / Apple MPS 会自动加速
首次运行需要联网从 HuggingFace 下载权重,之后可完全离线
安装
git clone <your-repo-url> laya-mcp
cd laya-mcp
uv sync # 创建 .venv 并安装依赖自检
uv run laya-mcp doctor输出示例:
laya-mcp 0.2.0
python: 3.13.x
settings: models=['english', 'multilingual'] device=auto port=8077
laya: 0.3.5 -> .../laya/__init__.py
torch: 2.x.x cuda=False
HF 缓存: 2.20 GB,包含 laya 检查点: True
fastapi: ok
uvicorn: ok
mcp: ok
pydantic_settings: ok
结论: 环境可用启动服务
uv run laya-mcp serveINFO laya_mcp.engine loading laya router models=['english','multilingual'] device=cpu
INFO laya_mcp.engine laya router ready loaded=['english','multilingual']
INFO laya_mcp.api engine ready in 67.8s
INFO mcp.server... StreamableHTTP session manager started
Uvicorn running on http://127.0.0.1:8077OpenAPI 文档:http://127.0.0.1:8077/docs
MCP 端点:
http://127.0.0.1:8077/mcp
⏳ 首次启动约 60~90 秒(在 CPU 上加载两个 checkpoint)。之后所有请求都是热态, 单次决策通常 10~50ms。
第一条请求
curl -s http://127.0.0.1:8077/v1/decide \
-H 'content-type: application/json' \
-d '{
"state": "Fix a typo in README, one file, no interface change",
"preset": "code_change_risk"
}' | python -m json.tool或者完全不启服务,直接一次性决策:
uv run laya-mcp ask \
--state "Fix a typo in README, one file, no interface change" \
--preset code_change_risk接入 MCP Agent
服务启动后,任何支持 MCP 的客户端都可以直接接入。关键信息只有一条:
MCP Streamable HTTP 端点:http://127.0.0.1:8077/mcpClaude Code
claude mcp add --transport http laya http://127.0.0.1:8077/mcpCodex / Cursor / 通用 mcp.json
{
"mcpServers": {
"laya": {
"type": "streamable-http",
"url": "http://127.0.0.1:8077/mcp"
}
}
}只给单个 Agent 用(stdio,无 HTTP)
不想起 HTTP 服务时,可以用 stdio 传输,由客户端自己拉起进程:
{
"mcpServers": {
"laya": {
"command": "uv",
"args": ["run", "--directory", "/abs/path/to/laya-mcp", "laya-mcp", "serve", "--transport", "stdio"]
}
}
}stdio 模式下进程由客户端按需拉起,建议启动后先让 Agent 调一次
preload_warmup, 否则第一次决策会等待模型加载。HTTP 模式则无需关心——服务启动时已预热。
让 Agent「自动决策」的推荐接入方式
在系统提示里注入用法。服务已经在 MCP
instructions里写了使用指南, 另外可以让 Agent 读取laya://guide资源获取完整的问题编写规范。强制 Agent 先查契约。约定:任何有副作用的动作前,先调用
decide/should_escalate, 并只在contract.grade == "auto_execute"时直接执行。按风险等级传
threshold。删库、发钱、改线上配置这类动作,传"high"或"critical"。用
preset起步,再按需内联问题。preset 与内联questions会合并到同一次前向传播, 边际成本几乎为零。
MCP 工具参考
共 9 个工具,全部返回结构化结果(文本 + structuredContent 字段同名)。
工具 | 用途 | 是否加载模型 | 典型耗时 |
| 主力工具:对 | 是 | 10~50ms |
| 最省成本:单个是非问题 + 契约 | 是 | ~10ms |
| 候选标签极多(几十~上百)时,先 embedding 粗筛 top-k 再决策 | 是 | 20~80ms |
| 只做路由:判断语言 / 脚本 / 工作流,返回该用哪个 checkpoint 及原因 | 否 | <1ms |
| 列出全部内置问题集及其 question id | 否 | <1ms |
| 回显并校验一组问题(id、kind、候选、等级),不推理 | 否 | <1ms |
| 一站式:给定动作,直接回答「是否需要人工介入」 | 是 | 10~50ms |
| 就绪状态、已加载 checkpoint、缓存条数、延迟分位 | 否 | <1ms |
| 强制加载 checkpoint(仅 stdio 或关闭预热时需要) | 是(慢) | 30~90s |
调用顺序建议(先便宜后昂贵)
route_request < list_presets / describe_questions < decide_yes_no < decide < decide_shortlist
(<1ms) (<1ms) (~10ms) (10~50ms) (20~80ms)decide 参数速查
参数 | 类型 | 说明 |
| string | object | array | 必填。模型真正读到的上下文:代码 / diff / 工单 / trace / 工具返回值。放原始材料,不要放摘要,摘要会掉准确率 |
| object | array | 内联问题集,见 Question Schema |
| string | 内置问题集名;可与 |
| number | string |
|
| string | 强制 checkpoint: |
| string | 提示 |
| string | 强制使用 |
| bool | 是否复用相同请求的缓存结果,默认 |
返回值结构
{
"contract": { /* 见上文「决策契约」 */ },
"grade": "auto_execute", // 整体 grade,Agent 只需看这个
"autonomous": true, // 是否可完全自主执行
"requires_human": false,
"summary": "all 5 answers clear the auto-execute bar; proceed",
"answers": {
"blast_radius": {
"kind": "choice",
"value": "single_file",
"confidence": 0.9642,
"margin": 0.8713,
"grade": "auto_execute",
"reason": "confidence 0.96 clears 0.80 and beats the runner-up by 0.87",
"requires_human": false,
"probabilities": { "single_file": 0.9642, "module": 0.0929, "...": 0.0 },
"raw": { /* Laya 原始返回,未加工 */ }
}
},
"routing": { "model": "english", "reason": "English Latin text", "repo": "convaiinnovations/laya" },
"latency_ms": 31.4,
"cached": false,
"usage": {}
}Resources 与 Prompts
除了工具,服务还暴露 MCP Resources(可被 Agent 当知识库读取)与 Prompts(可复用模板)。
Resources
URI | 内容 |
| 问题编写指南、kind 说明、grade 语义(强烈建议 Agent 读取) |
|
|
| 全部内置问题集(JSON) |
| 实时就绪状态与延迟,等同 |
Prompts
名称 | 用途 |
| 把「打算执行的动作」转成一组问题集的填空模板 |
| 对一条用户消息做分诊(意图 / 紧急度 / 情绪 / 流失风险) |
内置问题集(Preset)
Preset 是问题的骨架,不是答案。它们覆盖了 Agent 最常遇到的五类场景:
名称 | 问题数 | 适用场景 | 包含的 question id |
| 5 | 代码改动 前置风险评审:爆炸半径、可回滚性、数据影响、测试覆盖、是否需人工 |
|
| 4 | 客服工单分诊:意图、紧急度、情绪、流失风险 |
|
| 5 | 不可信输入的护栏:越狱、提示注入、敏感数据、危害程度、话题 |
|
| 4 | 模型/算力路由:难度、领域、是否需要工具、敏感性 |
|
| 5 | Agent trace 可观测性(与 |
|
# 查看全部 preset
uv run laya-mcp presets
# 查看单个 preset 的完整问题定义
curl -s http://127.0.0.1:8077/v1/presets/code_change_risk | python -m json.toolpreset + 自定义问题可以合并,一次前向传播全部回答:
{
"state": "<diff 或工单>",
"preset": "code_change_risk",
"questions": {
"target_release": {
"kind": "choice",
"instructions": "这个改动更适合进哪个版本?",
"criteria": {
"hotfix": "紧急修复,必须立刻发",
"next_minor": "下个小版本即可",
"backlog": "不急,排入待办"
}
}
}
}问题(Question)Schema
Laya 原生词汇是 choice / score / noul(且使用 t / ins / crit 短键)。
本工程定义了一套稳定、自描述的公开 schema,调用方不需要了解底层格式:
{
"question_id": {
"kind": "choice", // choice | score | yes_no
"instructions": "选哪个方案?", // 给模型的指令
"criteria": { // 见下方「三种 kind」
"redis_cache": "适合读多写少、容忍秒级延迟",
"db_index": "适合查询条件固定、数据量中等"
},
"weight": 1.0, // 可选,建议性权重
"description": "..." // 可选,仅在罗列时展示
}
}三种 kind
| 对应 Laya 类型 |
| 说明 |
|
|
| N 选 1。每个候选的适用条件要写具体、可区分 |
|
|
| 有序等级列表,返回等级下标 |
|
|
| 是非问题。陈述句要正向表述(true 表示「可以继续」) |
编写要点(直接影响准确率)
一个问题只问一件事。不要问「这个改动安全吗」,而是拆成爆炸半径 / 可回滚 / 数据影响。
候选之间必须互斥且穷尽。留一个
other兜底往往比硬选更好。criteria 写判定条件,不写形容词。✅「涉及共享 schema,会被其他团队消费」 ❌「影响很大」。
state放原始材料,代码、diff、完整工单原文,不要放你自己的总结。问题数量 2~6 个最佳(上限由
LAYAMCP_MAX_QUESTIONS控制,默认 32), 每个问题的边际成本极低——多问一个几乎不增加耗时。不要用 Laya 做生成任务。它只做分类/打分,不是写作模型。
校验问题集
写好后先让服务回显校验,避免把无效问题发到模型:
curl -s http://127.0.0.1:8077/v1/questions/describe \
-H 'content-type: application/json' \
-d '{"preset": "input_guard"}' | python -m json.toolREST API
Base path:/v1。完整交互式文档见 /docs。
决策
方法 | 路径 | 说明 |
|
| 回答一组带类型的问题(等同 MCP |
|
| 单个是非问题 |
|
| 大规模候选:embedding 粗筛 + 决策 |
|
| 只做路由,不推理(微秒级) |
目录 / 运维
方法 | 路径 | 说明 |
|
| 全部内置问题集(含问题定义) |
|
| 单个问题集(规范化后的最终形态) |
|
| 校验并回显一组问题 |
|
| 可用 checkpoint 及加载状态 |
|
| 存活检查(模型加载失败返回 200 + |
|
| 引擎状态、阈值、缓存、延迟分位 |
|
| 原始计数与延迟数据(供 Prometheus 等抓取) |
|
| 服务版本与 laya 版本 |
示例
# 1) 是非问题:这个动作能直接执行吗?
curl -s http://127.0.0.1:8077/v1/decide/yes-no \
-H 'content-type: application/json' \
-d '{
"state": {"action": "DROP TABLE users", "env": "production"},
"statement": "this action is safe to run without human review",
"threshold": "critical"
}' | python -m json.tool
# 2) 中文工单分诊(自动路由到 multilingual checkpoint)
curl -s http://127.0.0.1:8077/v1/decide \
-H 'content-type: application/json' \
-d '{"state": "客户反馈发票被重复扣款,语气很急", "preset": "support_triage"}' \
| python -m json.tool
# 3) 只路由,不推理
curl -s http://127.0.0.1:8077/v1/route \
-H 'content-type: application/json' \
-d '{"state": "给我翻译一下这段中文"}' | python -m json.tool
# {"model":"multilingual","reason":"non-Latin script (han, 100% of letters); ..."}统一错误格式
所有错误都是同一形状,并带 request_id 便于排查:
{
"error": "invalid_question",
"message": "at most 32 questions per call",
"details": { "count": 40, "limit": 32 },
"request_id": "0f3a9c1b2d4e5f60"
}命令行 CLI
uv run laya-mcp --help命令 | 说明 |
| 启动服务(默认 FastAPI + MCP Streamable HTTP) |
| 仅提供 MCP,走标准输入输出,适合本地 Agent |
| 一次性决策,不启服务(等价于 |
| 列出内置问题集 |
| 列出 MCP 工具(不加载模型) |
| 检查环境、依赖与 checkpoint 缓存 |
serve 常用参数
uv run laya-mcp serve \
--host 0.0.0.0 \
--port 8077 \
--models english,multilingual,typed-decisions \
--device auto \
--log-level INFO参数 | 说明 |
|
|
| 监听地址与端口(默认 |
| 常驻 checkpoint,逗号分隔,或 |
|
|
| 开发热重载(仅 HTTP) |
| 启动时不加载模型(首次请求会变慢) |
ask 示例
# N 选 1(--choices 是快捷写法)
uv run laya-mcp ask \
--state "高并发下数据库查询变慢,Query 全表扫描,QPS 5000/s" \
--instructions "根据性能瓶颈选择最合适的优化方案" \
--choices "redis_cache,db_index,read_write_split,search_engine"
# 读取文件作为 state(@ 前缀)
uv run laya-mcp ask --state @./diff.patch --preset code_change_risk
# 输出完整 JSON(含契约)
uv run laya-mcp ask --state "..." --preset input_guard --json
# 是非问题
uv run laya-mcp ask --state "<发布计划>" \
--statement "this can ship without human approval"ask 的可读输出示例:
5/5 answers fail the auto-execute bar (...): ask a human before acting
blast_radius: single_file (choice, confidence 0.64, margin 0.82, ask_human)
needs_human_review: True (yes_no, confidence 0.60, margin 0.21, ask_human)
checkpoint: english - English Latin text配置项(环境变量)
所有配置项前缀为 LAYAMCP_,也可写在项目根目录的 .env 里
(见 .env.example)。
服务
变量 | 默认值 | 说明 |
|
| 监听地址 |
|
| 监听端口 |
|
|
|
|
| MCP 挂载路径 |
|
| REST 前缀 |
|
| 是否开启 |
推理
变量 | 默认值 | 说明 |
|
| 常驻 checkpoint,逗号分隔或 |
|
|
|
|
| 默认 checkpoint |
|
| 自动识别语言/脚本/工作流并路由 |
|
| 启动时预热加载模型 |
|
| 推理线程池大小(负载高时调大) |
| 无 | 私有镜像或受限下载时使用 |
策略(决策契约的阈值)
变量 | 默认值 | 说明 |
|
| 达到此置信度才可能判为 |
|
| top1−top2 低于此值视为「近距平局」→ |
|
|
|
|
| 单次调用最大问题数 |
|
| 单个 choice 问题最大候选数 |
日志与指标
变量 | 默认值 | 说明 |
|
|
|
|
| 结构化 JSON 日志(接入 ELK / Loki 时保留) |
|
| 是否采集指标 |
|
| 请求 id 头名称 |
置信度等级
threshold 参数/配置支持用名字代替数字,便于按风险分级:
等级 | 数值 | 建议场景 |
| 0.60 | 无副作用的分类、打标签 |
| 0.70 | 内部草稿、低风险建议 |
| 0.80(默认) | 常规 Agent 决策 |
| 0.90 | 生产变更、影响用户的操作 |
| 0.95 | 删数据、动钱、安全相关 |
Docker 部署
# 参考 Dockerfile
FROM ghcr.io/astral-sh/uv:python3.13-bookworm-slim
WORKDIR /app
ENV UV_COMPILE_BYTECODE=1 \
UV_LINK_MODE=copy \
LAYAMCP_HOST=0.0.0.0 \
LAYAMCP_PORT=8077 \
HF_HOME=/models
COPY pyproject.toml uv.lock README.md ./
RUN uv sync --frozen --no-dev --no-install-project
COPY src ./src
RUN uv sync --frozen --no-dev
EXPOSE 8077
CMD ["uv", "run", "--no-dev", "laya-mcp", "serve"]docker build -t laya-mcp .
docker run --rm -p 8077:8077 -v laya-models:/models \
-e LAYAMCP_MODELS=english,multilingual \
laya-mcp建议把
/models挂成 volume:首次下载约 2.2 GB,重建镜像不必重下。 生产环境请在服务前置反向代理并加鉴权——MCP 端点本身无内建认证。
开发与测试
uv sync
uv run laya-mcp doctor # 环境自检
uv run laya-mcp tools # 列出工具(不加载模型,适合快速验证)
uv run pytest -q # 运行测试不加载模型做冒烟测试
# 仅启动 app,不预热(秒级返回)
LAYAMCP_WARMUP=false uv run python -c "
from fastapi.testclient import TestClient
from laya_mcp.api.app import create_app
with TestClient(create_app()) as c:
print(c.get('/').json())
print(c.get('/v1/healthz').json())
"扩展新工具
在
src/laya_mcp/core/里实现与传输无关的逻辑。在
src/laya_mcp/mcp_server/server.py里注册@server.tool(...),返回 Pydantic 模型。若要暴露 REST,在
src/laya_mcp/api/routes.py加端点,复用同一函数,不要重写逻辑。补充
laya://guide里的说明,让 Agent 知道何时该用这个新工具。
常见问题 FAQ
正常。CPU 上加载 english + multilingual 两个 checkpoint 约需 6090 秒。
这是一次性成本——服务预热后单次决策仅 1050ms。
临时验证接口、不想等加载:--no-warmup(或 LAYAMCP_WARMUP=false),
此时首次决策会变慢。只验证工具列表可用 uv run laya-mcp tools(完全不加载模型)。
检查 routing.model 是否为 multilingual。Laya 的 english checkpoint 是
ModernBERT-large(512 token,只懂英文),中文会被自动路由到 multilingual(mmBERT-base,1024 token)。
若自动路由不符合预期,可显式指定:model="multilingual" 或 lang="zh"。
注意 multilingual 上下文更长(1024)但推理略慢,纯英文场景用 english 更快更准。
Laya 返回的是校准后的置信度,但候选数量多时原始 softmax 仍会被抬高,
contract.py 里的 confidence_from_probs 已经做了校正。
另外注意启动时的告警:
laya: this checkpoint ships temperatures outside [0.5, 5] ... Treat confidence
from the affected buckets as uncalibrated.命中这类 bucket 时,不要只看 confidence,务必参考 margin 与整体 grade——
这正是契约存在的意义。
两个常见原因:
挂载方式错误。不能用
streamable_http_app()再mount到 FastAPI(子应用 lifespan 不会执行, session manager 起不来)。请直接用create_app()。客户端连接方式不对。MCP 是有状态的 Streamable HTTP,客户端要用
streamable-http传输类型 指向http://<host>:<port>/mcp;用 SSE 或旧版 HTTP 传输连不上。
若前置了反向代理,请确保允许流式响应(关闭 buffering、开启 chunked、放行 mcp-session-id 头)。
可以。权重下载一次后会缓存在 HF_HOME(默认 ~/.cache/huggingface)。
用 uv run laya-mcp doctor 确认 包含 laya 检查点: True,之后设置 HF_HUB_OFFLINE=1 即可离线。
不会。模型是进程级单例,推理跑在有界线程池里,相同请求命中 LRU 缓存。 横向扩展只需多起几个进程/容器(每个进程会各自加载模型,注意 CPU/显存与内存占用)。
场景 | 用 Laya | 用 LLM |
「走哪条路 / 是不是 X / 严重程度几级」 | ✅ | ❌ 太贵太慢 |
需要生成文本、写代码、长链推理 | ❌ | ✅ |
高频、对延迟敏感(<100ms) | ✅ | ❌ |
候选集封闭、要求可复现 | ✅ | ⚠️ 不稳定 |
实践上两者是互补关系:用 Laya 做前置判断与分流,只在 grade != auto_execute 时才唤起 LLM。
直接把 questions 内联传给 decide 即可,无需改代码:
{
"state": "...",
"questions": {
"priority": {
"kind": "choice",
"instructions": "这条需求应该排什么优先级?",
"criteria": {
"p0": "线上故障,立刻处理",
"p1": "本周内必须完成",
"p2": "可以排入下个迭代"
}
}
}
}需要长期复用时,在 src/laya_mcp/core/presets.py 的 PRESETS 中注册一个新条目。
GET /v1/healthz— 存活(模型失败也返回 200,status=degraded)GET /v1/status— 就绪、已加载 checkpoint、缓存、延迟分位GET /v1/metrics— 原始计数与延迟,可直接被抓取所有响应带
x-request-id与x-response-time-ms头
参考
许可证
见仓库根目录 LICENSE。使用 Laya 权重时请同时遵守其 HuggingFace 页面上的许可条款。
Available Tools
9 toolsdecideAnswer a typed question set about some contextARead-onlyIdempotent
Answer a typed question set about state in one forward pass.
Returns calibrated answers plus a contract directive. Branch on
contract.grade: act only when it is auto_execute. Compose a preset
with inline questions to answer both in the same pass.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Hint the state's language to skip detection (e.g. 'zh', 'en') | |
| task | No | Force the typed-decisions checkpoint for its four workflows | |
| model | No | Force a checkpoint: english | multilingual | typed-decisions | |
| state | Yes | Context to decide about: a string, or an object/list for structured input (code, ticket, diff, agent trace, tool result). Include the raw material, not a summary of it. | |
| preset | No | Named question set: code_change_risk, support_triage, input_guard, routing, typed_decisions_agent_trace | |
| questions | No | Questions to answer, keyed by question id. Each: {kind: choice|score|yes_no, instructions, criteria, weight?}. Choice criteria: {label: when-to-pick}. Score criteria: ordered level list. yes_no criteria: {true: statement-that-should-hold}. | |
| threshold | No | Confidence required to grade an answer auto_execute: a number in [0,1] or a named level (trivial=0.60, low=0.70, normal=0.80, high=0.90, critical=0.95). Defaults to the server's normal level. | |
| use_cache | No | Reuse an identical in-process answer |
Output Schema
| Name | Required | Description |
|---|---|---|
| grade | Yes | Overall grade; act autonomously only when this is auto_execute |
| usage | No | |
| cached | No | True when served from the in-process cache |
| answers | Yes | |
| routing | Yes | |
| summary | Yes | |
| contract | Yes | Directive: overall grade, per-question grades, and the reasons for them |
| autonomous | Yes | True when every answer cleared the auto-execute bar |
| latency_ms | Yes | Server-side inference time |
| requires_human | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld=false, so the bar is low, and the description adds real context: answers are 'calibrated', a `contract` directive is returned, and the grade gates action. It also notes a one-forward-pass execution model, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, and each sentence carries information (what it does, what it returns, how to compose). No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be spelled out, yet the description still explains the `contract.grade` field's significance. Combined with 100% schema coverage on 8 parameters, the definition is nearly complete for correct invocation; only sibling routing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented structurally, establishing a baseline of 3. The description adds one genuine semantic point – that a `preset` can be composed with inline `questions` in the same pass – but otherwise repeats the schema's job.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource ('Answer a typed question set about `state`') and clarifies the scope ('in one forward pass'). It is distinguishable from siblings like decide_yes_no and decide_shortlist by the word 'typed', but it never explicitly contrasts itself with them, so the differentiation is inferential rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage direction is about consuming the result: 'act only when it is `auto_execute`'. That is useful downstream guidance, but there is no when-to-use-this-vs-a-sibling guidance despite three closely related decide_* tools existing. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decide_shortlistAnswer a choice question over many labels (coarse-to-fine)ARead-onlyIdempotent
Pick one of many labels without degrading the prompt budget.
Laya scores every option inside a fixed token budget, so a 200-label
choice loses resolution. This embeds the state and each label, keeps the
top k, then runs the normal decision pass on those. Use it above ~40
labels, or whenever labels are long. k above the label count is a no-op
and reports shortlisted: false.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Labels kept from the shortlist stage | |
| state | Yes | Context used both to shortlist and to decide | |
| choices | Yes | Choice labels as {label: when-to-pick}; may be hundreds of labels | |
| threshold | No | As in `decide` | |
| use_cache | No | Reuse an identical answer | |
| instructions | Yes | What the choice means |
Output Schema
| Name | Required | Description |
|---|---|---|
| grade | Yes | Overall grade; act autonomously only when this is auto_execute |
| usage | No | |
| cached | No | True when served from the in-process cache |
| answers | Yes | |
| routing | Yes | |
| summary | Yes | |
| contract | Yes | Directive: overall grade, per-question grades, and the reasons for them |
| autonomous | Yes | True when every answer cleared the auto-execute bar |
| latency_ms | Yes | Server-side inference time |
| requires_human | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and closed-world behavior, so the safety profile is covered. The description adds real value beyond that: the two-stage embedding/shortlist mechanics, the fixed token budget rationale, and the shortlisted:false no-op outcome for oversized k.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the mechanism, then the usage threshold and edge case, with little waste. The line breaks and title repeat add minor friction but no meaningful bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present (return values need not be explained) and annotations covering the safety profile, the description supplies the remaining essentials: mechanism, when-to-use threshold, and the no-op edge case. Naming the decide sibling directly would close the last gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents parameters. The description adds meaning for k (top-k kept from the shortlist stage, no-op above the label count) and ties threshold back to the decide tool, going beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (pick one label from many) and frames the tool's distinctive mechanism (shortlist-then-decide) versus a plain decision pass. Sibling differentiation is implicit rather than explicit: it gestures at the 'normal decision pass' (i.e. decide) and a label-count threshold, but never names the alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete activation condition ('above ~40 labels, or whenever labels are long') plus an edge-case rule (k above the label count is a no-op returning shortlisted: false). It does not explicitly say which sibling to use below the threshold, so the routing is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decide_yes_noAsk one yes/no questionARead-onlyIdempotent
The cheapest useful call: one statement, a calibrated yes/no and a directive.
Phrase statement positively (true = the action you want to take is
safe/appropriate). Prefer decide when you have several questions, since
extra questions cost almost nothing once the model is loaded.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Force a checkpoint | |
| state | Yes | Context to evaluate the statement against | |
| statement | Yes | Statement phrased so that TRUE means 'go ahead' / 'this holds' | |
| threshold | No | As in `decide` | |
| use_cache | No | Reuse an identical answer |
Output Schema
| Name | Required | Description |
|---|---|---|
| grade | Yes | Overall grade; act autonomously only when this is auto_execute |
| usage | No | |
| cached | No | True when served from the in-process cache |
| answers | Yes | |
| routing | Yes | |
| summary | Yes | |
| contract | Yes | Directive: overall grade, per-question grades, and the reasons for them |
| autonomous | Yes | True when every answer cleared the auto-execute bar |
| latency_ms | Yes | Server-side inference time |
| requires_human | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, closed-world), and the description adds genuinely new operational context: this is the cheapest call and marginal questions are nearly free once the model is loaded. That cost/latency framing is exactly the kind of trait annotations cannot express. It does not discuss caching behavior directly, though `use_cache` is documented in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the tool's value proposition, then the phrasing rule, then the alternative. No filler, though the hard line breaks inside sentences make it slightly less scannable than it could be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description still previews the shape ('a calibrated yes/no and a directive'). For a five-parameter tool with full schema coverage and annotations, the remaining description content is sufficient; only the threshold semantics (deferred to `decide`) are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters, including the meaning of `statement` and `state`. The description reinforces the phrasing convention for `statement` ('phrase positively, true = go ahead'), but that largely duplicates the schema's own wording, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description makes clear this tool evaluates a single statement and returns a calibrated yes/no plus a directive, and it explicitly contrasts that scope with the multi-question sibling `decide`. It stops short of a clean verb+resource phrasing (e.g. 'evaluate one statement'), but an agent can confidently distinguish it from the other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the positive-phrasing convention and gives an explicit routing rule: 'Prefer `decide` when you have several questions, since extra questions cost almost nothing once the model is loaded.' That names the alternative and the condition that selects it, though it does not spell out when this tool should NOT be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_questionsShow what a question set expands toARead-onlyIdempotent
Echo a question set back: ids, kinds, options and levels.
Call this when you need the exact option labels before deciding (for example to map the answer back onto code branches).
| Name | Required | Description | Default |
|---|---|---|---|
| preset | No | Preset name | |
| questions | No | Inline questions |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description's 'Echo' wording is consistent with those annotations, but it adds no extra behavioral detail such as error behavior, auth needs, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste: the first states the purpose, the second states when to use it. Information is front-loaded and appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only expansion tool with rich annotations, full schema description coverage, and an output schema, the description covers purpose and usage sufficiently. It could be slightly more complete by clarifying how preset and inline questions relate or which one takes precedence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters ('Preset name' and 'Inline questions'). The description mentions output fields, not input semantics, and adds no clarification about how preset and questions interact; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Echo') and resource ('a question set'), and enumerates the expanded fields (ids, kinds, options, levels). It is clear what the tool does, though it does not explicitly distinguish itself from siblings like list_presets or decide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: 'Call this when you need the exact option labels before deciding.' This implicitly positions it ahead of decision-making siblings, but it does not name an alternative tool or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_presetsList the curated question setsARead-onlyIdempotent
Question sets you can pass as preset instead of writing your own.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and a closed-world scope, so the safety profile is covered. The description adds the useful linkage that values are passed as `preset`, but says nothing about the size or shape of the returned list beyond what the output schema presumably provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the resource and its purpose. It is arguably too terse to be maximally useful, but nothing in it is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, the description need not explain return values. Still, for a discovery tool in a set of nine decision-related siblings, the description does not say enough to route the agent confidently here versus describe_questions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (curated question sets) and its role as a `preset` value, so an agent understands it enumerates usable presets. However it never states the verb 'list' explicitly and does not distinguish itself from the sibling `describe_questions`, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'instead of writing your own' implies usage context – use this to discover presets rather than author your own question set. But there is no explicit when-to-use vs alternatives guidance, and no mention of how the returned presets feed into other tools such as decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preload_warmupLoad checkpoints now (slow, once)AIdempotent
Force checkpoints into memory so later decisions are fast.
Rarely needed: the HTTP server warms up at start. Call it only if you connect over stdio or the server was started with warmup disabled. Loading downloads hundreds of MB the first time, so expect a slow call.
| Name | Required | Description | Default |
|---|---|---|---|
| models | No | Checkpoints to keep resident: english, multilingual, typed-decisions |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it downloads hundreds of MB on first use and will be slow. It does not describe memory-residency duration or failure modes, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste. The core action is front-loaded, followed by the rarity caveat and the cost warning in descending order of importance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter warmup tool with an output schema and full annotation coverage, the description covers what it does, when to use it, and its performance cost. Nothing an agent needs to decide whether and how to invoke it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional 'models' parameter is documented in the schema with its valid checkpoint values. The description adds no syntax, default, or selection guidance beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and effect: 'Force checkpoints into memory so later decisions are fast.' That is a clear verb+resource+outcome. It does not explicitly contrast with any sibling, but the siblings (decide, route_request, server_status) are functionally distinct enough that differentiation is not critical.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-not guidance ('Rarely needed: the HTTP server warms up at start') and a precise triggering condition for when to use it ('only if you connect over stdio or the server was started with warmup disabled'). This is exactly the when/when-not structure that earns a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
route_requestWhich checkpoint would serve this, and why (no inference)ARead-onlyIdempotent
Detect language/script and workflow, and name the checkpoint that fits.
Microseconds of work, no model download or forward pass. Useful before deciding whether a request needs the multilingual checkpoint at all.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Force the language branch | |
| model | No | Force a checkpoint | |
| state | Yes | Context to classify | |
| questions | No | Optional question set; its ids can match a typed-decisions workflow |
Output Schema
| Name | Required | Description |
|---|---|---|
| repo | No | |
| model | Yes | |
| reason | Yes | |
| script | No | |
| language | No | |
| workflow | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and closed-world behavior, so safety is covered. The description adds a genuinely useful performance trait -- 'microseconds of work, no model download or forward pass' -- that an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and then the cost/positioning rationale. No filler, though the final sentence could be tightened into the first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and annotations cover the safety profile. The description supplies purpose and cost context; only the meaning of the required 'state' payload is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents lang ('force the language branch'), model, state, and questions. The description adds nothing about parameter syntax or interaction, which is acceptable at full coverage but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'detect language/script and workflow, and name the checkpoint that fits.' The 'no model download or forward pass' framing implicitly separates it from the inference-based siblings (decide, decide_yes_no), though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a positional hint -- 'useful before deciding whether a request needs the multilingual checkpoint at all' -- which implies this is a pre-flight step ahead of a decide-family call. However, no alternative tool is named and no when-not condition is given, so the routing guidance stays implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_statusEngine status and latencyBRead-onlyIdempotent
Readiness, loaded checkpoints, cache size and latency percentiles.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ready | Yes | |
| device | No | |
| counters | No | |
| requests | No | |
| latencies | No | |
| load_error | No | |
| cache_entries | No | |
| models_loaded | No | |
| models_requested | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety and side-effect profile is fully covered by structured data. The description contributes only a list of data categories, which is largely redundant with the output schema and adds no fresh behavioral context such as freshness, cost, or anything about the readiness semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short phrase with no filler, and the most decision-relevant item (readiness) is front-loaded. It is slightly under-structured as a bare noun list with no verb, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, an output schema covering return values, and annotations covering safety, the only remaining burden on the description is context, which the short content list supplies adequately. It could say a little more about what 'readiness' means, but nothing essential to calling the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and the empty schema is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment enumerates exactly what the tool exposes (readiness, loaded checkpoints, cache size, latency percentiles), which is specific to a status/inspection resource and clearly separable from the decision-oriented siblings. It lacks an explicit verb like 'returns' or 'reports', so it reads as a content list rather than a stated action, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of when an agent should check server status versus calling preload_warmup or the decide_* tools, and no prerequisites or exclusions. The agent must infer the usage context entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
should_escalateVerdict on whether to hand this to a humanARead-onlyIdempotent
One call, one boolean answer: does a human need to see this?
Asks three questions at once (risk, reversibility, whether a human is required), then reduces them to the strictest verdict. Use it as a gate in front of an irreversible step.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | The action, plan, diff or message under consideration | |
| threshold | No | As in `decide` |
Output Schema
| Name | Required | Description |
|---|---|---|
| grade | Yes | Overall grade; act autonomously only when this is auto_execute |
| usage | No | |
| cached | No | True when served from the in-process cache |
| answers | Yes | |
| routing | Yes | |
| summary | Yes | |
| contract | Yes | Directive: overall grade, per-question grades, and the reasons for them |
| autonomous | Yes | True when every answer cleared the auto-execute bar |
| latency_ms | Yes | Server-side inference time |
| requires_human | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, closed-world behavior, so the description need not restate safety. It adds genuine behavioral detail the schema cannot: the internal composition of three questions and the reduction rule ('strictest verdict'), which tells the agent how the boolean is derived.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core answer and followed by mechanism and usage. No filler; each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage, rich annotations, and an output schema, the description only needs to convey purpose, mechanism, and usage, all of which it does. The only soft spot is the unexplained `threshold` cross-reference, which is minor given the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are documented structurally, and an output schema exists. The description adds no syntax or semantics for `state` or `threshold` beyond a cross-reference ('As in `decide`'), so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific outcome ('one boolean answer: does a human need to see this?') that separates it from the decide_* siblings, which produce richer verdicts. It does not explicitly name which sibling to prefer, but the boolean framing is distinctive enough for selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it as a gate in front of an irreversible step' gives a clear trigger condition. No explicit when-not or named alternative (e.g., 'if you want the full rationale, use decide'), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.0- First observed
decide - First observed
decide_shortlist - First observed
decide_yes_no - First observed
describe_questions - First observed
list_presets - First observed
preload_warmup - First observed
route_request - First observed
server_status - First observed
should_escalate
TDQS
Scored across 9 tools
Most tools have clearly distinct roles: decide variants differ by input shape (typed questions, single statement, many labels), while route_request, list_presets, describe_questions, and server tools are separate concerns. However, should_escalate and decide_yes_no both produce gate-like boolean verdicts, and an agent could confuse when to use one versus the other despite the descriptions.
All names use snake_case, and most follow a predictable verb_noun or verb pattern (decide_yes_no, decide_shortlist, route_request, list_presets, describe_questions, preload_warmup). Minor deviations are should_escalate (modal verb phrase) and server_status (noun phrase), but these are still readable and consistent enough.
Nine tools is well-scoped for a decision/inference server. The surface covers decision making, routing, preset introspection, and server health without excessive redundancy.
The set covers core decision workflows (single yes/no, typed questions, shortlisting), routing, preset listing/description, status, and warmup. Minor gaps include no explicit tool to inspect or configure checkpoints beyond status, and no direct way to answer multiple independent statements in one call, but these are workable around.
Maintenance
Related MCP Connectors
Decision Layer for AI Agents — 58+ tools, Advisor, MCP. Free key: POST /v1/register {}.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to make confident decisions using calibrated probabilistic tools for screening, verification, ranking, classification, and gating, fully self-hosted as an optional MCP server.510 npm1MIT
- AlicenseNot gradedqualityAmaintenanceProvides an MCP interface to the Laya decision model, enabling typed queries (yes/no, multiple choice, score) with preflight token-budget reporting, honest confidence calibration, and structured error handling.490 PyPI1Apache 2.0
- AlicenseAqualityCmaintenanceEnables running Jev AI typed decisions from any MCP client, returning choices, scores, or yes/no with calibrated probabilities.3159 npmMIT
- AlicenseAqualityBmaintenanceEnables local Laya-MLX typed decisions (classification, scoring, risk routing, and yes/no) for coding agents like Codex, Claude Code, DeepSeek Harness, and Pi via MCP, running entirely on Apple Silicon Macs.12MIT