kaoyan-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kaoyan-mcp按我的弱项生成一份408练习题并安排复习"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
11408Advance MCP
项目状态与封存说明(2026-09-21): 本项目已完成预定功能架构与内容包交付:涵盖考研 11408 核心知识库(58 章、2455 原子考点、696 考点组合,出处对应主流教材精确页码)、自适应遗忘曲线调度、标准出卷与交付流水线,且通过 Jev 多轮交叉检验达到 100% 高置信度,门禁测试 0 FAIL / 0 WARN。 现阶段作为完整交付版本正式封存,核心功能与数据包保持长期稳定,停止日常主动功能开发与管线变动。欢迎作为独立的 11408 知识库与本地 MCP 辅助工具离线使用或二次参考。
面向考研 11408(数学一 + 408 计算机统考)的结构化知识库与备考辅助工具,提供知识点索引、自适应间隔复习、出卷排版与作答分析。基于 Model Context Protocol (MCP) 构建,可直接接入 Claude Desktop、Cursor、Codex 等 AI 客户端,亦可作为命令行工具独立使用。
项目采用 Python 标准库实现,零第三方运行时依赖(要求 Python 3.9+)。知识库内容与个人学习数据解耦,学习进度及复习账本完全保存在本地工作区,不上传网络。随仓库分发覆盖数学一与 408 考点的标准内容包。
知识库质量说明:随包分发的内容包覆盖 58 章、2455 个原子知识点与 696 个考点组合,出处精确对应主流教材(张宇基础 30 讲、王道 408 系列)的 PDF 页码。全库知识点均使用 Jev 工具进行过多轮独立交叉检验,对存疑条目逐项修正后全量达到 100% 高置信度,并经分层随机抽样人工复核确认。公开导出时已通过清洗管线剥离个人学习日志、批注与本地环境路径,考点状态统一初始化为未考察(⬜),门禁校验为 0 FAIL / 0 WARN。
目录
Related MCP server: FlashLearn
它由什么组成
四层,都能单独用:
知识库:一章一个 markdown 文件,里面两张表——原子表(最小可得分单元)与组合表(考点怎么拼成一道题)。带出处页锚、掌握状态与证据。
复习调度:遗忘曲线档位 1/2/4/7/15/30/60 天,分 A/B/C 三层并限每日配额,按分值档 → 弱度 → 逾期排序取用。
出卷与交付:题目与答案分离;卷面渲染成 A4/A5/Letter 的 HTML,再按本机能力交付——能打印就打印,不能就退成 PDF、HTML、纯文本。
学习算法:能力与难度做 MAP 拟合,参数化到「原子 / 档位 / 科目 / 遗忘 / 同日复测」;预测先写日志、答完做走前回测,另有混淆模式风险、讲法排名与模型变体消融。
快速开始
需要 Python 3.9 或更新版本。
git clone https://github.com/hanserdesu/11408Advance-mcp.git
cd 11408Advance-mcp
pip install -e .铺一个空工作区:
python -m kaoyan_mcp init ./我的备考
cd 我的备考再把自带的内容包装进来(不装也能跑,只是知识库是空的,见下一节):
python -m kaoyan_mcp pack list # 看仓库自带哪些内容包
python -m kaoyan_mcp pack install 11408 --into . # 把 11408 知识包装进当前工作区看看这台机器能干什么、知识库健不健康:
python -m kaoyan_mcp capabilities # 打印机 / 浏览器 / 可用通道
python -m kaoyan_mcp overview # 科目、章节、原子、状态分布
python -m kaoyan_mcp validate # 质量门禁
python -m kaoyan_mcp due # 今天到期的复习项
python -m kaoyan_mcp plan # 今日计划工作区不用每次指定:从当前目录逐级向上找 kaoyan.config.json 或 知识库 目录;也可以用 --root 或环境变量 KAOYAN_ROOT 指定。
内容包(知识库随包分发)
引擎、内容、个人数据是三样东西:kaoyan_mcp 是引擎,packs/<名字>/ 是内容包,学习状态存在你自己的工作区里。换考试方向等于换包,代码不动。
仓库自带一个包:
包名 | 覆盖范围 | 体量 |
| 考研数学一(高数 / 线代 / 概率)+ 408(数据结构 / 组成原理 / 操作系统 / 计算机网络) | 58 章、2455 个原子、696 个组合 |
出处一律用教材 PDF 页锚,状态全部从 ⬜ 起步——个人掌握进度不进包,你在自己的工作区里积累。
python -m kaoyan_mcp pack list # 列出自带包(含 schema_version)
python -m kaoyan_mcp pack install 11408 --into ./我的备考
python -m kaoyan_mcp pack verify 11408 # 只读门禁:FAIL / WARN 逐条列出装包做两件事:把 packs/<名字>/ 复制进工作区,并把 kaoyan.config.json 的 pack 指向它。包里的 kb_dir 声明知识库位置,于是知识库随包走;个人数据(状态、复习队列、事件流、试卷)留在工作区,不会被包覆盖。
工作区里已经写过章节时默认拒绝装包:加 --force 才覆盖,或在配置里用 dirs.kb 指定自己的知识库目录。
加一个方向不用改引擎,照这个契约放一份目录就行:
packs/<名字>/
pack.json # schema_version / name / title / exam / subjects / books / kb_dir
kb/ # 知识库:一章一个 markdown,格式见下文schema_version 是留给以后换版的兼容位:引擎读得懂自己支持的版本;遇到比它新的包会直接报错让你升级引擎,而不是猜着读。目前 11408 是唯一随仓库分发的包。
接到 AI 客户端
服务走 MCP 的 stdio 传输。典型配置:
{
"mcpServers": {
"kaoyan": {
"command": "python",
"args": ["-m", "kaoyan_mcp", "serve", "--root", "/绝对路径/我的备考"]
}
}
}Windows 上把 python 换成 python.exe 的绝对路径更稳;也可以不写 --root,改成给客户端设一个 KAOYAN_ROOT 环境变量。
除了工具,服务还带 4 段提示词(teaching / grading / modeling / paper),客户端支持提示词时可以直接取用——里面写的是教学与判卷的口径,比如「先考后讲、逐点推进」「按过程判、断哪讲哪」「出处必须可定位」。
没有打印机也能用
同一个出卷工具,按本机探测到的能力自动选通道。交付策略写在配置里:
策略 | 行为 | 适合谁 |
| 能打印就打印,否则依次退到 PDF、HTML、纯文本 | 大多数人 |
| 必须打印,打不出来或页数超估算就报错、不退化 | 靠「必须落笔」逼自己的备考纪律党 |
| 优先打印,失败则退到 PDF | 有打印机但不稳定 |
| 只出 PDF | 想自己拿去打印或平板批注 |
| 只出 HTML 文件 | 没有浏览器自动化、想自己 Ctrl+P |
| 把卷面直接交给对话 | 平板、纯聊天客户端、没有文件系统 |
通道与探测方式:
通道 | 依赖 | 说明 |
| Windows:装有 SumatraPDF;macOS/Linux:有 | 先用无头浏览器把 HTML 渲染成 PDF,再送去实体打印 |
| Edge / Chrome / Chromium 任一(可用 | 渲染完顺便核对页数 |
| 无 | 任何机器都能打开;离线可用 |
| 无 | 卷面转纯文本交给对话 |
capabilities 会把探测结果和「这条通道为什么不可用」一次说清,例如缺少 SumatraPDF 时会直接告诉你装什么,而不是静默失败。
打印的两条硬规矩:一是虚拟打印机不算打印机——OneNote、Microsoft Print to PDF、Fax 这类设备出不了纸,不会被自动选中,机器上只有这类设备时 capabilities 会直接说明,你要用就自己在 config 里写 deliver.printer;二是页数超估算不出纸——渲染出来比版面预算多页,就先退回 PDF 让你压缩书写区,不浪费纸(print_strict 下直接报错)。
工具清单
工具 | 做什么 |
| 探测本机交付能力与可用通道 |
| 知识库总览:科目、章节、原子、档位与状态分布 |
| 一次命中检索:按关键词找原子/组合/章节,返回 file:line 与状态 |
| 读单章:全文加结构化原子/组合行 |
| 原子清单,可按科目/档/状态过滤 |
| 质量门禁(只读) |
| 按章节文件重算并回写 |
| 到期复习项:分层 + 分值档 + 每日配额 |
| 判定写回:队列档位、章节状态格、事件流、账目对账 |
| 今日计划:到期复习 + 未清错题 + 断点 |
| 出一张卷:卷面 HTML 与答案 markdown(答案与题目分离) |
| 出卷并按环境交付:打印 / PDF / HTML / 纯文本 |
| 答题前预测通过概率(先预测后作答,写入预测日志) |
| 算法报告:能力与难度、校准、模式风险、讲法排名、变体消融 |
同一个引擎也有一套等价的命令行入口,方便不开客户端时直接跑,或写进脚本:
python -m kaoyan_mcp search 分布函数法
python -m kaoyan_mcp chapter 概率论 第02讲
python -m kaoyan_mcp grade GL-02-021 ✅ --mode CM-02
python -m kaoyan_mcp predict --next
python -m kaoyan_mcp paper 卷子.json --policy html
python -m kaoyan_mcp report --write
python -m kaoyan_mcp pack list
python -m kaoyan_mcp pack install 11408 --into ./我的备考
python -m kaoyan_mcp pack verify 11408工作区布局
kaoyan_mcp init 会铺出这样一套目录,每一块的职责都不重叠:
目录 / 文件 | 角色 |
| 一章一个 markdown,知识点状态的唯一来源 |
| 文件地图与状态统计(对账用) |
| 间隔复习账本 |
| 错因档案 |
| 卷面与答案,永远分开 |
| append-only 的事件账本,算法唯一数据源 |
| 先预测后作答的流水 |
| 每次回测的校准指标 |
| 当前位置:断点、挂题、队列 |
| 内容包:包清单 + 自带知识库目录 |
| 各目录名、交付策略、算法开关 |
知识包和个人数据是分开的:包可以分享给别人,个人数据(状态、队列、事件流)默认留在本地。
知识库格式契约
原子表按最小可得分单元切分,组合表记录考点是怎么拼成题的:
| ID | 原子知识点 | 档 | 得分范式 | 出处 | 状态 | 证据 / 备注 |
|---|---|---|---|---|---|---|
| GL-02-001 | 分布函数法:F_Y(y)=P(g(X) ≤ y) | S | M-构造 | §2.4@p57 | ⚠️ | 【实测 2026-09-13】方向反,已直讲 |
| ID | 组合模式 | 原子链 | 题目形态 | 断点 | 真题锚点 |
|---|---|---|---|---|---|
| C-GL-02-01 | C2 多考点串联 | GL-02-001 → GL-02-002 | 解答题 | GL-02-001(翻译即错) | 2015 真题 |ID 命名:原子
<科目码>-(NN|AP)-<三位序号>,组合C-<科目码>-(NN|AP)-<两位序号>。删除即作废、留「作废」字样,不用改号补位。档位:
S/A/B(高频深度 / 常规 / 边角),决定复习配额与出卷优先级。状态:
✅掌握 /⚠️不稳 /❌薄弱 /⬜未考察。出处:统一用 PDF 页锚
@pNNN,可加节号;不写「大概某页」。证据纪律:
【实测】与【建模推演】分开标;没实测过的不许写「已掌握」。写「已讲未验收」这类免责说明的记录,状态仍应是 ⬜。得分范式:数学用
M-*,408 用S-*,门禁会检查范式前缀与学科是否匹配。
质量门禁
kb_validate 把「写得对不对、引用通不通、账对不对得上」变成可执行判据,并且只读——修不修、怎么修由你决定。分两级:
FAIL(结构契约被破坏):kb.dir 目录缺失、id.format / combo.id ID 不合规范、id.duplicate ID 重复、tier.illegal 档位非法、paradigm.illegal / paradigm.subject 范式非法或与学科不符、status.illegal 状态非法、combo.chain 原子链为空、combo.ref 引用了不存在的原子、index.mismatch / index.summary_mismatch 统计与章节实际对不上。
WARN(规格建议未满足):source.page 缺页锚、source.range 页锚超出该书页数、evidence.missing 有状态无实测标记、evidence.unjudged 未考察却带实测记录、id.sequence 段内编号断档且无作废留痕、chapter.empty 没解析出原子、combo.mode / combo.breakpoint 组合信息缺失、coverage.s_tier S 档原子没被任何组合引用(背了不考)、table.stray_pipe 单元格里混进裸竖线、index.missing / index.row_missing 账目缺行。
账目对不上时用 kb_reconcile 按章节文件重算回写,而不是手改统计表——手改迟早再漂一次。
学习算法
模型在 logit 尺度上做 MAP 拟合:
$$ \mathrm{logit}\ P(\text{流畅通过}) = \theta + \theta_{\text{subj}} - d(a) - \gamma \cdot dt + \delta_{\text{retry}} \cdot \text{retry} $$
$$ d(a) = \mu_{\text{tier}} + \Delta_a $$
先验(冷启动不胡说):能力基座 N(0.9, 1.5²),科目偏置 N(0, 0.6²),档位均值按 S/A/B 给先验再各自收 N(·, 0.5²),原子偏离 N(上次判定修正, 0.6²),同日复测项 N(0, 0.5²);遗忘常数 gamma 默认固定 ln2/10(10 天半衰期),样本足够时才自由拟合。
优化用全批梯度加 Armijo 回溯线搜索,所以「收没收敛」是可判定、可复现、可报告的:每次报告都给出迭代次数、梯度无穷范数与目标值。先验在拟合前按事件顺序回溯算好,既没有未来信息泄漏,也把每次迭代从 O(n²) 降到 O(n)。
预测纪律:先预测、后作答——答题前 learn_predict 把 P(流畅通过) 与 90% 区间写进 分析/预测日志.csv;答完 learn_report 用走前回测(第 k 条只用前 k-1 条)算 Brier、LogLoss、基线 Brier 与 ECE,并按预测区间给出可靠性表。还有三样:
模式风险:按混淆模式码做 Beta 后验并按时近加权(默认 14 天半衰期),估计「这个坑再踩的概率」。
讲法排名:把直讲事件的讲法码归因到该原子后面那次判定上,比各种讲法的一次通过率(Wilson 区间)。
变体消融:
legacy/pooled/pooled_retry/full四个变体在同一事件流上跑走前回测对比,Brier 最低者优先;样本不到 20 条时明确告诉你差异不显著、先按默认。
判定少于 20 条时所有数字只当方向看,别当结论。
配置
工作区根目录放 kaoyan.config.json,全部字段可选:
{
"root": ".",
"pack": "demo",
"dirs": { "kb": "知识库", "queue": "复习" },
"deliver": { "policy": "auto", "printer": "", "paper": "A4", "copies": 1 },
"learn": { "fit_gamma": "auto" }
}环境变量:KAOYAN_ROOT(工作区根目录,KB_BASE 是同义别名)、KAOYAN_BROWSER(浏览器可执行文件路径)。
内容包清单 pack.json 可以声明学科、书目页数与自带知识库目录:
{
"name": "demo",
"title": "演示内容包",
"kb_dir": "kb",
"subjects": { "示例数学": { "kind": "math", "paradigm": "M" } },
"books": { "示例教材": { "pages": 200 } }
}books 里的页数用于校验出处页锚是否超出教材范围。
开发
pip install -e ".[dev]"
python -m pytest -q测试全部在临时工作区里跑(复制 examples/demo-workspace),不会碰真实数据。改完知识库相关逻辑,建议按「门禁 → 对账 → 测试」的顺序过一遍。
更新说明
2026-09-21 (v0.3.0):
知识库校对与 100% 高置信度达成:全库 58 章 3151 个考点(2455 个原子考点与 696 个考点组合)通过 Jev 多轮交叉检验与人工校对,修正概念表述偏差与考点引用链断点,全量达成 100% 高置信度,并通过引用完整性门禁。
导出清洗流程优化:完善打包导出流程,移除本地复习日志、个人批注与环境路径,知识库统一初始化为未考察状态并自动同步统计索引。
发布资产:发布 Release v0.3.0,提供 Wheel 安装包与独立内容包压缩包。
2026-09-20 (v0.2.0):
考点分级整理:对照历年统考真题考查重点,对考点完成 S/A/B 分级与深度归纳。
算法与交付优化:参数化拟合收敛至活跃考点集;出卷模块优化导入逻辑,降低冷启动耗时;优化实体打印就绪轮询逻辑。
2026-09-17 (v0.1.0):
基础 MCP 协议实现,支持通过标准输入输出接入各 AI 客户端;提供出卷与复习调度基本工具。
许可
MIT,见 LICENSE。
Available Tools
14 toolsenv_capabilitiesB
探查本机交付能力(打印机 / 浏览器 / PDF 工具 / 可用通道与策略)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. '探查' implies a read-only inspection, and the enumeration of probed categories hints at what comes back, but there is no statement of latency, permission needs, or side effects (printer probing can trigger OS-level calls) — a meaningful gap for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the capability and its four scopes with zero filler. It is terse rather than padded, though the terseness borders on under-specification for a tool with no annotations and no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey the return shape; listing the four probed categories does this at a high level, but it omits usage context and any indication of result structure. Adequate as a minimum-viable probe tool, not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and additionalProperties is false, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter-level value could be added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (探查/probe) and resource (本机交付能力/local delivery capabilities), then enumerates the probed domains: printer, browser, PDF tools, channels and policies. The domain is clearly distinct from the kb_*/review_*/paper_*/learn_* siblings, though no sibling is named for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when this should be called, no prerequisites, and no routing to alternatives. An agent gets no signal that this is typically a pre-flight check before paper_build or paper_deliver, which would be the natural usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_atomsC
原子清单(可按科目/档/状态过滤)
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | ||
| limit | No | ||
| status | No | ||
| subject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It never states that this is a read-only listing, what ordering or pagination applies, what happens when filters match nothing, or what 'limit' defaults to — significant gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short parenthetical line with no wasted words, and the resource is front-loaded. But at four undocumented parameters it is terse to the point of under-specification rather than efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, 0% schema description coverage, and four parameters, the description covers only a fraction of what an agent needs. Nothing says what an 'atom' record contains or how results are ordered or bounded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names three of the four parameters (subject, tier, status) as filters, which is genuinely useful, but adds no format or accepted-value detail and ignores 'limit' entirely. Partial compensation for a 0%-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a resource ('原子清单' / atom list) and a filtering capability, which is a recognizable verb+resource shape. However, 'atom' is unexplained domain jargon and the description does nothing to distinguish this from the overlapping siblings kb_search, kb_overview, or kb_chapter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this listing tool versus kb_search or kb_overview, and no exclusions or prerequisites. The filter hint implies a browsing use case but never says so explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_chapterC
读单章:返回章节全文与结构化原子/组合行
| Name | Required | Description | Default |
|---|---|---|---|
| chapter | No | 文件名或标题关键词 | |
| subject | No | ||
| max_chars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
无 annotations,描述需承载全部行为说明。仅说明了返回内容,但未提及权限、副作用、是否缓存、最大字符限制行为等。'读单章'暗示只读,但未明确安全性或错误处理。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
极简一句话,前置动词和返回内容,无冗余信息。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
对于 3 参数、无输出 schema、无 annotations 的工具,描述仅提供了基本用途。缺少参数说明、使用场景和只读保证,但核心功能清晰。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
3 个参数,schema 描述覆盖率仅 33%(仅 chapter 有描述)。description 未补充 subject 和 max_chars 的语义(例如 subject 是分类还是过滤条件,max_chars 如何影响输出)。在低覆盖率下,描述未起到补偿作用。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
明确说明动词和资源:读取单章并返回章节全文与结构化原子/组合行。能区分于 kb_overview、kb_search、kb_atoms 等兄弟工具(后者侧重概览、搜索或原子),但未提及其与 kb_atoms 的边界。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
没有说明何时使用本工具而非 kb_atoms 或 kb_search,也没有前提条件或排除场景。用户只能假设当需要完整章节内容时使用,但描述未明说。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_overviewA
知识库总览:科目、章节数、原子数、档位与状态分布
| Name | Required | Description | Default |
|---|---|---|---|
| subject | No | 只看某个科目(可选) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the shape of the returned aggregate, which is useful, but never states that the operation is read-only with no side effects, nor any cost, caching, or scope caveats for a tool that scans the whole knowledge base.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence naming the tool's domain followed by a compact field list. No filler, no redundancy, and the most important information (what you get back) leads the clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of enumerating the returned fields, which is the key thing an agent needs for an aggregate tool. It stops short of clarifying whether counts are scoped to the caller or the whole corpus, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – the single optional 'subject' filter is already documented in the schema as 只看某个科目(可选). The description adds nothing about the filter's semantics or accepted values, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (知识库) and enumerates exactly what the overview contains: subjects, chapter counts, atom counts, tier and status distribution. This clearly separates it from the narrower kb_chapter/kb_atoms/kb_search siblings, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 总览 (overview) implies 'use this to get a summary,' but there is no explicit when-to-use statement, no mention of alternatives like kb_search or kb_atoms, and no prerequisites. Usage is inferable only from the noun.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_reconcileA
把 _INDEX.md 的统计表按章节文件重算回写(幂等)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose two genuinely useful traits: the operation mutates _INDEX.md (a write, not a read) and it is 幂等 (idempotent), so re-running is safe. However it says nothing about preconditions (must chapter files exist?), failure behavior, or whether other index content is touched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the target (_INDEX.md statistics table) and the mechanism (recompute from chapter files) front-loaded, plus the idempotency qualifier in parentheses. Zero wasted tokens.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument tool with no output schema and no annotations, the description covers what it does, which file it mutates, and that the operation is repeatable. The remaining gap is the operational trigger — when in the kb workflow this should be invoked relative to kb_validate/kb_chapter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema confirms this with additionalProperties=false, so there is nothing for the description to disambiguate. The baseline of 4 applies; no parameter-level value could be added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb pair (重算回写 = recalculate and write back) and a concrete resource (_INDEX.md's statistics table), which is far more specific than the bare name kb_reconcile. It does not explicitly distinguish itself from neighbors like kb_validate or kb_overview, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the mention of recalculating from 章节文件 (chapter files) hints it should run after chapter content changes, but there is no explicit when-to-use statement, no when-not-to-use, and no named alternative among the many kb_* siblings. Adequate but leaves the agent to infer the trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_searchC
一次命中检索:按关键词找原子/组合/章节,返回 file:line 与状态
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | ||
| limit | No | ||
| query | No | 关键词(多个词是 AND) | |
| status | No | ||
| subject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the return content (file:line and status), which matters since there is no output schema, but says nothing about scope, matching behavior, permissions, or limits. Some value added, significant gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the action front-loaded and no wasted words. Terseness here is efficient, though it borders on under-specification rather than optimal brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters (0 required), 20% schema coverage, no annotations, and no output schema, a one-sentence description is not enough. It omits any explanation of tier/status/subject/limit and provides no usage context relative to the many kb_* siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only query is documented as AND-matched keywords). Four of five parameters (tier, limit, status, subject) are undocumented in both schema and description, and the description adds no semantics for them. It does not compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (检索/search) and resource (原子/组合/章节 = atoms/combinations/chapters), plus what it returns (file:line and status). It is clearly a keyword search, but it does not explicitly differentiate itself from siblings like kb_atoms or kb_chapter that also target those resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 一次命中检索 hints at a one-shot lookup, but there is no statement of when to use this versus kb_atoms, kb_chapter, or kb_overview, and no prerequisites or exclusions. The agent must infer selection from sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kb_validateC
知识库质量门禁:结构契约、引用完整性、状态账目对账、出处页锚
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| index_check | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It never says whether validation is read-only or mutates state, whether a failure blocks downstream tools, what permissions are needed, or what the pass/fail output looks like. 'Quality gate' hints at a verdict but discloses no actual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single colon-delimited fragment with no wasted words, but it is under-specified rather than genuinely concise: no verb, no sentence structure, and the reader must parse a bare noun list. Compact but not well-formed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two undocumented parameters, no annotations, and no output schema, the description is far too thin. It does not explain the pass/fail semantics of a gate, the scope of each check, or the interaction with siblings like kb_reconcile, leaving the agent unable to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters. The description is completely silent on 'limit' and 'index_check' — it never indicates whether limit caps findings, whether index_check toggles one of the four named checks, or what defaults apply. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment '知识库质量门禁' (knowledge-base quality gate) plus a list of four check categories conveys that this validates KB integrity, but no verb is stated and the name kb_validate has to carry the action. The four checks (structural contract, reference integrity, ledger reconciliation, provenance anchors) give an agent a rough idea of scope, but it is not sharply differentiated from kb_reconcile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to run this gate, whether it must precede a build/deliver step, or how it differs from kb_reconcile despite the overlapping '状态账目对账' (ledger reconciliation) concern. The agent is left to infer invocation conditions entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
learn_predictB
答题前预测通过概率(先预测后作答;预测写入预测日志)
| Name | Required | Description | Default |
|---|---|---|---|
| log | No | ||
| atom | No | ||
| hint | No | ||
| next | No | ||
| tier | No | ||
| event | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one real side effect: the prediction is written to a prediction log (预测写入预测日志). Beyond that it says nothing about permissions, reversibility, or what a prediction result looks like, leaving most behavioral traits undocumented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single parenthetical is tight and front-loads the core action, but at this length it is under-specified rather than genuinely concise for a six-parameter tool. Nothing is wasted, but too little is said.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no annotations, no output schema, and zero parameter documentation, the description is far too thin. An agent cannot determine what to pass for atom, tier, event, hint, or next, nor what the prediction returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Six parameters with 0% schema description coverage means the description must compensate, yet it explains none of them. The phrase about writing to a prediction log only faintly gestures at the `log` parameter; atom, hint, next, tier, and event are entirely opaque in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete action and object: predicting the pass probability before answering, with the prediction recorded to a log. This is a specific verb+resource rather than a restatement of the name. It does not, however, distinguish itself from siblings like review_grade or learn_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"先预测后作答" (predict first, then answer) implies the ordering context in which the tool is used, which is useful sequencing guidance. But no alternatives are named and no conditions for when-not-to-use it are given, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
learn_reportC
学习算法报告:能力/难度、校准、模式风险、讲法排名、变体消融
| Name | Required | Description | Default |
|---|---|---|---|
| write | No | ||
| ablation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, and it largely fails. It does not say whether the report is computed on demand, whether the write parameter persists state (a potential mutation), whether ablation triggers expensive recomputation, or anything about auth/permissions. Only the topic list hints at what is produced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single compact line with no wasted words, but it is a noun phrase fragment rather than a front-loaded sentence with a verb, so it reads more like a section list than an actionable description. Adequately sized but structurally weak.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter coverage, the description is the only source of information and it omits the write/ablation semantics and any indication of return shape. For a tool with side-effect-implying flags, this is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters. The description's mention of 变体消融 (variant ablation) loosely maps to the ablation flag, but the write flag is completely undocumented in both schema and description, leaving half the parameter surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (学习算法报告 / learning algorithm report) and enumerates the report's contents — ability/difficulty, calibration, pattern risk, teaching-method ranking, variant ablation — which is more than a tautology. However it lacks a verb (generate? return? compute?) and never distinguishes itself from siblings like learn_predict, so an agent must infer that this is the aggregate reporting endpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the many siblings (learn_predict, review_due, review_grade, etc.), nor any prerequisite or trigger condition. The agent is left to guess that this is the summary/report view after other learning steps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_buildC
出一张卷:生成卷面 HTML 与答案 md(答案与卷面分离)
| Name | Required | Description | Default |
|---|---|---|---|
| spec | No | title/subject/paper/questions[...](含 stem/answer/atoms/work_mm) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the answer and paper outputs are separated, but says nothing about where files are written, whether existing artifacts are overwritten, or what permissions/side effects apply to this generation operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that is front-loaded with the core action and then the outputs. Nothing is wasted, though the terseness leaves behavioral and usage gaps unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a complex nested spec object, no annotations and no output schema, the description should say what is returned (paths? inline content?) and how the generated artifacts are delivered. It states the artifact types but omits the return contract entirely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single nested spec object is documented inline (title/subject/paper/questions with stem/answer/atoms/work_mm), so the schema does the heavy lifting. The description adds no parameter-level meaning beyond what the schema already provides, which is the baseline 3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (出一张卷 / produce a paper) and names both output artifacts (卷面 HTML and 答案 md), so the agent knows what it creates. It does not differentiate from the sibling paper_deliver, leaving the boundary to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives. The obvious sibling paper_deliver is never referenced, so the agent gets no help deciding which of the two paper tools to invoke.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_deliverC
出卷并按环境交付:打印 / PDF / HTML / 纯文本
| Name | Required | Description | Default |
|---|---|---|---|
| spec | No | ||
| policy | No | ||
| printer | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. It never says that printing is an external, hard-to-reverse side effect, whether files are written and where, or what permissions/printer availability are required — significant omissions for a tool whose policy can trigger a physical print.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded phrase with no filler, and the formats are listed compactly. Its brevity, however, reflects under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a nested object parameter, no output schema, no annotations, and zero schema description coverage, so the description must do all the work — yet it only restates the formats. An agent cannot determine what 'spec' should contain or what happens on delivery.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across all three parameters. The description lists delivery formats that partially map to the 'policy' enum, but it never names the parameter, and the nested 'spec' object and 'printer' parameter remain completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb pair (generate paper + deliver) and a resource, plus the concrete delivery formats (print/PDF/HTML/text). It does not distinguish itself from the sibling paper_build, which sounds like it also produces a paper, so an agent could confuse the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no routing to alternatives such as paper_build. The phrase '按环境交付' hints that the channel is chosen by environment, but this is left for the agent to infer rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_dueC
到期复习项(分层 A/B/C + 分值档 + 每日配额)
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| quota | No | ||
| subject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, but it discloses nothing about read-only vs mutating behavior, whether quota consumption or state changes occur, or what the return shape is. It only hints at the tier/score structure of the items.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the resource, but it is a noun fragment in parentheses rather than a structured sentence. Its brevity reflects under-specification rather than efficient communication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter review/scheduling tool with no annotations and no output schema, the description supplies almost none of the information needed to invoke it correctly: no verb, no usage context, no parameter meaning, and no behavioral disclosure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters. '每日配额' loosely maps to `quota`, but `date` and `subject` are not addressed at all, and no formats, defaults, or constraints are provided, so the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('到期复习项' / due review items) and adds scope via the parenthetical ('分层 A/B/C + 分值档 + 每日配额'). However, it is a noun fragment with no verb, so it is unclear whether the tool lists, returns, or marks items, and it does not explicitly distinguish itself from the sibling review_grade.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. The description does not state the context in which this should be called, nor does it mention review_grade or any other alternative, so the agent must infer the selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_gradeC
判定写回:复习队列 + 章节状态格 + 事件流 + _INDEX 对账
| Name | Required | Description | Default |
|---|---|---|---|
| atom | No | ||
| mode | No | 混淆模式码,如 CM-02 | |
| note | No | ||
| teach | No | 讲法码,如 TH-03 | |
| dry_run | No | ||
| verdict | No | ✅/⚠️/❌ 或 流畅/勉强/失败 | |
| sync_index | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It does disclose that the write fans out to four targets (review queue, chapter status grid, event stream, _INDEX reconciliation), which hints at multi-effect mutation, but says nothing about auth requirements, reversibility, or what dry_run/sync_index actually change.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is extremely short, which is good in principle, but it is a compressed colon-plus-list fragment rather than a front-loaded sentence. Brevity here reflects under-specification rather than efficient communication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with zero annotations, no output schema, no required-field guidance, and 43% schema coverage, a single telegraphic line is wholly inadequate. An agent cannot safely determine what will be written or under what conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Seven parameters with only 43% schema coverage, and the description adds no parameter meaning at all. The four undocumented ones (atom, note, dry_run, sync_index) leave the agent guessing about the key inputs, including whether dry_run prevents any mutation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment '判定写回' does convey a specific action (writing a verdict back), and the following list names the targets touched. But the phrasing is telegraphic jargon with no sentence structure, and nothing separates it from siblings like review_due, learn_report, or kb_reconcile, which plausibly overlap with the review/queue/reconciliation domains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when this tool should be called instead of alternatives, nor any precondition, trigger, or exclusion. With siblings such as review_due and learn_report that could each be the right entry point for recording study outcomes, the absence of routing guidance is costly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
study_planC
今日计划:到期复习 + 未清错题 + 断点
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| quota | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It lists what the plan aggregates but says nothing about whether the call is read-only, whether it mutates any state, whether it requires auth, or how '断点' (breakpoint) is determined. Only content composition is disclosed, not behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single terse phrase with no wasted words, so it is concise and front-loaded. However, the brevity is achieved through under-specification rather than efficiency, leaving key concepts (quota, breakpoint) unexplained.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description is too thin: it neither explains the parameters nor clarifies the return/plan structure, and it does not address behavior at all. It is a content label rather than a usable tool definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter is referenced in the description. The roles of 'date' (plan target date?) and 'quota' (cap on items?) are entirely undocumented, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states what the tool produces (today's plan) and enumerates its composition: due reviews, uncleared wrong questions, and breakpoints/resume points. That is a specific resource with meaningful scope beyond the name. It stops short of differentiating itself from the sibling review_due, which appears to cover one of the same components.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as review_due or review_grade. The agent must infer that this is a daily-planning aggregator purely from the name and the component list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.3.0- First observed
env_capabilities - First observed
kb_atoms - First observed
kb_chapter - First observed
kb_overview - First observed
kb_reconcile - First observed
kb_search - First observed
kb_validate - First observed
learn_predict - First observed
learn_report - First observed
paper_build - First observed
paper_deliver - First observed
review_due - First observed
review_grade - First observed
study_plan
TDQS
Scored across 14 tools
Tools are grouped by domain and mostly distinct, but paper_build and paper_deliver both generate papers (deliver additionally handles output/delivery), and kb_validate partially overlaps with kb_reconcile on ledger reconciliation, which could cause misselection.
All names use lower snake_case with consistent domain prefixes (kb_, review_, paper_, learn_, etc.), but the set mixes noun-based names (kb_overview, study_plan) and verb-based names (kb_search, review_grade), so it is not a strict verb_noun convention.
14 tools is well within the effective range, and each tool maps to a distinct operation in the study workflow without obvious redundancy or trivial filler.
The surface covers knowledge-base reading/search/validation, review cycles, planning, paper generation, and learning analytics, but lacks explicit tools for authoring/editing knowledge-base content and for managing wrong questions (add/clear), which are referenced by study_plan.
Maintenance
Related MCP Connectors
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
- NibomoOAuthcom.nibomo
Read, write, and conversationally review open-source flashcards through split read/write MCP tools.
Read, write, and conversationally review open-source flashcards through split read/write MCP tools.
Build study flashcards and exam-prep decks from your AI chat, all stored locally.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that indexes exam PDFs and HTML documents for semantic search, topic frequency analysis, and study plan generation, fully local and private.MIT
- AlicenseNot gradedqualityCmaintenanceLocal Leitner flashcard MCP server enabling AI clients to create decks, manage cards, study with spaced repetition, and track progress.MIT
- FlicenseAqualityCmaintenanceA local MCP server that helps you maintain a personal Japanese learning knowledge base, including vocabulary, confusion relations, mistakes, and spaced-repetition reviews. It provides tools and prompts for managing and reviewing your Japanese learning data without calling external LLM APIs.6-
- AlicenseNot gradedqualityAmaintenanceEnables AI clients to create spaced-repetition flashcards from material being learned, quiz users aloud, grade answers, and schedule reviews using the FSRS algorithm through a self-hosted MCP server.2AGPL 3.0