grammar-kb-mcp
grammar-kb
英語文法教材のナレッジベース + 学習フロントエンドのモノレポプロジェクト。2層構成です:
grammar_kb/(Python バックエンド):PDF教材/講義をクリーニング・構造化し、検索可能・出典追跡可能な知識ポイントデータベースに変換します—— 自動で透かし除去、テーブル復元、知識ポイント分割、キーワードと関係の抽出を行い、ローカル SQLite(FTS5 全文検索)に格納します。 CLI、HTTP API、MCP サービスを提供します。web/(Vite フロントエンド):学生向けの学習インターフェース——コース / 語彙表 / 知識体系の3つの方法で講義を閲覧でき、 さらに課題成績の記録(CRUD、iCloud Drive に永続化、デバイス間で同期)も可能です。
「レイアウトが比較的統一され、ヘッダー/フッターの透かしがあり、テーブルを含む」教育用・技術用 PDF に適用できます。
特徴
🧹 透かし除去:フォント + 文字方向でヘッダー/フッター/斜め背景の透かしを除去(PDF サブセットフォントプレフィックス対応)
📊 テーブル復元:枠線のあるテーブルを自動検出し、GFM Markdown テーブルに復元
🧩 知識ポイント分割:見出し階層(章/節/小節/子項目/例文/練習)に従って独立して検索可能なユニットに分割
🏷️ 分類と関係:テーマ別に分類し、キーワード/シグナルワードと知識ポイント間の関係(「主節は未来形・従属節は現在形」「時制の呼応」など)を抽出
🎯 試験シグナル:各知識ポイントに試験の観点(時制/態/スペル/節…)を付与し、「試験観点から知識ポイントを逆引き」をサポート
📖 単語表:講義コーパスに基づいて単語表を生成(意味/品詞/語形変化/出典の追跡)
🔍 出典追跡可能:各知識ポイントに「講次 · 節パス · ページ番号」が付き、原文に戻れます
🗄️ 切り詰めなし:本文は SQLite
TEXTに保存(長さ上限なし)、FTS はヒット判定のみに使用🌐 HTTP API:REST サービス内蔵(FastAPI、
/docs対話ドキュメント付き)🔌 MCP 対応:MCP サービス内蔵、Claude などのクライアントから直接クエリ可能
Related MCP server: PDF RAG MCP Server
クイックスタート
uv sync # 安装依赖(含开发依赖)
uv run grammar-kb ingest ./pdfs # 导入一个 PDF 目录(全量重建,id 可复现)
uv run grammar-kb stats # 查看统计uv をインストールしていない場合:
curl -LsSf https://astral.sh/uv/install.sh | sh
よく使うコマンド
uv run grammar-kb ingest ./pdfs # 导入目录(或单个 PDF 文件)
uv run grammar-kb lecture 25 # 输出某讲的完整 Markdown(表格已还原)
uv run grammar-kb lecture 25 --format html # 输出某讲的 HTML(表格渲染为 <table>)
uv run grammar-kb kp 173 # 输出某知识点的完整 Markdown
uv run grammar-kb search "关键词" # 全文检索知识点
uv run grammar-kb search "since" --category 时态
uv run grammar-kb markers --category 时态 # 列出某类下所有关键词/标志词
uv run grammar-kb markers --tense 现在完成时 # 列出某时态的标志词
uv run grammar-kb relation 主将从现 # 按关系类型查知识点
uv run grammar-kb exam-signal 从句 # 按考点信号反查知识点(反之亦然)
uv run grammar-kb exam-signal --list # 列出所有考点信号维度
uv run grammar-kb words --limit 100 # 单词表(释义/词性/词形变化/来源)
uv run grammar-kb stats # 统计
uv run grammar-kb serve --port 8000 # 启动 HTTP 查询服务(见 http://127.0.0.1:8000/docs)デフォルトのデータベースは実行ディレクトリの data/grammar.db です。--db または環境変数 GRAMMAR_KB_DB で上書きできます。
ビルド済みデータセットを直接使う(任意)
自分で ingest したくない場合は、GitHub Releases から対応バージョンの grammar.db をダウンロードし、data/grammar.db に置く(または GRAMMAR_KB_DB でパスを指定する)だけで直接クエリできます。データセットのバージョン番号は release tag(例:data-v1)に記載されており、ライブラリ内の meta テーブルにもバージョンと生成時刻が記録されています。
アーキテクチャ
PDF ──► pdf_parser 去水印(字体+方向过滤)+ 重排行 + 还原表格
└─► structure 文本 → 大纲树 → 知识点切分(分类 + 关键词 + 关系)
└─► db SQLite(lecture / knowledge_point / marker / relation / block + FTS5)
└─► query 查询 API(CLI 与 MCP 共用)モジュール | 役割 |
| fitz で span(フォント/位置/方向)を抽出 → 透かしをフィルタリング → 行を再構成;pdfplumber でフィルタリング後の文字上でテーブルを復元 |
| 行の分類(節/小節/子項目/例文/練習)→ 知識ポイント分割 |
| 分類ルール、キーワード辞書、関係検出、試験シグナル(純関数) |
| コーパスに基づく単語表(意味/品詞/語形変化) |
| テーブル → GFM、知識ポイントと講義全体のレンダリング |
| schema + CRUD + FTS5(trigram, external-content)、切り詰めなし |
| 呼び出し向けのクエリ API |
| PDF → データベース格納(ディレクトリ取り込み = 全量再構築、id は再現可能) |
| 課題成績の独立 SQLite ライブラリ(CRUD;デフォルトは iCloud Drive) |
| コマンドライン |
| HTTP サービス(オプションの extra) |
| MCP サービス(オプションの extra) |
| 学習フロントエンド(Vite、詳細は下記「Web 学習フロントエンド」と |
データベース Schema(概要)
lecture(number UNIQUE, title, full_title, category, subcategory, source_file, page_count)
knowledge_point(lecture_id, lecture_number, title, category, section_path,
body_md, examples_md, table_md, is_table, source_page, source_bbox, tags_json, ord)
marker(kp_id, lecture_number, marker, marker_type, tense, note) -- 关键词/标志词
relation(kp_id, type, to_kp_id, note) -- 关系:主将从现/时态呼应…
lecture_block(lecture_id, page, seq, kind, text_md) -- 整讲还原用
-- 全文检索(external-content + trigram,中文子串命中)
CREATE VIRTUAL TABLE kp_fts USING fts5(title, body_md, examples_md, table_md,
content='knowledge_point', content_rowid='id', tokenize='trigram');データセットのカスタマイズ
このツールはデフォルトで「レイアウトが統一された教育用講義」向けに調整されています。データセットを変更する場合、通常は3箇所(すべて grammar_kb/ 内)を変更するだけです:
透かしフォント ——
pdf_parser.pyのWATERMARK_FONTS:新しいヘッダー/透かしフォント名を追加します。 新しい PDF のフォントを診断する簡単なスクリプト:uv run python -c "import fitz; d=fitz.open('某.pdf'); \ import collections; c=collections.Counter(s['font'] for b in d[0].get_text('dict')['blocks'] if b.get('type',0)==0 for l in b['lines'] for s in l['spans'] if s['text'].strip()); print(c)"分類ルール ——
classify.pyの_TITLE_RULES:見出しキーワードでテーマ分類をマッピングします。キーワード辞書 ——
classify.pyのTENSE_MARKERS(またはカスタムの同種辞書)。レイアウト正規表現 ——
structure.py:見出し階層で異なる記号(例:一、/(一))を使う場合は、対応する正規表現を調整します。
HTTP サービスとして
uv sync --extra server # 安装 server 依赖(fastapi + uvicorn)
uv run grammar-kb serve --port 8000 # 经由 CLI
# 或独立入口:
uv run grammar-kb-server --host 0.0.0.0 --port 8000起動後、http://127.0.0.1:8000/docs にアクセスして対話型 API ドキュメントを確認できます。エンドポイント:
メソッド | パス | 説明 |
GET |
| 統計とデータセットのメタ情報 |
GET |
| 講次一覧 |
GET |
| 特定の講義内容(テーブル復元) |
GET |
| 特定の知識ポイント |
GET |
| 全文検索 |
GET |
| シグナルワード |
GET |
| 関係で検索 |
GET |
| すべての試験シグナル次元 |
GET |
| 試験観点から知識ポイントを逆引き |
GET |
| 単語表(意味/品詞/語形変化) |
GET |
| 知識ポイントのテーマ体系ツリー(大分類→テーマ) |
GET |
| 任意の単語を検索(ECDICT 全量辞書) |
GET/POST |
| 課題成績:一覧 / 追加 |
PUT/DELETE |
| 課題成績:変更 / 削除 |
課題成績データの保存場所
成績は独立した SQLite ライブラリ(講義ライブラリ data/grammar.db とは別)に保存され、パスは次の順序で解決されます:
環境変数
GRAMMAR_KB_EXAM_DBiCloud Drive:
~/Library/Mobile Documents/com~apple~CloudDocs/grammar-kb/exam.db(macOS かつ iCloud が利用可能な場合)——データ量が少ないため、クラウドに置いて iCloud で複数デバイス間で同期しますフォールバック
data/exam.db
ライブラリは意図的に WAL モードを使用していません(単一ファイルで自己完結し、iCloud のファイル全体同期の方が信頼性が高いため)。他のデバイスでこのリポジトリをセットアップし、同じ iCloud アカウントにログインしてサービスを起動すると、同じ成績が読み込まれます。
例:
curl "http://127.0.0.1:8000/search?q=现在完成时&limit=3"
curl "http://127.0.0.1:8000/lectures/25?format=html"統一レスポンス形式:すべてのエンドポイントは {code, message, data} を返します。
// 成功(HTTP 200)
{ "code": 0, "message": "ok", "data": { "knowledge_points": 359, ... } }
// 错误(HTTP 与 code 一致)
{ "code": 404, "message": "第 99 讲不存在", "data": null }CORS:デフォルトですべてのオリジンを許可(Access-Control-Allow-Origin: *)、フロントエンドから直接クロスオリジン呼び出しが可能です。ホワイトリストを絞り込む場合:GRAMMAR_KB_CORS_ORIGINS=https://a.com,https://b.com grammar-kb-server。
MCP サービスとして
uv sync --extra mcp
uv run grammar-kb-mcp公開される tools:search_knowledge_points、get_knowledge_point、get_lecture_markdown、
list_lectures、list_markers、find_by_relation、stats。各 tool は Query の薄いラッパーです。
Claude Desktop の設定例:
{
"mcpServers": {
"grammar-kb": {
"command": "uv",
"args": ["run", "--directory", "/path/to/grammar-kb", "grammar-kb-mcp"],
"env": { "GRAMMAR_KB_DB": "/path/to/grammar-kb/data/grammar.db" }
}
}
}Web 学習フロントエンド(web/)
学生向けの学習インターフェースで、ローカルで実行中のバックエンドサービス(デフォルト http://127.0.0.1:8000``、開発中は Vite が /api/*` をプロキシ)に依存します。
# 终端 1:先起后端
uv sync --extra server && uv run grammar-kb-server
# 终端 2:再起前端
cd web && npm install && npm run dev # http://localhost:5180機能:
コース別ブラウズ:48講を文法体系(語法/時制/態/非述語/構文/総合復習)でグループ化し、クリックで講義全体を表示
語彙表:600+ の高頻度語(意味/品詞/語形変化/講義の出典)、品詞でフィルタリング、検索、並べ替え
知識体系:359 の散在する知識ポイントを「文法大分類 → テーマ」の2階層ツリーに集約、定型表現の早見表付き
🎯 試験シグナル(双方向):知識ポイント ↔ シグナルワード/時制の双方向ジャンプ——「この単語を見たら、どの知識ポイントが問われているか」
📝 課題成績:各講義に1枚の課題(35問、満点100)。問題番号をクリックして正誤を記録、スコアは自動計算; 複数回の回答はすべて保持され、修正・削除も可能;間違いノートは「講次+問題番号」で誤り回数を集計; データはバックエンドの
/exams経由で iCloud に保存(上記参照)、ブラウザ/デバイスをまたいでも失われず、旧 localStorage の記録は初回起動時に自動移行
技術スタック:Vite + ネイティブ ES Modules · marked(Markdown レンダリング)、フレームワーク依存なし。詳細は web/README.md を参照。
テスト
uv run pytest # 全部(含真实 PDF 集成)
uv run pytest -m "not integration" # 仅纯单测(无需 PDF,秒级)カバレッジ:透かしフィルタリング / 行の再構成 / テーブル復元 / 知識ポイント分割 / 分類 / キーワード抽出 / DB の切り詰めなし往復 / FTS の中国語・英語検索 / カスケード削除 / id 再構築の再現性 / クエリ / エンドツーエンド統合。
統合テストには PDF ディレクトリが必要で、環境変数 GRAMMAR_TEST_PDF_DIR で指定します。未指定または存在しない場合は自動的にスキップされます。
設計上のトレードオフと既知の制限
枠線のないテーブル:pdfplumber が枠線で検出できる ruled table のみ復元します。少数の枠線なしの多段対照は本文段落として保持されます(情報は失われません)。今後「列の空白揃え」によるフォールバック検出を追加予定です。
知識ポイント分割:統一レイアウトに基づくヒューリスティックです。特殊なレイアウトでは結合/分割が多少ずれる可能性があるため、
search+kpで確認できます。ディレクトリ取り込みは再構築:
ingest <ディレクトリ>はライブラリをクリアして再構築します(id は1から始まり、再現可能)。単一 PDF の取り込みはその講のみ更新します。
Available Tools
7 toolsfind_by_relationB
按关系类型查知识点,如 relation_type="主将从现"、"时态呼应"。
| Name | Required | Description | Default |
|---|---|---|---|
| relation_type | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It indicates a read-like lookup ('查') but does not disclose permissions, side effects, result limits, or other behavioral traits; output schema handles return shape but not operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and includes examples. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be described. But for a tool with no annotations and 0% schema coverage on its only parameter, the description leaves gaps: no sibling routing, no relation-type vocabulary, and no behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is one required parameter. The description compensates partially by naming 'relation_type' and giving example values ('主将从现', '时态呼应'), but it does not define the valid relation-type set or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb-resource pair ('查知识点') with the filter ('按关系类型') and example relation values. However, it does not distinguish itself from sibling search_knowledge_points, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative routing is provided. The relation-type filter is evident, but the agent is not told when this is preferable to search_knowledge_points.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_knowledge_pointA
按 id 获取单个知识点的完整 markdown(含解释、例句、表格、溯源)。
| Name | Required | Description | Default |
|---|---|---|---|
| kp_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the returned payload ('解释、例句、表格、溯源'), which adds value beyond the free-form object, and '获取' implies a read-only operation. It does not state permission requirements, behavior for a missing/invalid id, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words; the resource and the trigger condition (id) come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A one-parameter read tool with an output schema, so return-value shape is already covered by structured data and the description appropriately summarizes payload content. The main gap is the absence of any comparison to the search sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single kp_id parameter, which is typed only as an integer. The description partially compensates by indicating the parameter is a knowledge-point id ('按 id 获取单个知识点'), but adds no format, range, or resolution details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: retrieve a single knowledge point's full markdown by id. The word '单个' (single) implicitly contrasts with the sibling search_knowledge_points, but the description never names that alternative, so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '按 id 获取' implies the prerequisite that a kp_id must already be known, which is useful routing context. However, it gives no explicit when-to-use vs search_knowledge_points or any exclusion condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lecture_markdownB
获取某讲的完整 markdown 讲义(标题/正文/表格已还原为 GFM)。
例如 number=25 返回"第二十五讲 动词时态3"的完整 md。
| Name | Required | Description | Default |
|---|---|---|---|
| number | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state whether this is read-only (though the name implies it), what happens if an invalid number is given, whether output is cached, or how errors are surfaced. The format detail (GFM conversion) is helpful, but overall behavioral disclosure is thin for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that are front-loaded with the core purpose, followed by a concrete example. No wasted words. It could be slightly more structured (e.g., separating behavior notes), but it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values. However, with no annotations and a 0%-coverage parameter schema, the description should compensate more by clarifying read-only nature, error handling, or the relationship to list_lectures. As is, it is minimally complete for a simple read-by-id tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is only one parameter ('number'). The description adds meaning by giving a concrete example (number=25 -> lecture 25) and implying the parameter is the lecture number. This is better than nothing but not comprehensive, so a baseline 3 is appropriate for a single-param tool where the schema itself is silent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get/获取) and resource (complete markdown lecture notes), with the format detail that tables are converted to GFM. This distinguishes it from list_lectures (which presumably enumerates) and get_knowledge_point (different resource). The purpose is clear, though it doesn't explicitly contrast with siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The example ('number=25 returns 第二十五讲 动词时态3') implicitly signals when to use it: when you need the full markdown of a specific lecture by number. However, there is no explicit when-to-use vs. alternatives guidance, no mention of prerequisites (e.g., you must first know the lecture number via list_lectures), and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_lecturesA
列出已导入的全部讲次(讲号、标题、分类)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the returned fields and the 'all imported' scope, but does not mention ordering, permissions, pagination, or that the operation is read-only. For a simple zero-parameter list tool, this is minimally adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the action and scope. Every element earns its place and there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, has an output schema covering return values, and zero parameters. The description states what is listed and what fields are included, making it nearly complete. Missing are explicit usage context and any safety/read-only note, but these are minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds no parameter information, which is appropriate given there is nothing to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('列出' / list) and resource ('讲次' / lectures), and states the scope ('已导入的全部' / all imported) plus returned fields (lecture number, title, category). It clearly distinguishes itself from sibling tools by domain, but does not explicitly name or contrast with alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the tool's nature: use it to get a full list of imported lectures. However, there is no explicit guidance on when to prefer this over sibling tools like search_knowledge_points, nor are any exclusions or prerequisites stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_markersA
列出标志词/关键词,可溯源到讲次。
默认返回所有时态关键词(category="时态")。 可用 tense 限定具体时态,如 tense="现在完成时"。
| Name | Required | Description | Default |
|---|---|---|---|
| tense | No | ||
| category | No | 时态 |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It adds meaningful behavioral context by declaring the implicit default value of category, and the read-only nature is inferable from '列出'. However, it says nothing about auth needs, result size, or pagination, which remains a gap for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with what the tool returns, then the default, then the narrowing option. Zero wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and both parameters are addressed with defaults and an example. Only the missing enumeration of accepted values and any usage boundaries keep it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it documents the default for category ('默认返回所有时态关键词') and gives a concrete usage example for tense ('tense="现在完成时"'). It still omits the full set of valid tense values, so it is not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb + resource ('列出标志词/关键词') plus the traceability angle ('可溯源到讲次'), which clarifies what the listed markers link back to. It is clearly distinct from siblings like get_knowledge_point or list_lectures, though it never names an alternative to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It discloses the default behavior (returns all tense keywords with category="时态") and how to narrow results via tense, which implies usage. But it never states when to prefer this over search_knowledge_points or find_by_relation, nor any exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_knowledge_pointsB
按关键词检索语法知识点。
参数: query: 关键词(中文或英文,如 "现在完成时"、"主将从现"、"since")。 category: 可选,限定大类:词法/句法/时态/语态/非谓语/综合复习。 limit: 最多返回条数。 返回:知识点列表(标题、所在讲次、分类、标签)。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| category | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose whether this is a safe read-only operation, whether it requires permissions, whether results are paginated, or how the search behaves (e.g., exact match vs full-text). The bare return format '知识点列表' is thin behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The response is front-loaded with the core purpose, followed by a structured parameter list and a brief return summary. It is compact and every sentence serves to clarify invocation. Slightly verbose in enumerating category values, but that is necessary for an agent to pick valid values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and a 0% schema description coverage, the description steps in to cover all three parameters and the return shape, which is adequate. However, it does not describe pagination, ordering, or how to handle an empty result, leaving minor gaps for a search tool. Since an output schema exists, return values need not be fully explained, so the description's summary suffices.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains each parameter: query is a keyword in Chinese or English with concrete examples ('现在完成时', 'since'), category limits to specific large classes (词法/句法/时态/语态/非谓语/综合复习), and limit controls the maximum number returned. This adds substantial meaning beyond the schema, though limit's type and default are only in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: '按关键词检索语法知识点' (retrieve grammar knowledge points by keyword). It distinguishes from siblings like get_knowledge_point (singular retrieval) and list_lectures (a different resource), though it does not explicitly name them. The purpose is specific enough for an agent to know this is a search operation over knowledge points.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by '按关键词检索' (search by keyword), but there is no explicit statement of when to use this tool versus alternatives such as get_knowledge_point or find_by_relation. The parameter descriptions hint at refining searches with category, but no exclusion conditions or preferred scenarios are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statsC
返回知识库统计(讲次/知识点/标志词数量,按类别分布)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden, yet it says nothing about whether results are cached, whether the operation is read-only, its cost, or how the category distribution is structured. For a zero-parameter aggregation endpoint these omissions matter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the verb and resource, with the enumerated metrics acting as scope. No filler, though it is terse to the point of under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The existence of an output schema relieves the description of explaining return shapes, and there are no parameters to document. However, for a statistics endpoint over a complex knowledge base there is no mention of read-only behavior or when it should be preferred, leaving a gap given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Parameter count is zero, so there are no parameter semantics to convey and the baseline is 4; the description correctly does not fabricate parameter discussion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (return statistics) and resource (knowledge base) with enumerated sub-metrics (lecture/knowledge point/marker counts, distribution by category). This is clearer than a bare name but does not explicitly differentiate itself from siblings like list_lectures or list_markers, which also enumerate counts of the same entities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this aggregation tool versus the sibling list_* tools that would return the underlying items. An agent must infer that this is the 'counts-only' option.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.2.0- First observed
find_by_relation - First observed
get_knowledge_point - First observed
get_lecture_markdown - First observed
list_lectures - First observed
list_markers - First observed
search_knowledge_points - First observed
stats
TDQS
Scored across 7 tools
Each tool has a clearly distinct retrieval purpose: search, get by id, get lecture markdown, list lectures, list markers, find by relation, and stats. Overlap between search_knowledge_points and list_markers is minimal because markers are a specific entity type with their own filters.
Most tools follow a consistent verb_noun pattern (search_, get_, list_, find_by_). 'stats' is a minor deviation as a bare noun, but the overall naming remains predictable and readable.
7 tools is well-scoped for a knowledge base retrieval server. Each tool serves a clear function without redundancy, fitting comfortably within the ideal 3-15 range.
Core retrieval operations are covered: search, get by id, get lecture, list lectures, list markers, find by relation, and stats. Minor gaps exist, such as no direct way to list all knowledge points without a search term or to filter markers by lecture, but agents can work around these.
Maintenance
Related MCP Connectors
Search, read, cite, create, and safely update a user's private KeepFlash knowledge library.
Query and audit AppSheet apps in natural language via Knotrik's pre-scanned definitions.
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Read-only semantic search over Vedic scripture verses, commentaries, and recorded lectures.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables storage and retrieval of knowledge in a graph database format, allowing users to create, update, search, and delete entities and relationships in a Neo4j-powered knowledge graph through natural language.5-
- AlicenseNot gradedqualityDmaintenanceEnables intelligent search and question-answering over PDF documents using semantic similarity and keyword search. Supports OCR for scanned PDFs, persistent vector storage with ChromaDB, and maintains source tracking with page numbers.7MIT
- FlicenseCqualityBmaintenanceEnables academic literature management through PDF import, hybrid search, knowledge graph construction, and automated literature review generation. Combines full-text search with semantic vector search for comprehensive paper analysis.55-
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to query and interact with a graph database of markdown notes, extracting entities like wikilinks, mentions, and hashtags.MIT