twexam-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@twexam-mcpsearch questions about easement"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
twlawexam-mcp
台灣司法官/律師考試(司律一試、二試)歷屆試題 MCP 伺服器。 內建 109–114 年共 1,940 題(含測驗題標準答案、申論題 AI 擬答、法條反查、考點與爭點索引), 裝好就能在 Claude Code / Claude Desktop 裡查題、抽題練習、記錄作答、看弱點、排讀書計畫、讓 AI 逐爭點批改申論。
資料留在你本機,不上雲、不連網、不需要 API key。
安裝(一行)
需要 uv(沒有的話:pip install uv)。
Claude Code
claude mcp add twlawexam -- uvx twlawexam-mcpClaude Desktop — 在 claude_desktop_config.json 加入:
{
"mcpServers": {
"twlawexam": {
"command": "uvx",
"args": ["twlawexam-mcp"]
}
}
}其他 MCP 客戶端(Cursor、Codex、Cline…):指令都是 uvx twlawexam-mcp,stdio 傳輸。
裝好之後對 Claude 說「給我抵押權 5 題」「我常錯什麼」「剩 30 天怎麼念」就會動了。 伺服器自帶教學指令(怎麼出題、怎麼批改申論、怎麼安排複習),不用自己寫 prompt。
還沒上 PyPI 之前,或想用 GitHub 最新版:
uvx --from git+https://github.com/pass-and-prosper/twlawexam-mcp twlawexam-mcp
Related MCP server: xingce-solver
能做什麼
查題
工具 | 用途 |
| 全文搜尋歷屆考題(匹配題幹/選項/擬答) |
| 以 qid(年-考試-科目-題號)取得單題結構化內容 |
| 列出可查的考試別、年度範圍與科目 |
| 取整份試卷(某年・某考試・某科目全部題目) |
| 測驗題標準答案(題號 → 答案) |
| 申論題 AI 擬答(含免責聲明) |
| 按法條反查考過哪些題 |
| 法條考頻統計 |
考點與爭點
工具 | 用途 |
| 考點地圖:科目 → 子科目 → 考點層級,附各科實際題數 |
| 考點熱度排行(精準到「抵押權」層級) |
| 考點重點提示:核心法條/常考判決釋字/學說對立/易錯陷阱 |
| 申論爭點熱度排行(學說/實務交鋒點) |
| 依爭點找題,附各題的學說對立與實務見解 |
| 取某申論題拆出的所有爭點與實務字號 |
| 爭點重點包:辨識訊號、前置觀念、考點重點 |
| 爭點脈絡圖:一題多個爭點的先決問題鏈 |
練習與批改
工具 | 用途 |
| 依條件隨機抽題(可隱藏答案) |
| 依考點抽題(「給我抵押權 5 題」) |
| 考點申論題卷(模擬考) |
| 記錄作答、自動批改、更新間隔重複排程 |
| 申論批改評分表:爭點 checklist、學說對立、實務字號、五維評分準則 |
進度、弱點、讀書計畫
工具 | 用途 |
| 學習總覽:作答數、答對率、今天到期複習數 |
| 弱點地圖:各考點答對率由弱到強 |
| 弱點練習:優先出到期複習與最弱考點 |
| 錯誤類型診斷:觀念混淆/掉陷阱/粗心 |
| 考試就緒度:依考點頻率加權推估分數與覆蓋率 |
| 讀書計畫:綁考試日,排出攻擊順序與今日任務 |
| 清空作答記錄與複習排程 |
練習記錄放哪裡
作答記錄與複習排程寫在你本機的 progress.db,與題庫分開存放,是單機、個人的資料,不會被打包或上傳。
自 0.6 起它放在使用者資料夾、不在套件目錄裡,所以升級套件不會清掉進度:
平台 | 位置 |
Windows |
|
macOS |
|
Linux |
|
要換位置(例如放進雲端同步資料夾),設環境變數 TWEXAM_PROGRESS_DB 指到檔案路徑即可。
手機也能用
把伺服器改成 HTTP 模式,配 Cloudflare Tunnel 接到 Claude.ai 的自訂連接器,手機 App 就能用同一份題庫。 步驟見 docs/phone-setup.md。
從原始碼開發
git clone https://github.com/pass-and-prosper/twlawexam-mcp
cd twlawexam-mcp
uv venv
uv pip install -e ".[dev]"
uv run pytest -qClaude Code 會自動讀專案根目錄的 .mcp.json,開發中的版本即載入為 twexam 伺服器。
該檔預設指向 Windows 的 .venv\Scripts\python.exe,macOS / Linux 請改成 .venv/bin/python。
打包發佈:python scripts/build_release.py。它會先確認題庫裡沒有個人作答資料才打包。
重建題庫(一般使用者不需要):pip install -e ".[ingest]" 後執行 python -m twexam_mcp.ingest.run,從考選部公開資料重新下載、解析。
資料來源
考選部公開資訊(政府公開資料),網域:wwwq.moex.gov.tw、wwwc.moex.gov.tw。
免責聲明
get_model_answer與批改素材中的 AI 擬答為機器自動生成,並非官方解答或任何主管機關之見解。本工具所有內容不得作為應試依據、法律意見或任何正式文件之引用。
考選部公告之標準答案與評分準則以官方公告為準。
使用者應自行判斷資訊之準確性,作者及貢獻者不承擔任何因使用本工具所生之損失或法律責任。
授權
MIT License — 詳見 LICENSE。
Available Tools
29 toolsessay_exam_by_topicA
考點申論題卷(模擬考):把某考點或某子科目的所有申論題一次考出來,預設直接附 AI 擬答。 topic_point 或 topic_subject 至少給一個;show_answer=False 可先自己作答再看擬答。
| Name | Required | Description | Default |
|---|---|---|---|
| show_answer | No | ||
| topic_point | No | ||
| topic_subject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly discloses that all essay questions are generated in one sitting, AI model answers are included by default, and show_answer=False postpones the answers. Minor side effects such as whether the attempt is recorded or how both topic parameters are combined are not disclosed, but the key behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with the core purpose front-loaded and the parameter constraint immediately after. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter tool with no output schema, the description makes the main behavior and return content (questions plus AI answers) understandable. Minor omissions such as error behavior when neither parameter is provided and possible progress-recording side effects keep it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description fully compensates: it maps topic_point to '考點', topic_subject to '子科目', states the at-least-one requirement, and explains show_answer's default and self-testing behavior. This adds significant meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it generates a topic-based essay mock exam containing all essay questions for an exam point or sub-subject, with AI model answers attached by default. This clearly differentiates it from sibling practice/search tools by emphasizing the 'all questions at once' mock-exam scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: use this when you want a comprehensive essay exam on a topic_point or topic_subject, and it explains how to use show_answer=False for self-testing. It does not explicitly name alternative sibling tools or state when not to use it, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_answer_keyB
取某份試卷的測驗題標準答案(題號→答案)。
| Name | Required | Description | Default |
|---|---|---|---|
| year | Yes | ||
| subject | Yes | ||
| exam_code | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the returned data form (question number → answer), but it does not mention read-only behavior, failure modes, missing answer keys, or any limitations. The '取' wording implies retrieval, but little else is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler or redundancy. It front-loads the core purpose and includes the useful output mapping in parentheses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required parameters, no annotations, and no output schema, the description covers only the basic purpose and result shape. It omits parameter semantics, usage context, disambiguation from siblings, and error/edge-case behavior, so it is not complete enough for reliable agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three undocumented parameters. It only refers to '某份試卷' without explaining how year, exam_code, and subject combine or what formats are expected. An agent would have to infer meaning purely from parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('取' / get), the resource (standard answers for the test questions of a specific exam paper), and the output shape (question number → answer). This distinguishes it from sibling tools like get_model_answer, which suggests a different kind of answer output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus related alternatives such as get_model_answer, get_exam_paper, or get_question. There are no exclusions, prerequisites, or context signals to help an agent choose it over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_error_diagnosisA
錯誤類型診斷(不只正確率,要『為什麼錯』):回傳每筆答錯題帶——你反覆選的錯選項+正解 +逐項詳解+該題錯幾次+考點,並標『同考點群聚』。據此歸納誤解類型:①觀念混淆(把A當B) ②掉陷阱(選最誘人的錯選項)③粗心(偶發)。wrong_times≥2 或群聚=系統性,優先補。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| q_type | No | mcq |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and is remarkably transparent. It discloses the returned fields, the same-knowledge-point clustering marker, the three misunderstanding categories, and the threshold for treating errors as systematic. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the opening parenthetical, and the rest is a dense but purposeful enumeration of output fields, classification categories, and a priority rule. Every clause adds relevant information with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers the return shape and interpretation logic well. It falls short only in omitting input-parameter semantics and explicit usage guidance, but the defaults and self-explanatory parameter names keep it reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to explain limit and q_type, but it mentions neither. An agent cannot tell that limit caps the number of wrong-question records or that q_type selects the question type; the input semantics are entirely undocumented beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the resource ('錯誤類型診斷' / error-type diagnosis) and the verb ('回傳'), and enumerates the returned fields: wrong option, correct answer, detailed explanation, wrong count, and knowledge point. It also distinguishes itself from accuracy-only views with '不只正確率,要『為什麼錯』', though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool's context is implied: it diagnoses why wrong answers are wrong and provides interpretation rules such as 'wrong_times≥2 或群聚=系統性,優先補'. However, it never states when to choose this tool over siblings like get_weak_topics or get_progress, and it gives no exclusions or explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exam_mapA
考點地圖:各科目層級與核心考點(trial='sl1'/'sl2'/None=全部)。
| Name | Required | Description | Default |
|---|---|---|---|
| trial | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the filtering behavior for trial ('sl1'/'sl2'/None=all), which is beyond the schema. However, it does not mention whether the operation is read-only, any side effects, auth requirements, or rate limits. Given the 'get' name, some safety is implicit, but richer disclosure would be better.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the resource and content first, then the parameter. There is no filler; every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description conveys the core return content (subject hierarchy and core exam points) and the filtering behavior. It could specify the output format more precisely, but 'map' and 'hierarchy' provide enough guidance for an agent to understand what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, but the description fully explains the only parameter 'trial', including allowed values ('sl1'/'sl2') and the meaning of None (全部/all). This completely compensates for the schema's lack of detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('考點地圖', exam point map) and the content ('各科目層級與核心考點', subject hierarchy and core exam points). It also specifies the trial filter values, making the purpose concrete. It does not explicitly name a sibling, but it is distinct enough from related tools like get_topic_distribution or get_issue_distribution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: an agent needing a hierarchical map of core exam points would use this tool. The description explains the trial parameter meaning, but there is no explicit guidance on when to choose this over siblings, nor any when-not-to-use conditions. It is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exam_paperB
取整份試卷(某年·某考試·某科目全部題目)。
| Name | Required | Description | Default |
|---|---|---|---|
| year | Yes | ||
| subject | Yes | ||
| exam_code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It states the core behavior — retrieving all questions in a full exam paper — which implies a read-only fetch, but it does not mention any output limits, ordering, or whether secondary materials like answers are included. This is adequate for a simple getter but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler. The main action is front-loaded and the parenthetical adds the necessary scoping. Every word contributes to the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a field of 28 sibling tools, this description does not situate the tool relative to question-level or search-level alternatives. It also lacks parameter semantics and usage guidance, so an agent has only the name and a one-line gloss to decide when this tool is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three undocumented parameters. It only paraphrases them as '某年·某考試·某科目', adding minimal meaning beyond the property names year, exam_code, and subject. It does not explain acceptable formats, ranges, or how to discover valid codes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('取' / get) and a precise resource ('整份試卷' / entire exam paper), and clarifies the scope as all questions for a year, exam, and subject. It does not explicitly name a sibling tool, but the '全部題目' phrasing inherently distinguishes it from get_question.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use this tool versus alternatives like search_questions, get_question, or get_exam_map. There are no conditions, exclusions, or references to sibling tools, so the agent must infer usage from the tool name and general phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_grading_rubricA
申論批改評分表(學生寫完申論作答後呼叫):一次備齊批改該題所需素材——爭點 checklist +學說對立+實務字號+🔍辨識訊號/📐前置觀念+滿分擬答+五維評分準則。供你逐爭點對照 學生作答給形成性回饋(爭點漏抓/要件未涵攝/學說未選邊/字號漏引/架構),不打硬分數。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that the tool bundles all necessary grading materials in one call, is meant for qualitative formative feedback, and intentionally does not return a hard score. It does not mention side effects or error behavior, but for a read-only resource tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: it announces the resource type and call timing before listing contents. The long single sentence with emojis and slash-separated items is compact and mostly earns its place, though clearer line breaks would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description carries the burden of explaining return value and behavior. It lists the major rubric components, the intended grading use, and the 'no hard score' constraint. It does not describe the exact return structure or how to obtain qid, but it is sufficiently complete for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter qid has 0% schema coverage, and the description never explicitly defines qid or its format. It only refers indirectly to '該題' (that question), which implies qid is the essay question identifier. This adds some meaning beyond the bare schema but does not fully compensate for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as an essay-grading rubric ('申論批改評分表') and specifies exactly what it returns: issue checklist, doctrine conflicts, case citations, recognition signals, model answer, and five-dimension criteria. It also distinguishes itself from simple answer retrieval by framing the output as formative feedback material and explicitly noting it does not produce a hard score ('不打硬分數').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger condition: call after the student completes an essay answer ('學生寫完申論作答後呼叫'), and states the purpose is point-by-point formative feedback on missed issues, incomplete elements, etc. It does not name alternative sibling tools or explicitly say when not to use other tools, so it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issue_chainB
爭點脈絡圖:一題的多個爭點不是孤立的——回傳它們的『先決問題』鏈(誰是誰的先決問題、 事實哪個轉折引爆下一個爭點)。破解「孤立考點背熟、遇綜合題卻看不出脈絡」。帶使用者解綜合 申論前先看脈絡;單爭點題或尚未編脈絡者 chain 為空。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It does disclose an important behavior: the chain is empty for single-issue questions or questions with no compiled context. However, it does not mention read-only semantics, side effects, failure modes, or the output format, relying on the 'get' prefix to imply safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the core function, then purpose, then usage and the empty-chain condition. The motivational phrase about memorizing isolated test points is mildly promotional, but it reinforces the intended use context rather than being pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter retrieval tool, the description explains the output concept and the empty case clearly. However, because there is no output schema, it does not specify the exact structure of the chain, such as whether it is a list of prerequisite edges or how issue identifiers are represented, leaving some ambiguity for an agent interpreting the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never explicitly defines the 'qid' parameter or its expected format. The phrase '一題' and '單爭點題' hint that qid identifies a question, but this is indirect compensation for the only required parameter. An agent must infer qid conventions primarily from sibling tool names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: it returns the prerequisite-question chain ('先決問題鏈') among multiple issues of a question. It also distinguishes itself by noting the chain is empty for single-issue or un-linked questions, which helps separate it from generic issue-lookup siblings like get_issues or get_issue_primer, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells users to consult the chain before solving comprehensive essay questions ('帶使用者解綜合申論前先看脈絡') and gives a when-not condition: single-issue questions or questions without a compiled context return an empty chain. It lacks explicit alternative tool names, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issue_distributionB
爭點熱度排行:申論真實考點(學說/實務交鋒點)考過幾題,由多到少。可選子科目篩選。
| Name | Required | Description | Default |
|---|---|---|---|
| topic_subject | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It usefully discloses the sorting order ('由多到少'), the scope ('申論真實考點'), and the optional filter. However, it does not explain data sources, whether counts are unique questions, or how the optional filter affects the output, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that leads with the core purpose ('爭點熱度排行'), then explains the basis, ordering, and filtering in one breath. Every phrase adds useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, has only one optional parameter, and an output schema exists, so the description does not need to detail return values. Still, it lacks guidance on parameter value semantics and does not clarify how this tool relates to the many sibling tools, leaving some context-dependent decisions to the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare parameter definition. It does communicate that the single parameter is an optional sub-subject filter ('可選子科目篩選'), which adds meaning beyond the schema. It does not, however, define what a sub-subject is or how values should be formatted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this tool provides a ranking of exam issues by how often they appear, sorted from most to least frequent, and specifies the scope as essay questions on real points of contention (學說/實務交鋒點). It is more specific than a generic 'get distribution' label, though it does not explicitly differentiate itself from the similar sibling get_topic_distribution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as get_topic_distribution, get_statute_frequency, or search_by_issue. It only mentions an optional sub-subject filter, leaving the selection context entirely implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issue_primerA
申論爭點重點包(做題前必讀):🔍辨識訊號(怎麼從事實認出此爭點)+📐前置觀念 (要先懂的定義/法理)+考點重點。issue 可給標準爭點名或原始爭點字串(自動正規化)。
| Name | Required | Description | Default |
|---|---|---|---|
| issue | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains that it performs automatic normalization of the issue string, which is a useful behavioral detail beyond the schema. Since no annotations are provided, this partial transparency is helpful but still lacks details on what the primer includes in terms of length or format. It is not contradictory, but it does not fully disclose the behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, using emojis and symbols to structure the content, and front-loads the purpose. It is efficient, though the use of emojis might reduce formality, but the info density is high and every sentence serves a purpose. It is concise and structured well.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one parameter and no output schema, the description covers the essential aspects: what it does (primer content), how to use it (issue input), and a key behavior (normalization). It is sufficiently complete for the tool's complexity, though it could mention the return format or length, but that is minor for a primer tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides the parameter name and type, with 0% cover in the schema description. The description compensates by explaining that 'issue' can be a standard issue name or raw string, and that it will be automatically normalized. This adds meaningful semantics beyond the schema, making the parameter usage clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: it prepares a study primer for a specific issue, including signal identification, prerequisite concepts, and exam points. It distinguishes itself from siblings like get_topic_primer and get_issue_chain by focusing on the primer aspect, though it does not explicitly name them. The verb and resource are specific (get primer for issue), and the action is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says it is a 'must-read before doing questions' and indicates that the issue parameter can be a standard name or raw string, which implies when to use it. However, it does not explicitly state when not to use it or compare with alternative tools like get_topic_primer. The usage context is implied but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issuesA
取某申論題拆出的所有爭點:爭點名+學說對立+實務見解(判例/決議/釋字/憲判字號)。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. The verb '取' implies a read operation, and the description lists what is returned, but it does not explicitly state side-effect-free behavior, authentication requirements, or error handling. The read intent is reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the action and resource, then lists output contents. No filler or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single required parameter, an existing output schema, and a simple read operation, the description sufficiently covers the core purpose. It could mention edge cases (e.g., no issues found) but that is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description associates qid with a specific essay question ('某申論題'), adding context beyond the raw parameter name. However, it does not specify the format, source, or any constraints for qid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('取' – retrieve) and resource ('某申論題拆出的所有爭點' – all issues extracted from an essay question), and enumerates the return contents (issue name, doctrinal opposition, practical opinions). This clearly distinguishes it from sibling tools like get_issue_primer or get_issue_chain, which have different scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: to fetch issues for a particular qid. No explicit guidance is given on when to choose this tool over siblings or when not to use it, but the context of retrieving all issues for a specific question is reasonably inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_answerB
取申論題 AI 擬答(含免責聲明)。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose one useful behavioral trait: the returned AI answer includes a disclaimer. However, it does not explain the response format, error behavior, or whether the answer is non-authoritative beyond the vague '免責聲明' note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler and front-loads the key object and action. It earns its place, though the overall under-specification keeps it from a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple one-parameter retrieval tool with no output schema, so the description need not be long, but it should at least clarify the return value and the role of qid. It states the return is an AI model answer with a disclaimer, which is a minimal but adequate core for the simplest calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter 'qid' with 0% description coverage, and the tool description never mentions it. The word '申論題' weakly implies qid is an essay-question ID, but no format, source, or usage detail is provided. The description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('取' / retrieve) and a specific resource ('申論題 AI 擬答' / essay-question AI model answer), and it adds the detail that a disclaimer is included. It is clear enough to distinguish from 'get_answer_key' and 'get_grading_rubric' in general meaning, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus siblings like get_answer_key, get_question, or get_grading_rubric. There are no prerequisites, exclusions, or alternative conditions. The only implicit signal is that an AI model answer is desired.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_progressC
學習總覽:總作答數、答對率、已練題數、今天到期複習數。
| Name | Required | Description | Default |
|---|---|---|---|
| q_type | No | mcq |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of conveying behavior. It does communicate that this is a read-only summary and lists the exact data points, but it does not mention data freshness, scope, or any side effects. For a getter, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It presents the key metrics in a compact parallel list and is exactly as concise as the content allows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter getter with no output schema, the description supplies a useful summary of the return content. However, it omits any explanation of the q_type parameter, usage context relative to siblings, and potential caveats about the metrics. This leaves a meaningful gap for an agent trying to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, q_type, has zero schema description coverage and is never mentioned in the description. The agent cannot determine what q_type values are valid or whether it filters by question type, mode, or something else. The default 'mcq' and title 'Q Type' give a weak hint, but the description itself adds no meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('progress') and the specific purpose ('learning overview'), and enumerates the concrete metrics returned: total answers, accuracy rate, practiced questions, and reviews due today. This distinguishes it from mutating siblings like reset_progress, though it could more explicitly contrast with get_readiness or get_topic_distribution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool instead of related siblings such as get_readiness, get_topic_distribution, or reset_progress. The phrase 'learning overview' is an implicit hint at best; no explicit conditions, exclusions, or alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_questionA
以 qid(年-考試-科目-題號)取得單題結構化內容。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of conveying behavioral traits. The verb '取得' implies a read-only retrieval, and 'structured content' hints at the return nature, but the description does not disclose possible errors, exact response shape, or any side-effect-related caveats. This is adequate but minimal for a simple fetch-by-ID tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler or repetition. It front-loads the key parameter format before stating the result, making it easy to scan and parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one required parameter, no nested objects, no output schema), the description is nearly complete: it defines the input format and the type of result. It only lacks explicit guidance about when to prefer this over the many sibling tools, but for invoking the tool with a known qid, the essential information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only labels the parameter as 'Qid' with type string, while the description adds essential semantics: qid means 年-考試-科目-題號 (year-exam-subject-question-number). This fully compensates for the 0% schema description coverage and tells the agent exactly what value to supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb '取得' and a clear resource '單題結構化內容' (structured content of a single question), and explicitly defines the qid format as 年-考試-科目-題號. This makes the tool's function unambiguous and distinguishes it from sibling tools like search_questions or get_exam_paper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as get_model_answer or search_questions. There are no conditions, exclusions, or mentions of prerequisite context that would help an agent choose correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_readinessC
考試就緒度:依考點頻率加權推估分數、覆蓋率、最拖分考點、每日覆蓋進度(target=及格參考線)。
| Name | Required | Description | Default |
|---|---|---|---|
| daily | No | ||
| q_type | No | mcq | |
| target | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It does reveal meaningful behavior: frequency-weighted estimation, coverage calculation, weak-point identification, and daily progress tracking. However, it does not explicitly state that the tool is read-only, what data it depends on, or how the output is presented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the main purpose and then lists the key outputs. It avoids fluff, though the dense list of metrics could be slightly clearer with structural separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no parameter descriptions, the description is incomplete. It lists output concepts but not their format, and it fails to define 'daily' and 'q_type' or explain how the readiness score should be interpreted beyond the target reference. An agent would need to guess important details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only explains the 'target' parameter as the passing reference line. The 'daily' and 'q_type' parameters remain unexplained, leaving their meaning and acceptable values ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's purpose: estimating exam readiness via weighted scores, coverage, weak points, and daily progress. It is specific about the resource and outputs, though it lacks an explicit verb and does not directly distinguish itself from close siblings like get_progress or get_weak_topics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a general 'check your readiness' use case but provides no guidance on when to choose this tool over alternatives such as get_progress, get_weak_topics, or get_study_plan. No exclusions or conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statute_frequencyC
法條考頻統計(可選 exam_code)。
| Name | Required | Description | Default |
|---|---|---|---|
| exam_code | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It fails to state whether the operation is read-only, what dimension of frequency is counted (per exam, per question, aggregated across all exams), what the response shape is, or how the optional exam_code changes the result. For a data-retrieval tool, this is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no waste. However, this brevity is closer to under-specification than earned conciseness — the few words present do not carry enough information to justify the size.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no annotations and no output schema, the description was the last line of defense for the agent. It leaves open what the output contains, how exam_code filtering works, and how this tool relates to overlapping siblings. An agent would likely have to call the tool blindly to learn these details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only repeats that exam_code is optional (可選), which is already visible in the schema's default: null. It does not explain what an exam_code is, where to obtain valid values (e.g., from list_exams), or how the filter affects the aggregation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description "法條考頻統計(可選 exam_code)" clearly states the tool's resource (statute/legal provision exam frequency) and action (statistics/aggregation). The purpose is understandable on its own, though it does not differentiate from siblings such as get_topic_distribution or get_issue_distribution, which are also statistical reporting tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage hint is "可選 exam_code" (optional exam_code), which tells the agent a filter exists but provides no guidance on when to choose this tool over the many sibling statistical/reporting tools (get_topic_distribution, get_issue_distribution, get_exam_map). No exclusions or alternative routing is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_planA
讀書計畫引擎(綁考試日的處方):把「弱點×考點頻率」排成攻擊順序,依剩餘天數排出分相日程 (掃弱點→申論模擬→衝刺複習)+今日該做什麼+進度夠不夠的誠實判斷。days_remaining 請用 使用者的考試日期減今天算出。非及格保證,作答數據越多越準。
| Name | Required | Description | Default |
|---|---|---|---|
| daily | No | ||
| q_type | No | mcq | |
| target | No | ||
| days_remaining | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the disclosure burden. It honestly states limitations ('非及格保證') and data dependency ('作答數據越多越準'), and describes the scheduling behavior. It does not mention side effects or permission needs, but for a planning tool the key behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each clause earns its place: purpose, scheduling logic, output, parameter guidance, and caveat. It is front-loaded with the core purpose and does not repeat schema field names. Slightly long, but efficient for the information conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description gives enough context on output ('今日該做什麼+進度夠不夠'), the main parameter, and the tool's behavior. What it lacks is explanation of the other three parameters and any return structure details, but the most essential context for calling the tool is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only explains days_remaining ('use exam date minus today'), leaving daily, q_type, and target entirely unexplained. This is a significant gap for an agent trying to select correct values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a study plan engine that turns weaknesses × exam-point frequency into an attack order, creates a phased schedule, and tells what to do today plus progress judgment. This distinguishes it from siblings like get_progress or get_weak_topics, which only retrieve raw data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to compute days_remaining and sets expectations (not a pass guarantee, accuracy depends on answer data), but it does not explicitly say when to use this tool versus alternatives like get_weak_topics or practice_weak. The context is helpful but the when/not-when advice is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_topic_distributionA
考點熱度排行:各 (子科目, 考點) 考過幾題(q_type='mcq'/'essay',exam_code='sl1'/'sl2')。
| Name | Required | Description | Default |
|---|---|---|---|
| q_type | No | ||
| exam_code | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose that this is a read-only ranking of question counts per topic, with optional q_type and exam_code values. It does not explicitly state null-filter behavior, but the schema defaults and optionality make this inferable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The core behavior and filter values are communicated in a compact and scannable way.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read-only tool with two optional parameters and an output schema, the description is nearly sufficient. The only missing detail is explicit handling of omitted parameters, which is adequately implied by schema defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate; it does by specifying allowed values for q_type ('mcq'/'essay') and exam_code ('sl1'/'sl2'), which are absent from the schema. It does not fully explain null/default semantics, but the schema already indicates defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific aggregation task: a topic-heat ranking counting questions per (subsubject, topic). It also names the relevant dimensions and filter values, making the tool's scope distinct from siblings like get_issue_distribution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternative siblings such as get_issue_distribution or get_weak_topics. The filters are described, but no context is given for choosing this tool in a workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_topic_primerC
考點重點提示:做題前必讀的核心法條/常考判決釋字/學說對立/易錯陷阱。
| Name | Required | Description | Default |
|---|---|---|---|
| topic_point | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes the content (core statutes, judgments, doctrines, traps) but doesn't disclose behavioral traits such as whether it returns a summary, whether it requires prior progress, whether it's read-only, or what the output format is. For a tool that likely just returns study material, the lack of behavioral detail is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the purpose ('考點重點提示') and lists the content categories. It is efficient and easy to parse, though it could benefit from a brief usage note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one parameter, no output schema, and no annotations, the description should provide more context about how to invoke it correctly. It tells the agent what the primer contains but not how to specify the topic, what the return value looks like, or how this differs from get_issue_primer. The description is incomplete for an agent to use it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the parameter 'topic_point' at all. The parameter name suggests it's a topic identifier, but the description doesn't clarify what values are valid, how to format them, or how they relate to the tool's purpose. With zero schema coverage and no parameter explanation, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific purpose: providing key points (core statutes, common judgments/interpretations, doctrinal conflicts, common traps) before answering questions. It clearly identifies the resource (topic primer) and its content. However, it doesn't explicitly differentiate from the sibling get_issue_primer, which likely serves a similar role for issues rather than topics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says '做題前必讀' (must-read before doing questions), which implies when to use it: before attempting practice questions. It doesn't explicitly state when not to use it or name alternatives like get_issue_primer. The usage context is implied but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_weak_topicsC
弱點地圖:依個人作答正確率,由弱到強列出各考點(含答對率)。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| q_type | No | mcq | |
| min_attempts | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
無注解提供,描述承擔全部責任。它說明了排序方向和包含答對率,但未說明數據來源、是否需要登入、返回格式或任何副作用。對於一個檢索型工具,行為揭露不足。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述非常簡短,一句話完成,沒有冗餘。但簡短導致資訊不足,未達到適當的規格程度,因此僅給3分。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
有輸出schema,但描述未解釋返回結構。參數未說明,使用情境未提供,整體對於一個3參數工具來說不完整。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
schema描述覆蓋率為0%,描述完全未提及參數(limit、q_type、min_attempts)的含義或用法。描述未補償schema的不足,參數語義完全缺失。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明確指出工具列出考點並按個人正確率由弱到強排序,包含答對率。動詞和資源明確,能與部分兄弟工具(如get_topic_distribution)區分,但未完全排除所有替代工具。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述未說明何時使用此工具而非其他工具,也未提及任何情境或前置條件。僅陳述功能,沒有使用指引。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_examsA
列出可查的考試別與年度範圍。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the burden. It conveys a read-only listing operation (non-destructive) implicitly but lacks explicit details like return format or no side effects. It's adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, informative sentence that is front-loaded with the main action and scope. No unnecessary words, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with no parameters and has an output schema (though its content is unknown). The description is sufficient for an agent to call it, but it could be more explicit about the format of the return value, which the output schema may or may not cover.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no semantic details. The description defines the scope of what is returned (exam types and year ranges), which is essential. This compensates fully for the absence of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool lists available exam types and year ranges, which is a clear verb-resource pair. It is distinguishable from siblings like list_subjects, though it doesn't explicitly name the alternative it differs from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to discover exam options) but does not provide explicit guidance on when not to use it or mention alternatives. Since it's a discovery tool, context is clear but exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_subjectsB
列出科目(可選 exam_code 篩選)。
| Name | Required | Description | Default |
|---|---|---|---|
| exam_code | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only mentions a list operation with an optional filter. It does not disclose whether the operation is read-only, whether pagination or ordering applies, or what the response looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. The core action and the optional filter are front-loaded, and every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list operation with one optional parameter and an existing output schema, the description covers the essential purpose and filter semantics. Additional details like pagination or ordering would be nice, but they are not critical for this tool's simple contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameter. It does add meaning by stating that exam_code is an optional filter, but it does not explain the expected format or allowed values, providing only partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states '列出科目' (list subjects), a specific verb plus resource, and adds the optional exam_code filter. It is distinguishable from sibling tools like list_exams by resource type, but it does not explicitly name alternatives, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus related siblings such as list_exams or get_exam_paper. The description only states what the tool does, leaving scenario selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
practice_by_topicB
依考點抽題練習(topic_point 為必填,如「抵押權(普通/最高限額)」)。
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| q_type | No | ||
| hide_answer | No | ||
| topic_point | Yes | ||
| topic_subject | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it only says 'practice by drawing questions' and does not disclose side effects, whether progress is recorded, or how question selection works. It also does not mention read-only vs. state-changing behavior, which matters for a practice tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence communicates the core action and the required parameter without wasted words. The parenthetical example is efficient and directly relevant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too sparse for a 6-parameter tool with no annotations: it covers only topic_point, leaving usage context, selection behavior, and non-topic parameters unaddressed. Even though an output schema exists, the overall context is incomplete for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds semantics only for topic_point by noting it is required and giving a format example. The other five parameters (n, seed, q_type, hide_answer, topic_subject) are left entirely to schema titles/defaults, and with 0% schema-description coverage the description fails to compensate for that gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('抽題練習' – practice by drawing questions) and resource ('依考點' by exam topic), and highlights the required topic_point. It is clear what the tool does, though it does not explicitly differentiate from siblings like random_practice or practice_weak.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that topic_point is required and provides a concrete example ('抵押權(普通/最高限額)'), implying the tool is for topic-specific drill. However, it does not mention when to prefer this over random_practice or practice_weak, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
practice_weakA
弱點練習:優先出「今天到期複習」與「最弱考點」的題目(hide_answer 預設隱藏答案)。
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| q_type | No | mcq | |
| subject | No | ||
| hide_answer | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the prioritization behavior (due-for-review and weakest points) and the hide_answer default, which are meaningful traits. It stops short of stating side-effect/read-only status, but nothing contradicts the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence front-loads the core behavior and adds one parenthetical default. There is no filler or repetition beyond the short label '弱點練習', which the following clause immediately justifies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core selection logic is clear and, since an output schema exists, return values are not a gap. However, with no annotations and several sibling practice tools, the lack of usage routing and undeclared n/q_type/subject semantics leaves the definition only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It only clarifies hide_answer's default; n, q_type, and subject semantics are left to inference from names and defaults, and q_type values are not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific function: generate practice questions while prioritizing today's due reviews and the weakest exam points. This distinguishes it from random_practice and practice_by_topic without naming them explicitly, though the verb is somewhat embedded in '出題'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: use it when the user wants weak-spot/due-review drilling. It provides no explicit when-to-use or when-not-to-use guidance, nor does it point to alternatives such as random_practice or practice_by_topic, despite many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
random_practiceC
依條件抽題練習(hide_answer 可隱藏答案/擬答)。
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| q_type | No | ||
| subject | No | ||
| exam_code | No | ||
| hide_answer | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It only notes that 'hide_answer' can hide the answer, but doesn't disclose behavioral details like whether it's read-only or mutates progress, whether it's deterministic with a seed, or what output format is returned (output schema exists, but that might convey structure). It provides minimal behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence, which is efficient. It is front-loaded with the core function. It doesn't waste words, but could be slightly more informative without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 6 parameters (though all optional), a relatively high complexity, and an output schema exists. However, with zero schema description coverage, the description should compensate by explaining what each parameter filters on and what the output represents. It omits critical context like the range of 'n', what 'q_type' values are valid, and how the output is structured, leaving the agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 6 parameters have 0% schema description coverage, meaning the schema provides no descriptive text. The description only mentions 'hide_answer' functionality, leaving 'n', 'seed', 'q_type', 'subject', and 'exam_code' without any semantic explanation. Since coverage is low, the description must compensate but fails.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says '依條件抽題練習' which indicates it is for practice by drawing questions based on conditions, but it doesn't specify the resource or the exact conditions. It doesn't differentiate itself from sibling tools like 'practice_by_topic' or 'practice_weak', which are also practice-related, leaving ambiguity about the specific use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies that it can be used for conditional practice, but does not give any guidance on when to use this tool versus the many sibling practice tools. It doesn't provide criteria for selection (e.g., 'use for random whole-pool practice, not for topic-specific drills').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_answerB
記錄一次作答並自動批改+更新間隔重複排程。MCQ 自動對答案;申論可傳 self_correct 自評。
| Name | Required | Description | Default |
|---|---|---|---|
| qid | Yes | ||
| answer | No | ||
| self_correct | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool mutates state (records an answer, updates scheduling) and that it auto-grades MCQ, with optional self-grading for essays. It does not disclose side effects like whether progress is overwritten, whether scheduling changes are irreversible, or whether repeated calls for the same qid are idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main action, then adds a useful conditional detail about MCQ vs essay. Every sentence earns its place, though the second sentence could be slightly more explicit about parameter usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description gives the core behavior but omits important context: what happens on repeated calls, whether grading is synchronous, and what the return value indicates. The sibling list shows many read-only tools, so the mutation nature is clear, but an agent might still need more detail to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains self_correct's role for essay questions, which adds meaning beyond the schema. However, it does not explain the 'answer' parameter's format or how qid relates to question types, leaving some parameter semantics to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('記錄' record) and resource ('作答' answer), and clearly distinguishes itself from sibling tools by mentioning automatic grading and spaced-repetition scheduling. It also differentiates MCQ vs essay behavior, which helps an agent understand the tool's core function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when recording an answer that should be graded and scheduled. It does not explicitly state when not to use it or name alternatives, but the sibling list includes many read-only tools, so the mutation purpose is fairly clear. However, it lacks explicit guidance on when to use self_correct or how this relates to other practice tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_progressA
清空所有作答記錄與複習排程(重新開始)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so well by explicitly stating that all answer records and review schedules are cleared. This is a clear destructive action, though it stops short of explicitly warning that the action is irreversible or describing what happens afterward.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that delivers the core destructive action immediately, with a clarifying parenthetical. Every part earns its place and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool, the description is largely complete: it names the precise resources affected and the intended restart purpose. However, because there are no annotations and no output schema, a brief note on irreversibility or confirmation would have made it fully complete for an agent deciding whether to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics to clarify. The description does not need to compensate for schema gaps; the schema is empty and complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action with a specific verb and resource: 清空所有作答記錄與複習排程 (clear all answer records and review schedule). The parenthetical '重新開始' reinforces the intent, and the destructive nature distinguishes it from the read-only sibling tools like get_progress and get_readiness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by '重新開始' (restart), suggesting the tool should be used when a user wants to begin anew. However, it does not explicitly say when not to use it, mention confirmation requirements, or name any alternatives, so guidance is only implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_by_issueA
依爭點找題:某爭點(如「不能未遂之判斷標準」)考過哪幾題,附各題的學說對立與實務見解。
| Name | Required | Description | Default |
|---|---|---|---|
| issue | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It describes the output as including scholarly debates and judicial opinions per question, which adds behavioral context. However, it does not specify return format, pagination, or whether results are exhaustive, but the existence of an output schema covers some of this.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loads the purpose and the example. The single-sentence structure is efficient, but it could be slightly more structured to separate the purpose from the output details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter search tool with an output schema, the description covers the core use case and what results include. The absence of usage guidelines is a minor gap, but the output schema and example provide enough context for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that the 'issue' parameter takes a legal issue keyword (e.g., '不能未遂之判斷標準'), which clarifies the expected input. However, it doesn't delve into syntax or format, but the example is helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (search) and resource (questions by issue), and gives an example of a specific issue. While it doesn't explicitly name a sibling tool, it is distinguishable from other search tools like search_by_statute and search_questions based on the issue-centric focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a user wants to find which questions cover a specific legal issue, and it mentions that results include doctrinal conflicts and judicial views. However, it does not explicitly contrast with sibling tools like search_questions or search_by_statute, leaving the exact selection criteria to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_by_statuteB
按法條反查考過哪些題。
| Name | Required | Description | Default |
|---|---|---|---|
| statute | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden, but it only states the search purpose. It does not mention whether the tool returns a list, how exact or fuzzy the statute matching is, whether it is read-only, or any side effects. The behavior is minimally implied but not transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short, direct sentence with no filler or redundant phrasing. It is front-loaded and easy to parse, though its brevity contributes to some missing contextual detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one required parameter and an output schema is present, so the description does not need to explain return structure. However, given the absence of annotations and 0% schema parameter coverage, the single-sentence description leaves meaningful gaps around parameter format and when to use this instead of related search tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented 'statute' parameter. It does indicate that the statute is the search key, but it does not explain expected format, whether article numbers or full statute names are needed, or whether partial matches are allowed. This is only minimal added meaning beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description '按法條反查考過哪些題' clearly states the action: reverse-lookup previously examined questions by statute. It names the resource (statute) and the outcome (which questions have been tested), which distinguishes it from sibling tools like search_by_issue, though it does not explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when a user has a statute and wants to find related past exam questions. However, it gives no explicit guidance on when to prefer this over search_questions, search_by_issue, or get_statute_frequency, and provides no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_questionsB
全文搜尋歷屆考題(匹配題幹/選項/擬答)。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full disclosure burden. It discloses the matching scope (stem/options/model answers), which is useful, but says nothing about read-only behavior, pagination, result ordering, rate limits, or auth requirements. For a search tool the read-only nature is inferable, but the description adds little beyond the match fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the core purpose (全文搜尋歷屆考題) and appends the matching detail in parentheses. Zero waste; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 params, output schema present, no nested objects), and the description covers the search behavior adequately. The main gap is the absence of guidance distinguishing it from search_by_statute and search_by_issue, which an agent navigating 28 siblings would rely on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It adds meaning to 'query' by specifying that it matches against the stem/options/model answers. The 'limit' parameter, however, receives no additional semantic detail — though its default of 20 and integer type make it fairly self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (全文搜尋/full-text search), a resource (歷屆考題/past exam questions), and the matched fields (題幹/選項/擬答). This scope signals a general full-text search tool and implicitly contrasts with siblings like search_by_statute and search_by_issue, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or exclusions are provided. With specialized sibling searches (search_by_statute, search_by_issue) available, the agent is not told when to prefer this free-text search over those alternatives. Usage is only implied by the word 'search' in the title and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
29 tool updates
v0.6.0- First observed
essay_exam_by_topic - First observed
get_answer_key - First observed
get_error_diagnosis - First observed
get_exam_map - First observed
get_exam_paper - First observed
get_grading_rubric - First observed
get_issue_chain - First observed
get_issue_distribution - First observed
get_issue_primer - First observed
get_issues - First observed
get_model_answer - First observed
get_progress - First observed
get_question - First observed
get_readiness - First observed
get_statute_frequency - First observed
get_study_plan - First observed
get_topic_distribution - First observed
get_topic_primer - First observed
get_weak_topics - First observed
list_exams - First observed
list_subjects - First observed
practice_by_topic - First observed
practice_weak - First observed
random_practice - First observed
record_answer - First observed
reset_progress - First observed
search_by_issue - First observed
search_by_statute - First observed
search_questions
TDQS
Scored across 29 tools
Most tools have clearly distinct purposes, but a few pairs could be confused, such as get_issue_primer vs get_topic_primer and get_issue_distribution vs get_topic_distribution, which both offer preparatory/statistical content centered on either issues or topics. Descriptions are detailed enough to resolve most ambiguities, but the overlap is notable.
The vast majority follow snake_case verb_noun conventions (e.g., get_readiness, list_subjects, search_questions, practice_by_topic). Minor deviations like random_practice and essay_exam_by_topic break the pattern slightly, but overall the naming is predictable and readable.
At 29 tools, the surface is on the heavy side, but each tool serves a distinct function within a comprehensive exam preparation workflow—from question retrieval and practice to analytics, grading, and study planning. The count is justified by the breadth of the domain, though consolidation of a few overlapping utilities could make it tighter.
The tool set covers the full exam prep lifecycle: exam and question discovery, practice in multiple modes, answer recording and grading, progress tracking, weakness analysis, study planning, and detailed issue breakdowns. There are no obvious dead ends or missing core operations for the stated purpose.
Maintenance
Related MCP Connectors
Taiwan legal research MCP: 判決書、全國法規、釋字/憲判與立法歷程查詢,12 個工具,回應均附官方出處 URL。
Taiwan legal research: court judgments, statutes, and interpretations. 台灣判決、法條、函釋、釋字搜尋。
台灣繁中:一個 MCP 端點串接 19 個台灣資料站工具,並可搜尋 MCP 伺服器與 x402 付費 API。
Remote MCP server for deterministic educational practice-assessment score conversions.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that provides tools to search for study questions and answers from a knowledge base, enabling exam preparation through random questions and guided learning.5-
- AlicenseAqualityAmaintenanceMCP server for structured civil service exam question solving based on Huasheng's methodology, providing question routing, method retrieval, and guided analysis prompts.1526MIT
- AlicenseNot gradedqualityDmaintenance一个优化过的台湾法规查询MCP服务器,提供高效的法规搜索、条文查询和关键字搜索功能,支持摘要模式减少token消耗。MIT
- FlicenseNot gradedqualityDmaintenanceA server that queries Taiwan uniform invoice winning numbers and performs automatic prize matching.-