esq-builder-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@esq-builder-mcpvalidate and auto-fix my ESQ question bank, then package it as a ZIP"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
esq-builder-mcp
ESQ 1.0 题库包 MCP 工具链:把 [esq-question-bank-import] 技能的确定性环节(构建/校验/上传/词表分析)固化为 MCP 工具,供任意 MCP 客户端(ZCode / Claude Desktop / Codex 等)调用。
为什么
技能(SKILL.md)传的是流程知识,LLM 每次执行都可能踩坑(ASCII key、双花括号、上传路径 405……)。本 server 把这些坑固化进工具代码——调用方不会再遇到它们。
工具 | 作用 | 固化的坑 |
| 校验 + 打包 ESQ ZIP(可选 | externalKey 纯 ASCII 3-200 位; |
| 校验 ESQ 包(默认内置校验器,可选官方 CLI 对账) | 双轨校验(见下) |
| 上传 + 发布到刷题机后端 | 路径写死 |
| kajweb/dict JSONL 高频词解析(.jsonl 或 book zip) | 逐行 json.loads(整文件 load 报 Extra data);wordRank 排序 |
| 真题 passage 热点词统计 | 近两年过滤;去停用词; |
auto_fix:机械性坑自动修复
esq_build_package(auto_fix=true) 在校验前自动修复「纯机械」的坑,修复明细记录在返回值 fixes 数组(审计):
含中文/非法字符的 packageId/paperKey/unitKey/questionKey →
cn.xxx.y2021.u1风格重建,answers 两级键自动同步改名单花括号
{blank:N}→ 双花括号{{blank:N}}缺失的
unit.sequence补 index+1;不达标 blockKey(如 2 位的p1)归一为block-{index}
判断性问题(空位数≠题数、答案不在选项中)不会被静默修复,仍走「拒绝 + 可行动错误」。默认 false 保持严格行为。
双轨校验
esq_validate_package 有两条通道,返回值 validator 字段标明所用通道:
默认:内置校验器(
esq_validator.py,vendor 自 backend/app/services/esq.py 校验子集,import 调用)——零外部依赖,PyPI/uvx/PyInstaller 分发可用;对账:官方 CLI——显式传
validator_path或设ESQ_VALIDATOR_PATH时走 subprocess 调官方校验器。
两条通道的一致性由 tests/test_validator_conformance.py 守护(本机有刷题机仓库时自动执行;后端校验逻辑变更后先跑它再同步 vendored 副本)。
Related MCP server: data-transformer
安装与运行
# PyPI(任意 MCP 客户端, 无需 clone)
uvx esq-builder-mcp # stdio 模式
# Windows 单文件 exe: 到 Releases 下载 esq-builder-mcp.exe, 客户端 command 直指该 exe
# 源码方式
cd D:/esq-builder-mcp
uv venv && uv pip install -e ".[dev]"
uv run esq-builder-mcp发布新版本
bump
pyproject.toml的version(PyPI 不允许同版本重传)git tag v0.1.1 && git push origin v0.1.1→ GitHub Actions 自动 build + 发布(Trusted Publishing,无 token)Windows exe:
uv run python scripts/build_exe.py,产物dist/esq-builder-mcp.exe,附到对应 Release
一次性配置: PyPI 项目 Settings → Publishing 配 Trusted Publisher(Owner=mo9652962-ai / Repository=esq-builder-mcp / Workflow name=publish.yml / Environment=pypi)
注册到 MCP 客户端
ZCode(~/.zcode/cli/config.json → mcpServers)或其他客户端:
{
"mcpServers": {
"esq-builder": {
"command": "uv",
"args": ["--directory", "D:/esq-builder-mcp", "run", "esq-builder-mcp"]
}
}
}Windows 下 MCP 命令参数一律用正斜杠路径(Codex config.toml 转义坑的同款规避)。
环境变量
变量 | 默认 | 说明 |
| (未设) | 设定后 |
测试
uv run pytest -v # 35 项;含 vendored vs 官方 CLI 一致性对账(无刷题机环境自动 skip)后续演进
发布到 PyPI✅ 已发布 pypi.org/project/esq-builder-mcp,uvx esq-builder-mcp一行接入(实测冷启动 stdio 握手 5 工具齐全)。ESQ 1.1 examType:manifest.papers[].examType 已在官方校验器支持,构造器暂未暴露。
Windows 单文件 exe:走 PyInstaller(复用刷题机发布经验)。
与技能的关系
上游技能:
~/.agents/skills/esq-question-bank-import/SKILL.md(流程与数据源)本 server 是其「确定性环节」的工具化;AI 标注答案(基元律动)等 LLM 判断环节仍在技能侧。
演进记录
2026-09-28:校验改双轨(vendored 默认 + 官方 CLI 对账),解除对刷题机仓库路径的运行时依赖,PyPI 分发解锁;
esq_build_package增加auto_fix通道;esq_parse_wordlist支持 book zip 输入。
Available Tools
5 toolsesq_build_packageEsq Build PackageC
校验并构建 ESQ 1.0 题库包 ZIP。
| Name | Required | Description | Default |
|---|---|---|---|
| papers | Yes | [{paperKey, year, units:[{unitKey, type(cloze|reading|part_b), title, sequence, passage?, candidates?, questions:[...]}]}] - cloze: passage.blocks[].text 用 {{blank:N}} 双花括号标记空位(不是单花括号!); 词库题用 unit.candidates=[{key:"A",content:"..."}], questions 不写 options - reading: questions 写 options=[{key:"A",content:"..."}] - question: {questionKey, number, type:"single_choice", stem, score} | |
| answers | Yes | {paperKey: {questionKey: {correctOption:"B", score: 2}}} —— 每空必填, correctOption 必须存在于该题选项 | |
| auto_fix | No | true 时先自动修复机械性坑再构建——含中文的 externalKey 重建为 cn.xxx.y2021.u1 风格(answers 键同步改名)、单花括号 {blank:N} 转双花括号、 缺失 sequence 补默认值。修复明细在返回值 fixes 数组; 判断性问题(空位数≠题数、 答案不在选项中)仍会报错。默认 false(拒绝 + 可行动错误)。 | |
| manifest | Yes | {packageId, contentVersion, title, subject, publisher, license:{notice}, source:{description}} | |
| output_path | Yes | 输出 ZIP 绝对路径 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but discloses almost nothing: it does not say whether validation failures abort the build, whether an existing ZIP at output_path is overwritten, or what permissions/auth are needed. The auto_fix semantics live only in the schema, not here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero filler and the action front-loaded. It is efficient, though its brevity borders on under-specification for a tool with five parameters and nested input.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema fully documents inputs. But for a mutating, file-producing tool with no annotations, the description omits failure behavior, overwrite semantics, and the validate/build pipeline relationship — gaps that matter when an agent decides whether to call this or the validator.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the deeply nested papers/answers/manifest structures and the auto_fix behavior. The description adds no parameter meaning beyond the schema, which is the correct baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific compound verb (校验并构建 / validate and build) and a precise resource (ESQ 1.0 题库包 ZIP), so an agent knows this produces a ZIP artifact rather than just checking one. However, it never contrasts itself with the sibling esq_validate_package, leaving the agent to infer that this tool is the validate-then-build composite.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance at all: nothing says when to prefer this over esq_validate_package (dry-run validation) or when to follow it with esq_upload_and_publish. The relationship between validation and building is only implied by the word 构建.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esq_hot_wordsEsq Hot WordsB
真题 passage 文本 → 热点词频统计(去停用词, [a-zA-Z][a-zA-Z'-]{3,})。
| Name | Required | Description | Default |
|---|---|---|---|
| texts | Yes | [{year: 2024, text: "passage 正文"}, ...](或纯字符串列表) | |
| top_n | No | 取前 N 个(默认 300) | |
| since_year | No | 只统计该年份之后的真题(默认 2023, 即"近两年热点") | |
| extra_stopwords | No | 额外停用词 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. It discloses stopword removal and the token regex, but omits side-effect/auth info, output shape (covered by output schema), and determinism/read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise single line, front-loading the core input-to-output transformation before the filtering details. No wasted words, though the parenthetical filter is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a complete input schema and an output schema, the description covers the core operation and filtering logic. It remains thin on integration context, but the structured fields cover the parameter and return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are documented in the schema. The description adds no parameter-specific semantics beyond the processing rule, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Describes a specific transformation from exam passage text to hot word frequency statistics, including stopword removal and a regex filter. It does not explicitly distinguish itself from siblings like esq_parse_wordlist, but the computation is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named. The only contextual signpost is the input type '真题 passage 文本'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esq_parse_wordlistEsq Parse WordlistC
解析 kajweb/dict 高频词表 JSONL(wordRank=真题词频排序)。
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | 标注级别名(如 "四级·高频") | |
| top_n | No | 只取前 N 个(如四级核心词取 1162) | |
| jsonl_path | Yes | kajweb book/*.zip 解压后的 JSONL 文件路径 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and mostly does not meet it: it does not state that the operation is a read-only parse, whether it writes anything to disk, or how it behaves on malformed JSONL. It only clarifies the meaning of wordRank, which is domain metadata rather than behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with the resource front-loaded and no filler; the parenthetical about wordRank earns its place. It is arguably too terse to be maximally useful, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and parameters are fully documented, so return-value explanation is unnecessary. What is missing is pipeline context: which sibling precedes and follows this parse step, and what precondition the jsonl_path implies. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already documents level, top_n and jsonl_path. The description adds only the wordRank semantics, which is a useful gloss on the output ordering rather than meaning for any specific parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (解析/parse) and resource (kajweb/dict 高频词表 JSONL), so the agent knows exactly what is being processed. It does not, however, distinguish this tool from the nearest sibling esq_hot_words, leaving the boundary between 'parse a wordlist' and 'get hot words' to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no stated prerequisite (e.g. that the JSONL must first be extracted from book/*.zip), and no reference to any alternative such as esq_hot_words or esq_build_package. The agent must guess where this fits in the esq_* pipeline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esq_upload_and_publishEsq Upload And PublishA
上传 ESQ 包到刷题机后端并(可选)发布。上传前先 esq_validate_package。
| Name | Required | Description | Default |
|---|---|---|---|
| publish | No | 是否发布(默认 true) | |
| base_url | No | 后端地址(默认 http://127.0.0.1:8765, 需先启动后端) | http://127.0.0.1:8765 |
| zip_path | Yes | ESQ 包路径 | |
| profile_id | No | 题库 profile(不传则用后端当前激活级别) |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the basic action and an optional flag; it does not describe side effects (e.g., whether publishing replaces existing content, permission requirements, or error behavior), nor does it mention that the backend must be running (that detail is only in the schema). For a mutation tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and followed by a prerequisite. Every sentence earns its place with no waste or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there are no annotations and this is a mutation tool (upload and optionally publish), the description should ideally disclose more about side effects, permissions, and the meaning of 'publish.' It covers the basic workflow and names a prerequisite, but misses behavioral context that an agent would need to invoke it safely. The presence of an output schema means return values need not be explained, which partially offsets the gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including defaults and the base_url prerequisite. The description adds no parameter syntax or format details beyond what the schema provides. The baseline of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'upload ESQ package to the quiz machine backend and (optionally) publish.' It names the sibling tool esq_validate_package as a prerequisite, which helps distinguish it from other build/validate siblings. However, it does not explicitly differentiate from esq_build_package, leaving some ambiguity about whether this tool also builds packages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it should be used after running esq_validate_package. It also notes that publishing is optional. But it does not state when not to use this tool or list alternative tools for other scenarios, so it stops short of explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esq_validate_packageEsq Validate PackageB
校验 ESQ 包(默认内置校验器, 与刷题机官方 esq.py 同规则)。
| Name | Required | Description | Default |
|---|---|---|---|
| zip_path | Yes | ESQ 包路径 | |
| validator_path | No | 可选, 指定后改用官方校验器 CLI(backend/tools/validate_question_bank.py) subprocess 对账——默认内置实现已与其保持一致并由一致性测试守护 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral disclosure. It adds that the default built-in validator follows the same rules as the official esq.py, but does not disclose side effects, permissions, error handling, or whether the operation is read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence states the purpose and default behavior with no redundant text. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, has a full input schema and an output schema, so return values and parameters are covered elsewhere. However, with no annotations and no usage context, the description leaves gaps around when to choose it and its read-only/safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the two parameters are already documented in the input schema. The description adds no parameter-level meaning beyond the schema, which matches the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (校验) and resource (ESQ 包), making the tool's action clear. It does not explicitly contrast with sibling tools such as esq_build_package or esq_upload_and_publish, so it stops short of full 5-level differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this tool instead of building, uploading, parsing, or hot-word siblings, and no prerequisites. It only describes a default validator behavior, not selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.1- First observed
esq_build_package - First observed
esq_hot_words - First observed
esq_parse_wordlist - First observed
esq_upload_and_publish - First observed
esq_validate_package
TDQS
Scored across 5 tools
esq_build_package and esq_validate_package overlap slightly since build also validates, but the descriptions clarify that build produces a ZIP while validate is a standalone check used as a pre-upload step. esq_parse_wordlist and esq_hot_words both deal with word frequency but target clearly different inputs (wordlist JSONL vs passage text).
All tools share the esq_ prefix and mostly follow a verb_noun pattern (build_package, validate_package, upload_and_publish, parse_wordlist). esq_hot_words breaks the pattern as a bare noun phrase, a minor deviation.
Five tools is well-scoped for a package builder/publisher, and each tool earns its place in the build-validate-upload pipeline plus two auxiliary analysis utilities.
The core lifecycle (build, validate, upload/publish) is covered, and the word-frequency helpers round out the domain. Minor gaps exist for listing/managing previously built packages, but the primary workflow has no dead ends.
Maintenance
Related MCP Connectors
- FormastyOAuthcom.formasty
Create, edit, validate, publish, and inspect Formasty forms and quizzes over authenticated MCP.
- KajabiOAuthcom.kajabi
Manage Kajabi from any MCP client — products, pages, contacts, offers, emails, analytics.
Structured analysis API and remote MCP tool for text, JSON records and numeric series.
Production-grade cryptography toolkit with 31 MCP tools for classical, PQC, and KMS workflows.
Related MCP Servers
- AlicenseAqualityAmaintenanceZero-dependency stdio bridge to Moltline Studio's fleet of 14 hosted MCP servers covering code review, time operations, data transforms, business ops, education, research, outreach and more. Free tier requires no registration; premium tools unlock with a license. Independently audited, MCPize Verified A.10MIT
- AlicenseAqualityBmaintenanceProvides safe, deterministic inspection, transformation, validation, and diffing of structured data (JSON, CSV, YAML, Parquet) via schema-aware MCP tools.4Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables local-first SEO evidence tooling by analyzing Screaming Frog exports, running bounded live and infrastructure checks, and producing structured audits, task backlogs, and reports via CLI and MCP.MIT
- FlicenseNot gradedqualityCmaintenanceEnables creating and administering multiple quizzes, adding questions, registering participants, validating answers, reviewing responses, and managing leaderboards through MCP, accessible locally via stdio or remotely over HTTP.-