Skip to main content
Glama
mo9652962-ai

esq-builder-mcp

by mo9652962-ai

esq-builder-mcp

ESQ 1.0 题库包 MCP 工具链:把 [esq-question-bank-import] 技能的确定性环节(构建/校验/上传/词表分析)固化为 MCP 工具,供任意 MCP 客户端(ZCode / Claude Desktop / Codex 等)调用。

为什么

技能(SKILL.md)传的是流程知识,LLM 每次执行都可能踩坑(ASCII key、双花括号、上传路径 405……)。本 server 把这些坑固化进工具代码——调用方不会再遇到它们。

工具

作用

固化的坑

esq_build_package

校验 + 打包 ESQ ZIP(可选 auto_fix)

externalKey 纯 ASCII 3-200 位;{{blank:N}} 双花括号;option/candidates key 单大写字母;correctOption 必须存在于选项;cloze 空位数=题数;manifest 必填字段 + semver

esq_validate_package

校验 ESQ 包(默认内置校验器,可选官方 CLI 对账)

双轨校验(见下)

esq_upload_and_publish

上传 + 发布到刷题机后端

路径写死 /api/question-banks/imports(/upload 会 405);503 重试 3 次间隔 10s;publish 失败时提示用 job_id 单独重试

esq_parse_wordlist

kajweb/dict JSONL 高频词解析(.jsonl 或 book zip)

逐行 json.loads(整文件 load 报 Extra data);wordRank 排序

esq_hot_words

真题 passage 热点词统计

近两年过滤;去停用词;[a-zA-Z][a-zA-Z'-]{3,}

auto_fix:机械性坑自动修复

esq_build_package(auto_fix=true) 在校验前自动修复「纯机械」的坑,修复明细记录在返回值 fixes 数组(审计):

  • 含中文/非法字符的 packageId/paperKey/unitKey/questionKey → cn.xxx.y2021.u1 风格重建,answers 两级键自动同步改名

  • 单花括号 {blank:N} → 双花括号 {{blank:N}}

  • 缺失的 unit.sequence 补 index+1;不达标 blockKey(如 2 位的 p1)归一为 block-{index}

判断性问题(空位数≠题数、答案不在选项中)不会被静默修复,仍走「拒绝 + 可行动错误」。默认 false 保持严格行为。

双轨校验

esq_validate_package 有两条通道,返回值 validator 字段标明所用通道:

  • 默认:内置校验器(esq_validator.py,vendor 自 backend/app/services/esq.py 校验子集,import 调用)——零外部依赖,PyPI/uvx/PyInstaller 分发可用;

  • 对账:官方 CLI——显式传 validator_path 或设 ESQ_VALIDATOR_PATH 时走 subprocess 调官方校验器。

两条通道的一致性由 tests/test_validator_conformance.py 守护(本机有刷题机仓库时自动执行;后端校验逻辑变更后先跑它再同步 vendored 副本)。

Related MCP server: data-transformer

安装与运行

# PyPI(任意 MCP 客户端, 无需 clone)
uvx esq-builder-mcp              # stdio 模式

# Windows 单文件 exe: 到 Releases 下载 esq-builder-mcp.exe, 客户端 command 直指该 exe
# 源码方式
cd D:/esq-builder-mcp
uv venv && uv pip install -e ".[dev]"
uv run esq-builder-mcp

发布新版本

  1. bump pyproject.toml 的 version(PyPI 不允许同版本重传)

  2. git tag v0.1.1 && git push origin v0.1.1 → GitHub Actions 自动 build + 发布(Trusted Publishing,无 token)

  3. Windows exe: uv run python scripts/build_exe.py,产物 dist/esq-builder-mcp.exe,附到对应 Release

一次性配置: PyPI 项目 Settings → Publishing 配 Trusted Publisher(Owner=mo9652962-ai / Repository=esq-builder-mcp / Workflow name=publish.yml / Environment=pypi)

注册到 MCP 客户端

ZCode(~/.zcode/cli/config.json → mcpServers)或其他客户端:

{
  "mcpServers": {
    "esq-builder": {
      "command": "uv",
      "args": ["--directory", "D:/esq-builder-mcp", "run", "esq-builder-mcp"]
    }
  }
}

Windows 下 MCP 命令参数一律用正斜杠路径(Codex config.toml 转义坑的同款规避)。

环境变量

变量

默认

说明

ESQ_VALIDATOR_PATH

(未设)

设定后 esq_validate_package 改走官方校验器 CLI(对账/仲裁通道);默认内置校验器,不需要此变量

测试

uv run pytest -v          # 35 项;含 vendored vs 官方 CLI 一致性对账(无刷题机环境自动 skip)

后续演进

  • 发布到 PyPI ✅ 已发布 pypi.org/project/esq-builder-mcp,uvx esq-builder-mcp 一行接入(实测冷启动 stdio 握手 5 工具齐全)。

  • ESQ 1.1 examType:manifest.papers[].examType 已在官方校验器支持,构造器暂未暴露。

  • Windows 单文件 exe:走 PyInstaller(复用刷题机发布经验)。

与技能的关系

  • 上游技能:~/.agents/skills/esq-question-bank-import/SKILL.md(流程与数据源)

  • 本 server 是其「确定性环节」的工具化;AI 标注答案(基元律动)等 LLM 判断环节仍在技能侧。

演进记录

  • 2026-09-28:校验改双轨(vendored 默认 + 官方 CLI 对账),解除对刷题机仓库路径的运行时依赖,PyPI 分发解锁;esq_build_package 增加 auto_fix 通道;esq_parse_wordlist 支持 book zip 输入。

Available Tools

5 tools
esq_build_packageEsq Build PackageC

校验并构建 ESQ 1.0 题库包 ZIP。

ParametersJSON Schema
NameRequiredDescriptionDefault
papersYes[{paperKey, year, units:[{unitKey, type(cloze|reading|part_b), title, sequence, passage?, candidates?, questions:[...]}]}] - cloze: passage.blocks[].text 用 {{blank:N}} 双花括号标记空位(不是单花括号!); 词库题用 unit.candidates=[{key:"A",content:"..."}], questions 不写 options - reading: questions 写 options=[{key:"A",content:"..."}] - question: {questionKey, number, type:"single_choice", stem, score}
answersYes{paperKey: {questionKey: {correctOption:"B", score: 2}}} —— 每空必填, correctOption 必须存在于该题选项
auto_fixNotrue 时先自动修复机械性坑再构建——含中文的 externalKey 重建为 cn.xxx.y2021.u1 风格(answers 键同步改名)、单花括号 {blank:N} 转双花括号、 缺失 sequence 补默认值。修复明细在返回值 fixes 数组; 判断性问题(空位数≠题数、 答案不在选项中)仍会报错。默认 false(拒绝 + 可行动错误)。
manifestYes{packageId, contentVersion, title, subject, publisher, license:{notice}, source:{description}}
output_pathYes输出 ZIP 绝对路径

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but discloses almost nothing: it does not say whether validation failures abort the build, whether an existing ZIP at output_path is overwritten, or what permissions/auth are needed. The auto_fix semantics live only in the schema, not here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero filler and the action front-loaded. It is efficient, though its brevity borders on under-specification for a tool with five parameters and nested input.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the schema fully documents inputs. But for a mutating, file-producing tool with no annotations, the description omits failure behavior, overwrite semantics, and the validate/build pipeline relationship — gaps that matter when an agent decides whether to call this or the validator.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the deeply nested papers/answers/manifest structures and the auto_fix behavior. The description adds no parameter meaning beyond the schema, which is the correct baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific compound verb (校验并构建 / validate and build) and a precise resource (ESQ 1.0 题库包 ZIP), so an agent knows this produces a ZIP artifact rather than just checking one. However, it never contrasts itself with the sibling esq_validate_package, leaving the agent to infer that this tool is the validate-then-build composite.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all: nothing says when to prefer this over esq_validate_package (dry-run validation) or when to follow it with esq_upload_and_publish. The relationship between validation and building is only implied by the word 构建.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esq_hot_wordsEsq Hot WordsB

真题 passage 文本 → 热点词频统计(去停用词, [a-zA-Z][a-zA-Z'-]{3,})。

ParametersJSON Schema
NameRequiredDescriptionDefault
textsYes[{year: 2024, text: "passage 正文"}, ...](或纯字符串列表)
top_nNo取前 N 个(默认 300)
since_yearNo只统计该年份之后的真题(默认 2023, 即"近两年热点")
extra_stopwordsNo额外停用词

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral burden. It discloses stopword removal and the token regex, but omits side-effect/auth info, output shape (covered by output schema), and determinism/read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very concise single line, front-loading the core input-to-output transformation before the filtering details. No wasted words, though the parenthetical filter is dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a complete input schema and an output schema, the description covers the core operation and filtering logic. It remains thin on integration context, but the structured fields cover the parameter and return details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are documented in the schema. The description adds no parameter-specific semantics beyond the processing rule, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Describes a specific transformation from exam passage text to hot word frequency statistics, including stopword removal and a regex filter. It does not explicitly distinguish itself from siblings like esq_parse_wordlist, but the computation is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The only contextual signpost is the input type '真题 passage 文本'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esq_parse_wordlistEsq Parse WordlistC

解析 kajweb/dict 高频词表 JSONL(wordRank=真题词频排序)。

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNo标注级别名(如 "四级·高频")
top_nNo只取前 N 个(如四级核心词取 1162)
jsonl_pathYeskajweb book/*.zip 解压后的 JSONL 文件路径

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and mostly does not meet it: it does not state that the operation is a read-only parse, whether it writes anything to disk, or how it behaves on malformed JSONL. It only clarifies the meaning of wordRank, which is domain metadata rather than behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence with the resource front-loaded and no filler; the parenthetical about wordRank earns its place. It is arguably too terse to be maximally useful, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and parameters are fully documented, so return-value explanation is unnecessary. What is missing is pipeline context: which sibling precedes and follows this parse step, and what precondition the jsonl_path implies. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 and the schema already documents level, top_n and jsonl_path. The description adds only the wordRank semantics, which is a useful gloss on the output ordering rather than meaning for any specific parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (解析/parse) and resource (kajweb/dict 高频词表 JSONL), so the agent knows exactly what is being processed. It does not, however, distinguish this tool from the nearest sibling esq_hot_words, leaving the boundary between 'parse a wordlist' and 'get hot words' to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated prerequisite (e.g. that the JSONL must first be extracted from book/*.zip), and no reference to any alternative such as esq_hot_words or esq_build_package. The agent must guess where this fits in the esq_* pipeline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esq_upload_and_publishEsq Upload And PublishA

上传 ESQ 包到刷题机后端并(可选)发布。上传前先 esq_validate_package。

ParametersJSON Schema
NameRequiredDescriptionDefault
publishNo是否发布(默认 true)
base_urlNo后端地址(默认 http://127.0.0.1:8765, 需先启动后端)http://127.0.0.1:8765
zip_pathYesESQ 包路径
profile_idNo题库 profile(不传则用后端当前激活级别)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the basic action and an optional flag; it does not describe side effects (e.g., whether publishing replaces existing content, permission requirements, or error behavior), nor does it mention that the backend must be running (that detail is only in the schema). For a mutation tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action and followed by a prerequisite. Every sentence earns its place with no waste or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there are no annotations and this is a mutation tool (upload and optionally publish), the description should ideally disclose more about side effects, permissions, and the meaning of 'publish.' It covers the basic workflow and names a prerequisite, but misses behavioral context that an agent would need to invoke it safely. The presence of an output schema means return values need not be explained, which partially offsets the gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including defaults and the base_url prerequisite. The description adds no parameter syntax or format details beyond what the schema provides. The baseline of 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'upload ESQ package to the quiz machine backend and (optionally) publish.' It names the sibling tool esq_validate_package as a prerequisite, which helps distinguish it from other build/validate siblings. However, it does not explicitly differentiate from esq_build_package, leaving some ambiguity about whether this tool also builds packages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it should be used after running esq_validate_package. It also notes that publishing is optional. But it does not state when not to use this tool or list alternative tools for other scenarios, so it stops short of explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

esq_validate_packageEsq Validate PackageB

校验 ESQ 包(默认内置校验器, 与刷题机官方 esq.py 同规则)。

ParametersJSON Schema
NameRequiredDescriptionDefault
zip_pathYesESQ 包路径
validator_pathNo可选, 指定后改用官方校验器 CLI(backend/tools/validate_question_bank.py) subprocess 对账——默认内置实现已与其保持一致并由一致性测试守护

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. It adds that the default built-in validator follows the same rules as the official esq.py, but does not disclose side effects, permissions, error handling, or whether the operation is read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence states the purpose and default behavior with no redundant text. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, has a full input schema and an output schema, so return values and parameters are covered elsewhere. However, with no annotations and no usage context, the description leaves gaps around when to choose it and its read-only/safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the two parameters are already documented in the input schema. The description adds no parameter-level meaning beyond the schema, which matches the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (校验) and resource (ESQ 包), making the tool's action clear. It does not explicitly contrast with sibling tools such as esq_build_package or esq_upload_and_publish, so it stops short of full 5-level differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool instead of building, uploading, parsing, or hot-word siblings, and no prerequisites. It only describes a default validator behavior, not selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.1
    • First observedesq_build_package
    • First observedesq_hot_words
    • First observedesq_parse_wordlist
    • First observedesq_upload_and_publish
    • First observedesq_validate_package

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation4/5

esq_build_package and esq_validate_package overlap slightly since build also validates, but the descriptions clarify that build produces a ZIP while validate is a standalone check used as a pre-upload step. esq_parse_wordlist and esq_hot_words both deal with word frequency but target clearly different inputs (wordlist JSONL vs passage text).

Naming Consistency4/5

All tools share the esq_ prefix and mostly follow a verb_noun pattern (build_package, validate_package, upload_and_publish, parse_wordlist). esq_hot_words breaks the pattern as a bare noun phrase, a minor deviation.

Tool Count5/5

Five tools is well-scoped for a package builder/publisher, and each tool earns its place in the build-validate-upload pipeline plus two auxiliary analysis utilities.

Completeness4/5

The core lifecycle (build, validate, upload/publish) is covered, and the word-frequency helpers round out the domain. Minor gaps exist for listing/managing previously built packages, but the primary workflow has no dead ends.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Zero-dependency stdio bridge to Moltline Studio's fleet of 14 hosted MCP servers covering code review, time operations, data transforms, business ops, education, research, outreach and more. Free tier requires no registration; premium tools unlock with a license. Independently audited, MCPize Verified A.
    10
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Provides safe, deterministic inspection, transformation, validation, and diffing of structured data (JSON, CSV, YAML, Parquet) via schema-aware MCP tools.
    4
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables local-first SEO evidence tooling by analyzing Screaming Frog exports, running bounded live and infrastructure checks, and producing structured audits, task backlogs, and reports via CLI and MCP.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables creating and administering multiple quizzes, adding questions, registering participants, validating answers, reviewing responses, and managing leaderboards through MCP, accessible locally via stdio or remotely over HTTP.
    -