skill-maintenance-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@skill-maintenance-mcpback up my skill-evolution directory before I edit it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
skill-maintenance-mcp
技能库维护 MCP server:把 [skill-evolution] 技能的机械环节工具化(备份/损坏扫描/决策日志/体检),供任意 MCP 客户端调用。蒸馏、修剪等判断力环节仍在技能侧。
对应 skill-evolution 五步闭环:
工具 | 对应环节 | 说明 |
| §7 改前备份 | cp 技能目录到 |
| §7 损坏扫描(铁律) | formfeed( |
| §2 决策历史 | 四字段(诊断/修订/证据/结果)追加 |
| §2 | 条目标题列表,修订前先看防重复踩坑 |
| §3/§4 体检 | frontmatter 必填 + 损坏扫描 + 孤儿 reference(未挂链=不可发现)+ decision-log 概要 |
安装与注册
cd D:/skill-maintenance-mcp
uv venv && uv pip install -e ".[dev]"ZCode(~/.zcode/cli/config.json → mcp.servers):
{
"mcp": {
"servers": {
"skill-maintenance": {
"command": "uv",
"args": ["--directory", "D:/skill-maintenance-mcp", "run", "skill-maintenance-mcp"]
}
}
}
}Related MCP server: MCP Gatekeeper
推荐工作流(agent 修订技能时)
skill_backup → 修订 → skill_scan_corruption → skill_log_decision → skill_validate测试
uv run pytest -v # 含对真实技能库的只读联动测试(无技能库环境自动 skip)设计边界
不做自动字节修复:修复需要人工判断正确序列(skill-evolution §7 的修复是逐案例的),工具只定位到行号 + kind + snippet。
备份默认拒绝同日覆盖:改两次时手动挪走第一份,防止快照语义被静默破坏。
上游技能:
~/.agents/skills/skill-evolution/SKILL.md;骨架方法见mcp-server-craft技能。
Available Tools
5 toolsskill_backupSkill BackupA
改前备份技能目录到 ~/.agents/skill-backups/<今天>/(skill-evolution 铁律: 修订前必备份)。
| Name | Required | Description | Default |
|---|---|---|---|
| backup_root | No | 备份目标根(默认 ~/.agents/skill-backups) | |
| skill_names | Yes | 技能目录名列表(如 ["esq-question-bank-import"]) | |
| skills_root | No | 技能库根目录 | C:/Users/31954/.agents/skills |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral disclosure. It states the action and destination path, but does not clarify whether backups overwrite existing copies, whether the operation is idempotent, how errors are handled, or whether the entire directory tree is copied. The output schema covers return values, but behavioral edge cases remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with a parenthetical rule. Every part earns its place: the operation, the destination pattern, and the mandatory timing are all stated without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, 3-parameter tool with full schema coverage and an output schema, the description covers purpose, timing, and destination adequately. The only minor gap is richer behavioral detail about backup overwrite/idempotency, but the structured fields carry most of the remaining burden.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, including defaults and examples. The description adds the backup path convention but no additional per-parameter semantics, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('备份' / back up) and resource ('技能目录' / skill directories), plus the destination pattern '~/.agents/skill-backups/<今天>/'. It is clearly distinguishable from sibling tools like skill_scan_corruption, skill_log_decision, skill_read_decisions, and skill_validate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it: before modification ('改前备份', '修订前必备份'), and frames this as an iron rule. There is no explicit when-not guidance or named alternative, but the timing condition is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_log_decisionSkill Log DecisionA
向技能的 references/decision-log.md 追加决策记录(四字段: 诊断/修订/证据/结果——记「为什么改」, 未来 agent 不重新踩坑)。
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | 一句话标题(自动加日期前缀) | |
| result | Yes | 结果——接受/拒绝 + 原因 | |
| evidence | Yes | 证据——评估/实测结果 | |
| revision | Yes | 修订——改了哪里 | |
| diagnosis | Yes | 诊断——技能在什么场景失效 | |
| skill_path | Yes | 技能目录绝对路径 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Append' usefully signals an additive, non-destructive write to a specific file path, which is meaningful context. It does not state whether the file is created if absent, whether entries are deduplicated on repeat calls, or what permissions are required for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the verb and target file front-loaded, followed by a parenthetical that explains both the record structure and the reason the tool exists. Nothing is wasted and no padding is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers purpose, target artifact, and record shape. Remaining gaps are edge cases such as missing-file creation and duplicate handling, which matter for a write tool with no annotations but are secondary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, including the four content fields. The description restates the four field names (diagnosis/revision/evidence/result) but adds no format, length, or content guidance beyond what the schema descriptions provide. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: append a decision record to the skill's references/decision-log.md, and names the exact artifact and the four fields it contains. It is clearly distinguishable from the read-side sibling skill_read_decisions, though it never explicitly names that counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys the underlying rationale (record 'why a change was made' so future agents don't repeat the mistake), which implies the appropriate moment to use it. However, there is no explicit when/when-not guidance and no reference to the read sibling skill_read_decisions, so routing relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_read_decisionsSkill Read DecisionsB
读取技能 decision-log 概要(条目标题列表),修订前先看避免重复踩坑。
| Name | Required | Description | Default |
|---|---|---|---|
| skill_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the return shape (a summary of titles rather than full entries), but says nothing about permissions, scope of the log, or size limits. It is a safe read, which limits risk, but disclosure is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the action and resource, with the usage rationale appended. No wasted words, though the terseness contributes to the gaps noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and this is a simple read tool. Still, the one required parameter's expected value is never addressed, which leaves the definition under-complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never explains what skill_path should contain (a directory? a skill identifier? a file path?). With one undocumented required parameter and no compensating text, an agent must guess the expected format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: reading the skill decision-log summary, and clarifies the granularity (entry title list). It is distinguishable from the sibling skill_log_decision (which writes) by the read/revise framing, though it never names siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an implied usage condition — review before revising to avoid repeating past mistakes — which is useful context. However, it does not name alternatives (e.g., skill_log_decision for writing) or state when not to use it, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_scan_corruptionSkill Scan CorruptionA
技能文件损坏扫描(铁律: 每次写入/修订后必跑)。检测 formfeed(\x0c) / 真 tab / 行中 CR——JSON 序列化转义被解释的隐性损坏。
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | 文件或技能目录列表(目录按 SKILL.md/references/templates/scripts 展开扫描) | |
| skills_root | No | 传入相对路径时的根目录 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses the concrete corruption targets being checked, which is real context beyond the schema, but it never states that the scan is read-only/non-mutating or what happens on findings—important traits for a tool an agent is told to run after every write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the purpose is front-loaded and the mandatory-run rule plus the detection list follow immediately. Efficient, though the parenthetical iron-rule framing is slightly decorative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a scan tool with a full parameter schema and an output schema, the description covers what matters: what it does, the checks it performs, and when to invoke it. It only lacks the read-only/behavior-on-findings note, a minor gap given the existing structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both `paths` (directory expansion behavior) and `skills_root` (base for relative paths) are already documented in the schema. The description adds no parameter detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (skill file corruption scan) and enumerates exactly which corruption classes it detects (formfeed, real tab, mid-line CR from JSON escape interpretation). It is clearly distinguishable from skill_backup/skill_log_decision, though it never contrasts itself with the closest sibling, skill_validate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit, strong trigger: 'iron rule — must run after every write/revision'. That tells the agent exactly when to invoke it. It stops short of stating when NOT to use it or how it relates to skill_validate, so no exclusions or alternatives are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_validateSkill ValidateC
技能体检: frontmatter 必填(name/description/version)+ 损坏扫描 + 孤儿 reference(未挂链=不可发现)+ decision-log 概要。
| Name | Required | Description | Default |
|---|---|---|---|
| skill_path | Yes | 技能目录绝对路径 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It lists what is checked but never states whether the operation is read-only, whether it mutates or repairs anything, or what side effects occur. A diagnostic tool implies no mutation, but nothing in the text confirms it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the purpose and packs four distinct checks via separators, with zero filler. It is terse but every clause earns its place by naming a concrete check.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the checks are enumerated. What is missing is routing against the overlapping siblings, particularly skill_scan_corruption and the two decision-log tools, which the description neither differentiates from nor excludes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage ('技能目录绝对路径' / absolute path to the skill directory), so the schema already documents it fully. The description adds no syntax, format, or constraint detail beyond the schema, which is the baseline expectation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb (体检/validate) and enumerates the exact checks performed: required frontmatter fields, corruption scan, orphan references, and a decision-log summary. This tells the agent precisely what the tool inspects. However, it does not distinguish itself from skill_scan_corruption, which appears to cover one of the same checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance at all. With four siblings — skill_backup, skill_scan_corruption, skill_log_decision, skill_read_decisions — the agent must guess whether to call this aggregate validator or the narrower per-concern tools, especially since corruption scanning overlaps with skill_scan_corruption.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
skill_backup - First observed
skill_log_decision - First observed
skill_read_decisions - First observed
skill_scan_corruption - First observed
skill_validate
TDQS
Scored across 5 tools
Each tool targets a distinct maintenance action: backup, corruption scan, decision logging, decision reading, and aggregate validation. However, skill_validate explicitly includes corruption scanning and a decision-log summary, so an agent may wonder when to call it versus skill_scan_corruption or skill_read_decisions.
All names use a consistent skill_ prefix and snake_case, which is predictable. Minor deviations exist: skill_backup and skill_validate are verb-only, while others use verb_noun, and log_decision/read_decisions differ in number.
Five tools is well-scoped for skill maintenance. Each tool covers a specific, non-redundant part of the backup-scan-validate-log workflow.
The core maintenance loop is present, but backup lacks a corresponding restore or list-backups operation, and decision-log reading only returns a summary. These are notable gaps for a server explicitly concerned with safe skill revision.
Maintenance
Related MCP Connectors
Manage portable AI agent playbooks, Agent Skills, MCP configurations, personas, and memory.
Scans remote MCP servers for protocol, security, and TLS issues; exposes scan tools via MCP.
The governed runtime for agent skills. Search the catalog and inspect a skill before running it.
Governed AI agent skills — one library, distributed to devs and exposed to remote agents over MCP.
Related MCP Servers
- AlicenseAqualityCmaintenanceConverts AI Skills (following Claude Skills format) into MCP server resources, enabling LLM applications to discover, access, and utilize self-contained skill directories through the Model Context Protocol. Provides tools to list available skills, retrieve skill details and content, and read supporting files with security protections.328Apache 2.0
- FlicenseNot gradedqualityCmaintenanceEnables users to validate MCP servers, skills, extensions, and packages for schema, security, functional, and semantic quality directly from their MCP client.-
- AlicenseNot gradedqualityCmaintenanceProvides read-only repository health scanning tools for drift detection, module reachability, prompt bloat, evidence calibration, and registration completeness, enabling agents to diagnose repositories via MCP.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables managing a canonical library of agent skills and MCP servers, syncing them across multiple development harnesses, and adding, importing, or configuring them through MCP tools.13 npmMIT