mcp-skill-server
Most coding assistants now support skills natively, so an MCP server just for skill discovery isn't necessary. Where this package adds value is making skills' execution deterministic and deployable — with a fixed entry point and controlled execution, skills developed in your editor can run in non-sandboxed production environments. It also supports incremental loading, so agents discover skills on demand instead of loading everything upfront.
MCP Skill Server
Build agent skills where you work. Write a Python script, add a SKILL.md, and your agent can use it immediately. Iterate in real-time as part of your daily workflow. When it's ready, deploy the same skill to production — no rewrite needed.
Why?
Most skill development looks like this: write code → deploy → test in a staging agent → realize it's wrong → redeploy → repeat. It's slow, and you never get to actually use the skill while building it.
MCP Skill Server flips this. It runs on your machine, inside your editor — Claude Code, Cursor, or Claude Desktop. You develop a skill and use it in your real work at the same time. That tight feedback loop (edit → save → use) means you discover what's missing naturally, not through artificial test scenarios. The premise is if the skill doesn't work well with Claude Code, it's unlikely to work with a less sophisticated agent.
How skills mature to survive in the outside world
Claude skills can already have companion scripts, but there's no formalized entry point — the agent decides how to invoke them. That works for local use, but it's not deployable: a production MCP server can't reliably call a skill if the execution path isn't fixed.
MCP Skill Server enforces a declared entry field in your SKILL.md frontmatter (e.g. entry: uv run python my_script.py). This gives you a single, fixed entry point that the server controls. Commands and parameters are discovered from the script's --help output — that's the source of truth, not the LLM's interpretation of your code.
1. Claude/coding agent skill → SKILL.md + scripts, but no fixed entry — agent decides how to run them
2. Local MCP skill (+ entry) → Fixed entry point, schema from --help, usable daily via this server
3. Production → Same skill, same entry — deployed to your enterprise MCP serverSharpen locally, then harden for production
Every agent that connects to the MCP server gets the same interface — list_skills, get_skill, run_skill — so the skill's description, parameter names, and help text are identical regardless of which agent calls them. That said, different agents have different strengths — a skill that works locally still needs testing with your production agent.
Use it yourself — build the skill, use it daily via Claude Code or Cursor. Fix descriptions and param names when the agent misuses the skill.
Test with a weaker model — try a smaller model to surface interface ambiguity.
Add a deterministic entry point — declare
entryin SKILL.md for reliable, secure execution. Useskill initto scaffold it,skill validateto check readiness.Test with your production agent — verify end-to-end in your target environment, then deploy.
Related MCP server: Skills MCP
Install
Claude Desktop (one-click)
After installing, edit the skills path in your Claude Desktop config to point to your skills directory.
Claude Code
claude mcp add skills -- uvx mcp-skill-server serve /path/to/my/skillsCursor
Add to .cursor/mcp.json in your project (or Settings → MCP → Add Server):
{
"mcpServers": {
"skills": {
"command": "uvx",
"args": ["mcp-skill-server", "serve", "/path/to/my/skills"]
}
}
}Manual install
# From PyPI (recommended)
uv pip install mcp-skill-server
# Or from source
git clone https://github.com/jcc-ne/mcp-skill-server
cd mcp-skill-server && uv sync
# Run the server
uvx mcp-skill-server serve /path/to/my/skillsThen add to your editor's MCP config:
{
"mcpServers": {
"skills": {
"command": "uvx",
"args": ["mcp-skill-server", "serve", "/path/to/my/skills"]
}
}
}Creating a Skill
Option A: Use skill init (recommended)
# Create a new skill
uv run mcp-skill-server init ./my_skills/hello -n "hello" -d "A friendly greeting"
# Or use the standalone command
uv run mcp-skill-init ./my_skills/hello -n "hello" -d "A friendly greeting"
# Promote an existing prompt-only Claude skill to a runnable MCP skill
uv run mcp-skill-init ./existing_claude_skillOption B: Manual setup
1. Create a folder with your script
my_skills/
└── hello/
├── SKILL.md
└── hello.py2. Add SKILL.md with frontmatter
---
name: hello
description: A friendly greeting skill
entry: uv run python hello.py
---
# Hello Skill
Greets the user by name.3. Write your script with argparse
# hello.py
import argparse
parser = argparse.ArgumentParser(description="Greeting skill")
parser.add_argument("--name", default="World", help="Name to greet")
args = parser.parse_args()
print(f"Hello, {args.name}!")That's it. The server auto-discovers commands and parameters from your --help output — no config needed.
Validating for Deployment
When a skill is ready to graduate to production:
uv run mcp-skill-server validate ./my_skills/hello
# or
uv run mcp-skill-validate ./my_skills/helloChecks:
Required frontmatter fields (name, description, entry)
Entry command uses allowed runtime
Script file exists
Commands discoverable via
--help
How It Works
MCP Tools
The server exposes four tools to your agent:
Tool | Description |
| List all available skills |
| Get details about a skill (commands, parameters) |
| Execute a skill with parameters |
| Reload skills after you make changes |
Schema Discovery
The server automatically discovers your skill's interface by parsing --help output:
# Subcommands become separate commands
subparsers = parser.add_subparsers(dest='command')
analyze = subparsers.add_parser('analyze', help='Run analysis')
# Arguments become parameters with inferred types
analyze.add_argument('--year', type=int, required=True) # int, required
analyze.add_argument('--file', type=str) # string, optionalOutput Files
Files saved to output/ are automatically detected. Alternatively, print OUTPUT_FILE:/path/to/file to stdout.
Plugins
Output Handlers
Process files generated by skills (upload, copy, transform, etc.):
from mcp_skill_server.plugins import OutputHandler, LocalOutputHandler
# Default: tracks local file paths
handler = LocalOutputHandler()
# Optional GCS handler (requires `uv sync --extra gcs`)
from mcp_skill_server.plugins import GCSOutputHandler
handler = GCSOutputHandler(
bucket_name="my-bucket",
folder_prefix="skills/outputs/",
)Response Formatters
Customize how execution results are formatted in MCP tool responses:
from mcp_skill_server.plugins import ResponseFormatter
class CustomFormatter(ResponseFormatter):
def format_execution_result(self, result, skill, command):
return f"Result: {result.stdout}"
# Use with create_server()
from mcp_skill_server import create_server
server = create_server(
"/path/to/skills",
response_formatter=CustomFormatter()
)Development
git clone https://github.com/jcc-ne/mcp-skill-server
cd mcp-skill-server
uv sync --dev
uv run pytest
uv run mcp-skill-server serve examples/Further Reading
Tool Design for LLMs — Why skills use a list/get/run pattern instead of exposing raw tools, and how it affects LLM accuracy
License
MIT
Available Tools
4 toolsget_skillA
Get skill details including commands and parameters
| Name | Required | Description | Default |
|---|---|---|---|
| skill_name | Yes | Name of the skill to get details for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, whether it requires specific permissions, what happens if the skill does not exist, or any side effects. The description only states the action without behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It states the action and the key content returned, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieval tool with one well-defined parameter and no output schema, the description adequately covers the purpose and the nature of the returned data (commands and parameters). It lacks error handling or format details, but those are less critical for a basic get operation, and the tool's simplicity keeps it reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single parameter 'skill_name', and the schema already explains its purpose. The tool description adds no additional parameter meaning beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') with a clear resource ('skill details') and adds detail ('including commands and parameters'). It clearly distinguishes from siblings like list_skills (which lists all), run_skill (which executes), and refresh_skills (which refreshes). An agent can immediately tell this is a single-item retrieval tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for fetching details of one skill by name, but does not explicitly state when to prefer it over list_skills or mention any alternatives. The usage context is implied from the sibling names, but no direct guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_skillsB
List all available skills
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears full responsibility for behavioral disclosure. It only states 'List all available skills' without mentioning return format, pagination, ordering, permissions, or any caveats about what 'available' means. For a read-only enumeration this is acceptable but sparse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence with no filler. It front-loads the verb and resource, making the purpose instantly scannable. This is appropriately concise for a parameterless list operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description is minimally sufficient. It tells the agent that it lists all skills, which is likely enough to invoke correctly. However, it omits any context about the nature of 'skills' or potential limitations (e.g., only active ones), which could matter in edge cases. Overall adequate but not rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema is empty (100% coverage by default). The description adds no parameter-specific semantics because none exist, so the baseline of 4 applies. No further detail is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and the resource ('all available skills'). It implies a broad enumeration, which distinguishes it from get_skill (retrieve a specific one) and run_skill (execute). However, it doesn't explicitly contrast with siblings, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. While it is intuitive that listing all skills is the purpose, the description doesn't mention when to prefer it over get_skill or refresh_skills. A sentence like 'Use this to enumerate skills, as opposed to fetching a single skill or refreshing the list' would elevate this.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_skillsA
Refresh the skill list (use after adding new skills)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. 'Refresh' is vague about side effects: whether it reloads from disk, invalidates a cache, mutates stored data, or is idempotent. The description focuses on the trigger rather than what actually happens during refresh.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence delivers the action and usage context with no filler. The verb is front-loaded and the guidance is compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema refresh utility, the description covers the main decision point: use it after adding skills. It could be slightly more complete by noting whether the refresh affects only the cached list or also individual skill details, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so there is no parameter semantics burden. The description appropriately adds the only actionable context needed: when to invoke the refresh.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('Refresh the skill list') and adds a concrete trigger ('use after adding new skills'). It is distinguishable from sibling tools like list_skills or run_skill, though it doesn't explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'use after adding new skills' provides explicit timing guidance for when to call this tool. It does not discuss exclusions or alternatives, but the when-to-use context is clear enough for a zero-parameter utility.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_skillC
Execute a skill command with parameters
| Name | Required | Description | Default |
|---|---|---|---|
| command | No | Command to execute (use 'default' for single-command skills) | default |
| parameters | No | Parameters to pass to the skill command | |
| skill_name | Yes | Name of the skill to run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits, but it only says 'Execute a skill command with parameters'. It does not indicate side effects, required permissions, reversibility, or what happens on execution. For a tool that likely has side effects (running a command), this is seriously under-disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no wasted words, but it is so minimal that it borders on under-specification rather than effective conciseness. It lacks the structure to front-load key information; it simply states the obvious.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has three parameters, one nested object, and no annotations or output schema, the description is far from complete. It does not explain what a 'skill command' is, how parameters are used, or what to expect after execution. An agent would struggle to call it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds no extra meaning beyond the word 'parameters', which is trivial. It does not clarify how parameters relate to the command or any syntax, so it stays at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the verb 'Execute' and the resource 'skill command with parameters', which clearly conveys the action. It is not a tautology and gives a specific purpose, though it does not differentiate from sibling tools like list_skills or get_skill, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the siblings (list_skills, get_skill, refresh_skills). No context is provided about typical scenarios or exclusions, leaving the agent to infer usage on its own. This is a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.2- First observed
get_skill - First observed
list_skills - First observed
refresh_skills - First observed
run_skill
TDQS
Scored across 4 tools
Each tool targets a distinct operation: listing, retrieving details, executing, and refreshing. There is no overlap or ambiguity between them.
All tool names follow a consistent verb_noun pattern (list_skills, get_skill, run_skill, refresh_skills), making the API predictable.
Four tools is a tight, well-scoped set for a skill management server, with each tool covering a clear and necessary operation.
The surface covers the full expected lifecycle: discover skills, inspect skill details, execute skills, and refresh the catalog. No obvious gaps for this domain.
Maintenance
Related MCP Connectors
- SkilderOAuthai.skilder
One place to build, share, and govern the skills and tools your AI agents use at work.
Agent-first skill marketplace with USK open standard for Claude, Cursor, Gemini, Codex CLI.
Git-backed platform for skills, tools, and context for AI agents
Build, deploy, and sell AI agents for local-service businesses - from your IDE.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables Claude to create, edit, run, and manage reusable skills stored locally, including executing scripts with automatic dependency management and environment variables. Works across all MCP-compatible clients like Cursor, Claude Desktop, and claude.ai.530MIT
- FlicenseNot gradedqualityDmaintenanceTransform any AI agent into a domain expert by giving it access to modular, reusable skills through the Model Context Protocol. Brings Claude's Skills format to any MCP-compatible agent, allowing you to create skills once and use them everywhere.47 npm29-
- AlicenseBqualityAmaintenanceTransform 17 source types into AI-ready skills and RAG knowledge, directly from Claude Code.404,375 PyPI14,963MIT
- AlicenseAqualityCmaintenanceTurn any YouTube video, article, PDF, or image into a reusable Claude Code skill — without leaving your editor.633MIT