mcp-minions
Allows using Ollama-hosted local models as the backend for classification, schema-based extraction, and summarization, with structured JSON output validated by the server.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-minionsClassify this support ticket as billing, technical, or other: 'I can't log in.'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-minions
An MCP server that hands narrow, well-defined jobs to a small local model instead of a big one: classify text, extract fields into a JSON Schema, and summarize. The model runs on your own machine through Ollama or LM Studio, so the text never leaves it.
The idea is "strong system, slim model": a small model is fine at reading and rewriting, so the server gives it one clear task, forces structured output, and validates the result in code before returning it. It does not trust the model to follow the format.
Tools
Tool | What it does |
| Picks exactly one of the given labels (1 to 50, unique). Returns |
| Fills an object matching your JSON Schema ( |
| Summary in at most |
| Same, but reading UTF-8 text from a local file path. |
Related MCP server: ollama-handoff
How it behaves
Schema first, then verify. Each call asks the backend for schema-constrained JSON, and then checks the reply with
jsonschema. Whether a backend actually enforces the schema varies, so a reply that is not valid JSON or does not match the schema is an error, not a guess.No silent fallback. If the backend is down you get
MINION_DOWN. Nothing is retried on another model.Untrusted input. Your text is wrapped in
<data>tags, a closing</data>inside it is neutralised, and the system prompt tells the model never to follow instructions found in the data. This reduces prompt-injection risk. It is not a guarantee, which is why the output schema is checked in code.Stable error codes:
MINION_DOWN,MINION_BACKEND_ERROR,MINION_SCHEMA,MINION_BAD_INPUT,MINION_BAD_CONFIG, and (v0.2)MINION_NO_QUORUM.
Install and run
Requires Python 3.10+ and uv. You also need Ollama or LM Studio running with a model loaded.
git clone https://github.com/MartinTheGuitarMan/mcp-minions
cd mcp-minions
uv sync
MINION_BACKEND=ollama MINION_MODEL=<your-model> uv run mcp-minionsExample MCP client entry (the exact file and format depend on your client):
{
"mcpServers": {
"minions": {
"command": "uv",
"args": ["--directory", "/path/to/mcp-minions", "run", "mcp-minions"],
"env": { "MINION_BACKEND": "ollama", "MINION_MODEL": "<your-model>" }
}
}
}Configuration (environment variables)
Variable | Default | Meaning |
| required | Model name as the backend knows it |
|
|
|
|
| OpenAI-compatible root, ending in |
|
| Seconds to wait for the model |
|
| Size cap for the |
| unset | Set to |
Roster, quorum and judge (v0.2)
Set MINION_ROSTER to inline JSON or a path to a JSON file to use several models. Without it, the single-model setup above applies unchanged.
{
"models": {
"g1": {"backend": "lmstudio", "model": "google/gemma-3-1b", "kind": "fast"},
"e4b": {"backend": "lmstudio", "model": "google/gemma-4-e4b", "kind": "fast"},
"q9": {"backend": "ollama", "model": "qwen3.5:9b", "kind": "reasoning", "max_tokens": 4096}
},
"routes": {
"classify": {"fast": ["g1", "e4b"], "quorum": 2, "judge": "q9"},
"extract": {"fast": ["e4b"], "judge": "q9"},
"summarize": {"fast": ["g1"]}
}
}kind: fastmodels use constrained decoding (schema requested, then validated by us).kind: reasoningmodels are called without constrained decoding: they think freely, the think block is stripped, the last JSON object is extracted and validated against the schema by us, with one retry that feeds the error back.Quorum (classify only). Every fast model in the route votes; a label wins with at least
quorumvotes (default: majority) and a strict lead.Judge. Runs only on disagreement or schema failure, never otherwise. It verifies the proposals (
confirmone, orcorrectit) instead of redoing the task; on a schema failure it repairs the fast model's invalid output. With no judge configured, disagreement raisesMINION_NO_QUORUMand a schema failure raisesMINION_SCHEMA.Hard failures are never routed around.
MINION_DOWNand backend errors from any model propagate; the judge is not used to mask them.Set
MINION_META=1to add a_minionkey (votes, judge used) to results.
Security note
The *_file tools read any UTF-8 text file the caller names (up to the size cap). Whoever can call this server can therefore read files your user account can read. Only connect clients you trust, and leave the file tools out of any setup where that is a problem.
Tests
uv run pytestThe unit tests use a fake backend. To try a real model, scripts/smoke.py runs every tool against a live backend and repeatedly probes whether the schema constraint holds, including hostile and prose-bait inputs:
uv run python scripts/smoke.py ollama <your-model> --repeat 5Status
Version 0.2. Early and small; expect changes.
License
MIT, see LICENSE.
Available Tools
6 toolsclassifyC
Classify text into exactly one of the given labels using a local model.
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| text | Yes | ||
| labels | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose two useful traits — processing is done by a 'local model' (no external API) and output is constrained to a single label — but says nothing about failure modes when text matches no label, determinism, or latency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the constraint front-loaded and no filler. It is appropriately sized, though the brevity is partly the source of the coverage gaps rather than pure economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter description coverage, the description should carry much more. It omits what the return value looks like (label only, label plus score?), how 'hint' behaves, and any error behavior — real gaps for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies that 'labels' is a candidate set from which exactly one is chosen, but the 'hint' parameter is never mentioned anywhere, leaving a required piece of the API undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (classify) and resource (text) plus the key constraint 'exactly one of the given labels'. The word 'text' implicitly separates it from the sibling classify_file, though it never names that alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this over classify_file, extract, or summarize. The only routing signal is the implicit 'text' vs file distinction, which the agent must infer from sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classify_fileB
Like classify, reading UTF-8 text from a local file (size-capped).
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| path | Yes | ||
| labels | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose two real traits: input must be UTF-8 text and the read is size-capped. It does not state the actual size limit, what happens on non-UTF-8 or oversized files, or any error/permission behavior, leaving meaningful gaps for a file-reading operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single tight sentence with no waste and the key constraint (file input) is front-loaded. Brevity here reflects under-specification rather than disciplined conciseness, since essential details are simply absent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with 0% schema coverage, no annotations, and no output schema, the description is too thin. It omits parameter meaning, the size cap value, error behavior, and how results relate to the sibling 'classify' output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the description must compensate, and it largely does not. It hints at 'path' (local file) but says nothing about 'labels' (required) or 'hint', leaving the agent to guess at their meaning from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose by anchoring to the sibling 'classify' and adding the file-based variant: it classifies UTF-8 text read from a local file. This distinguishes it from the plain 'classify' tool, though the verb itself is only implied via the name rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Saying 'Like classify, reading UTF-8 text from a local file' implicitly tells the agent to use this when input is a file rather than inline text. However, it gives no explicit when-to-use/when-not guidance and never mentions the other file siblings (extract_file, summarize_file) as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extractC
Extract fields from text into an object matching a JSON Schema (type: object).
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| text | Yes | ||
| schema | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no indication of whether extraction is model-driven, how it behaves on non-conforming or missing fields, whether it errors or returns partial results, or any cost/latency characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the output contract is stated up front. It is efficient, though the brevity comes at the cost of the missing detail noted in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, 0% schema coverage, and a nested schema parameter, the definition is too thin. It should at minimum explain the 'hint' parameter, the expected JSON Schema shape, and failure behavior when extraction cannot satisfy the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters. The description echoes 'text' and 'schema' and adds one useful constraint (the schema must be type: object), but the third parameter 'hint' is entirely undocumented in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extract) and resource (fields from text) plus the output contract (an object matching a JSON Schema). The phrase 'from text' implicitly separates it from extract_file, but it never names or contrasts with the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as extract_file for file inputs or classify/summarize for other tasks. The only routing signal is the incidental 'from text' phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_fileB
Like extract, reading UTF-8 text from a local file (size-capped).
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| path | Yes | ||
| schema | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two real traits: input must be UTF-8 decodable text and file size is capped. It says nothing about failure behavior on non-UTF-8 or oversized files, filesystem permission requirements, or what the extraction produces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the source constraint front-loaded and no filler. It is efficient, though the brevity borders on under-specification rather than true conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and 0% parameter coverage for a tool with a nested free-form 'schema' argument. The description should be compensating for that structural silence but instead offers only one clause, leaving an agent without enough to invoke it confidently beyond the happy path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters (path, schema, hint), one of which is a nested free-form object, and the description adds no parameter meaning at all. An agent gets no guidance on what 'hint' does or what shape the 'schema' object should take, which is the main ambiguity for this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the operation (extract) and the source (a local file), and its phrasing 'Like extract ... from a local file' positions it as the file-input counterpart of the sibling extract tool. It is clearly distinguishable from classify_file/summarize_file by the verb, though it never states what is extracted (e.g. structured data per the schema).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Like extract, reading UTF-8 text from a local file' implies usage (choose this when the input is a file rather than inline text), but there is no explicit when-to-use statement, no when-not, and no named alternative. Usage is left to inference from the sibling reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarizeC
Summarize text in at most max_words words.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| max_words | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and discloses almost nothing: no indication of input size limits, truncation behavior, whether the summary is extractive or abstractive, or what the return looks like. It is at least non-misleading, but it leaves the agent guessing on every behavioral question.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler or repetition. It is efficiently sized, though the brevity comes at the cost of substance rather than being tight around rich content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations, no output schema, and no parameter documentation, the description should carry more. It omits input constraints, output shape, and sibling differentiation, so an agent cannot call it confidently in ambiguous cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only restates max_words as a word cap (its default of 100 is already in the schema) while saying nothing about the required text parameter. It adds little beyond what the schema title already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Summarize) and resource (text), and constrains the output length via max_words. It does not, however, distinguish itself from the sibling summarize_file, leaving the agent to infer that this one takes inline text rather than a file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus classify, extract, or summarize_file, nor any prerequisites or exclusions. The agent gets no routing guidance at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_fileB
Like summarize, reading UTF-8 text from a local file (size-capped).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| max_words | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses three behavioral constraints: input must be UTF-8, input must be a local file, and size is capped. It does not say what happens when the cap is exceeded, whether read permissions are required, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the file-based scope is stated first. It is efficient, though 'size-capped' is too vague to fully earn its clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter coverage, the description should do more. It omits the max_words parameter, the size cap value, error behavior, and output shape, leaving an agent unable to predict results or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 2 parameters. The description only obliquely gestures at 'path' (local file) and says nothing at all about 'max_words' (which defaults to 100), so the parameter that controls output length is entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (summarize) and resource (local file), and explicitly anchors itself to the sibling 'summarize' by stating it works 'like' it but reads from a file. An agent can distinguish it from summarize, classify_file, and extract_file without opening schemas, though the distinction is by analogy rather than a direct statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Like summarize, reading UTF-8 text from a local file' implies the selection rule: use this when the source is a file rather than inline text. However, it never states when not to use it, nor what to do for non-UTF-8, non-text, or oversized files.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.2.0- First observed
classify - First observed
classify_file - First observed
extract - First observed
extract_file - First observed
summarize - First observed
summarize_file
TDQS
Scored across 6 tools
classify, extract, and summarize target three clearly distinct operations (classification, schema-based field extraction, summarization), and each _file variant is explicitly documented as the same operation with a file input source. There is no real ambiguity about which tool performs which action.
All names use a consistent snake_case verb pattern, with the _file suffix applied uniformly to denote the file-reading variant of each operation. The taxonomy is predictable and readable.
Six tools for three core text operations plus their file counterparts is tight and well-scoped; every tool earns its place with no redundancy beyond the intended input-source pairing.
The surface covers the three primary operations for both raw text and local files, forming a coherent lifecycle. Minor gaps exist (e.g. no batch, URL, or stdin input variants), but agents can work around them by reading content themselves.
Maintenance
Related MCP Connectors
JSON/YAML, regex, diff, JWT, SQL dialects — the keyless millisecond ops an agent needs mid-task.
150+ vertical AI expert bots as agent tools. $1 bots run on YOUR machine - your data stays yours.
Turn messy text into strict JSON schemas agents can trust (invoice, receipt, contact, resume).
Deterministic JSON repair, validate, example-gen, schema-coerce for agents. Zero LLM, sub-10ms.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables Claude to delegate coding tasks to local Ollama models, reducing API token usage by up to 98.75% while leveraging local compute resources. Supports code generation, review, refactoring, and file analysis with Claude providing oversight and quality assurance.383 npm25AGPL 3.0
- AlicenseAqualityBmaintenanceOffloads cheap work from cloud LLM agents to a local Ollama model, reducing costs and keeping frontier models focused on complex tasks.833 PyPI3MIT
- AlicenseNot gradedqualityDmaintenanceEnables Claude Code to delegate mechanical tasks (summaries, boilerplate, reformatting) to local models running in LM Studio.1MIT
- AlicenseAqualityCmaintenanceEnables Claude Code to offload routine code generation and text processing tasks to a local Ollama LLM, saving Cloud API tokens and costs with automatic model selection and security features.11124 npm4Apache 2.0