Skip to main content
Glama

mcp-minions

An MCP server that hands narrow, well-defined jobs to a small local model instead of a big one: classify text, extract fields into a JSON Schema, and summarize. The model runs on your own machine through Ollama or LM Studio, so the text never leaves it.

The idea is "strong system, slim model": a small model is fine at reading and rewriting, so the server gives it one clear task, forces structured output, and validates the result in code before returning it. It does not trust the model to follow the format.

Tools

Tool

What it does

classify(text, labels, hint="")

Picks exactly one of the given labels (1 to 50, unique). Returns {"label": ...}.

extract(text, schema, hint="")

Fills an object matching your JSON Schema ("type": "object"). Returns the object.

summarize(text, max_words=100)

Summary in at most max_words words (5 to 1000). Returns {"summary": ...}.

classify_file, extract_file, summarize_file

Same, but reading UTF-8 text from a local file path.

Related MCP server: ollama-handoff

How it behaves

  • Schema first, then verify. Each call asks the backend for schema-constrained JSON, and then checks the reply with jsonschema. Whether a backend actually enforces the schema varies, so a reply that is not valid JSON or does not match the schema is an error, not a guess.

  • No silent fallback. If the backend is down you get MINION_DOWN. Nothing is retried on another model.

  • Untrusted input. Your text is wrapped in <data> tags, a closing </data> inside it is neutralised, and the system prompt tells the model never to follow instructions found in the data. This reduces prompt-injection risk. It is not a guarantee, which is why the output schema is checked in code.

  • Stable error codes: MINION_DOWN, MINION_BACKEND_ERROR, MINION_SCHEMA, MINION_BAD_INPUT, MINION_BAD_CONFIG, and (v0.2) MINION_NO_QUORUM.

Install and run

Requires Python 3.10+ and uv. You also need Ollama or LM Studio running with a model loaded.

git clone https://github.com/MartinTheGuitarMan/mcp-minions
cd mcp-minions
uv sync
MINION_BACKEND=ollama MINION_MODEL=<your-model> uv run mcp-minions

Example MCP client entry (the exact file and format depend on your client):

{
  "mcpServers": {
    "minions": {
      "command": "uv",
      "args": ["--directory", "/path/to/mcp-minions", "run", "mcp-minions"],
      "env": { "MINION_BACKEND": "ollama", "MINION_MODEL": "<your-model>" }
    }
  }
}

Configuration (environment variables)

Variable

Default

Meaning

MINION_MODEL

required

Model name as the backend knows it

MINION_BACKEND

lmstudio

lmstudio or ollama

MINION_BASE_URL

http://127.0.0.1:1234/v1 (LM Studio), http://127.0.0.1:11434/v1 (Ollama)

OpenAI-compatible root, ending in /v1

MINION_TIMEOUT

120

Seconds to wait for the model

MINION_MAX_FILE_BYTES

1000000

Size cap for the *_file tools

MINION_OLLAMA_NATIVE

unset

Set to 1 to use Ollama's native /api/chat format instead of response_format

Roster, quorum and judge (v0.2)

Set MINION_ROSTER to inline JSON or a path to a JSON file to use several models. Without it, the single-model setup above applies unchanged.

{
  "models": {
    "g1":  {"backend": "lmstudio", "model": "google/gemma-3-1b",   "kind": "fast"},
    "e4b": {"backend": "lmstudio", "model": "google/gemma-4-e4b",  "kind": "fast"},
    "q9":  {"backend": "ollama",   "model": "qwen3.5:9b", "kind": "reasoning", "max_tokens": 4096}
  },
  "routes": {
    "classify":  {"fast": ["g1", "e4b"], "quorum": 2, "judge": "q9"},
    "extract":   {"fast": ["e4b"], "judge": "q9"},
    "summarize": {"fast": ["g1"]}
  }
}
  • kind: fast models use constrained decoding (schema requested, then validated by us). kind: reasoning models are called without constrained decoding: they think freely, the think block is stripped, the last JSON object is extracted and validated against the schema by us, with one retry that feeds the error back.

  • Quorum (classify only). Every fast model in the route votes; a label wins with at least quorum votes (default: majority) and a strict lead.

  • Judge. Runs only on disagreement or schema failure, never otherwise. It verifies the proposals (confirm one, or correct it) instead of redoing the task; on a schema failure it repairs the fast model's invalid output. With no judge configured, disagreement raises MINION_NO_QUORUM and a schema failure raises MINION_SCHEMA.

  • Hard failures are never routed around. MINION_DOWN and backend errors from any model propagate; the judge is not used to mask them.

  • Set MINION_META=1 to add a _minion key (votes, judge used) to results.

Security note

The *_file tools read any UTF-8 text file the caller names (up to the size cap). Whoever can call this server can therefore read files your user account can read. Only connect clients you trust, and leave the file tools out of any setup where that is a problem.

Tests

uv run pytest

The unit tests use a fake backend. To try a real model, scripts/smoke.py runs every tool against a live backend and repeatedly probes whether the schema constraint holds, including hostile and prose-bait inputs:

uv run python scripts/smoke.py ollama <your-model> --repeat 5

Status

Version 0.2. Early and small; expect changes.

License

MIT, see LICENSE.

Available Tools

6 tools
classifyC

Classify text into exactly one of the given labels using a local model.

ParametersJSON Schema
NameRequiredDescriptionDefault
hintNo
textYes
labelsYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It does disclose two useful traits — processing is done by a 'local model' (no external API) and output is constrained to a single label — but says nothing about failure modes when text matches no label, determinism, or latency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the constraint front-loaded and no filler. It is appropriately sized, though the brevity is partly the source of the coverage gaps rather than pure economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter description coverage, the description should carry much more. It omits what the return value looks like (label only, label plus score?), how 'hint' behaves, and any error behavior — real gaps for a 3-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clarifies that 'labels' is a candidate set from which exactly one is chosen, but the 'hint' parameter is never mentioned anywhere, leaving a required piece of the API undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (classify) and resource (text) plus the key constraint 'exactly one of the given labels'. The word 'text' implicitly separates it from the sibling classify_file, though it never names that alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this over classify_file, extract, or summarize. The only routing signal is the implicit 'text' vs file distinction, which the agent must infer from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classify_fileB

Like classify, reading UTF-8 text from a local file (size-capped).

ParametersJSON Schema
NameRequiredDescriptionDefault
hintNo
pathYes
labelsYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose two real traits: input must be UTF-8 text and the read is size-capped. It does not state the actual size limit, what happens on non-UTF-8 or oversized files, or any error/permission behavior, leaving meaningful gaps for a file-reading operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single tight sentence with no waste and the key constraint (file input) is front-loaded. Brevity here reflects under-specification rather than disciplined conciseness, since essential details are simply absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with 0% schema coverage, no annotations, and no output schema, the description is too thin. It omits parameter meaning, the size cap value, error behavior, and how results relate to the sibling 'classify' output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, so the description must compensate, and it largely does not. It hints at 'path' (local file) but says nothing about 'labels' (required) or 'hint', leaving the agent to guess at their meaning from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose by anchoring to the sibling 'classify' and adding the file-based variant: it classifies UTF-8 text read from a local file. This distinguishes it from the plain 'classify' tool, though the verb itself is only implied via the name rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Saying 'Like classify, reading UTF-8 text from a local file' implicitly tells the agent to use this when input is a file rather than inline text. However, it gives no explicit when-to-use/when-not guidance and never mentions the other file siblings (extract_file, summarize_file) as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extractC

Extract fields from text into an object matching a JSON Schema (type: object).

ParametersJSON Schema
NameRequiredDescriptionDefault
hintNo
textYes
schemaYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no indication of whether extraction is model-driven, how it behaves on non-conforming or missing fields, whether it errors or returns partial results, or any cost/latency characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the output contract is stated up front. It is efficient, though the brevity comes at the cost of the missing detail noted in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, 0% schema coverage, and a nested schema parameter, the definition is too thin. It should at minimum explain the 'hint' parameter, the expected JSON Schema shape, and failure behavior when extraction cannot satisfy the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters. The description echoes 'text' and 'schema' and adds one useful constraint (the schema must be type: object), but the third parameter 'hint' is entirely undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (extract) and resource (fields from text) plus the output contract (an object matching a JSON Schema). The phrase 'from text' implicitly separates it from extract_file, but it never names or contrasts with the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as extract_file for file inputs or classify/summarize for other tasks. The only routing signal is the incidental 'from text' phrasing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_fileB

Like extract, reading UTF-8 text from a local file (size-capped).

ParametersJSON Schema
NameRequiredDescriptionDefault
hintNo
pathYes
schemaYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two real traits: input must be UTF-8 decodable text and file size is capped. It says nothing about failure behavior on non-UTF-8 or oversized files, filesystem permission requirements, or what the extraction produces.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the source constraint front-loaded and no filler. It is efficient, though the brevity borders on under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and 0% parameter coverage for a tool with a nested free-form 'schema' argument. The description should be compensating for that structural silence but instead offers only one clause, leaving an agent without enough to invoke it confidently beyond the happy path.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters (path, schema, hint), one of which is a nested free-form object, and the description adds no parameter meaning at all. An agent gets no guidance on what 'hint' does or what shape the 'schema' object should take, which is the main ambiguity for this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the operation (extract) and the source (a local file), and its phrasing 'Like extract ... from a local file' positions it as the file-input counterpart of the sibling extract tool. It is clearly distinguishable from classify_file/summarize_file by the verb, though it never states what is extracted (e.g. structured data per the schema).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Like extract, reading UTF-8 text from a local file' implies usage (choose this when the input is a file rather than inline text), but there is no explicit when-to-use statement, no when-not, and no named alternative. Usage is left to inference from the sibling reference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarizeC

Summarize text in at most max_words words.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
max_wordsNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses almost nothing: no indication of input size limits, truncation behavior, whether the summary is extractive or abstractive, or what the return looks like. It is at least non-misleading, but it leaves the agent guessing on every behavioral question.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler or repetition. It is efficiently sized, though the brevity comes at the cost of substance rather than being tight around rich content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations, no output schema, and no parameter documentation, the description should carry more. It omits input constraints, output shape, and sibling differentiation, so an agent cannot call it confidently in ambiguous cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description only restates max_words as a word cap (its default of 100 is already in the schema) while saying nothing about the required text parameter. It adds little beyond what the schema title already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Summarize) and resource (text), and constrains the output length via max_words. It does not, however, distinguish itself from the sibling summarize_file, leaving the agent to infer that this one takes inline text rather than a file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus classify, extract, or summarize_file, nor any prerequisites or exclusions. The agent gets no routing guidance at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_fileB

Like summarize, reading UTF-8 text from a local file (size-capped).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
max_wordsNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses three behavioral constraints: input must be UTF-8, input must be a local file, and size is capped. It does not say what happens when the cap is exceeded, whether read permissions are required, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the file-based scope is stated first. It is efficient, though 'size-capped' is too vague to fully earn its clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage, the description should do more. It omits the max_words parameter, the size cap value, error behavior, and output shape, leaving an agent unable to predict results or failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 2 parameters. The description only obliquely gestures at 'path' (local file) and says nothing at all about 'max_words' (which defaults to 100), so the parameter that controls output length is entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (summarize) and resource (local file), and explicitly anchors itself to the sibling 'summarize' by stating it works 'like' it but reads from a file. An agent can distinguish it from summarize, classify_file, and extract_file without opening schemas, though the distinction is by analogy rather than a direct statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Like summarize, reading UTF-8 text from a local file' implies the selection rule: use this when the source is a file rather than inline text. However, it never states when not to use it, nor what to do for non-UTF-8, non-text, or oversized files.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.2.0
    • First observedclassify
    • First observedclassify_file
    • First observedextract
    • First observedextract_file
    • First observedsummarize
    • First observedsummarize_file

TDQS

B3.4/5.0

Scored across 6 tools

Disambiguation5/5

classify, extract, and summarize target three clearly distinct operations (classification, schema-based field extraction, summarization), and each _file variant is explicitly documented as the same operation with a file input source. There is no real ambiguity about which tool performs which action.

Naming Consistency5/5

All names use a consistent snake_case verb pattern, with the _file suffix applied uniformly to denote the file-reading variant of each operation. The taxonomy is predictable and readable.

Tool Count5/5

Six tools for three core text operations plus their file counterparts is tight and well-scoped; every tool earns its place with no redundancy beyond the intended input-source pairing.

Completeness4/5

The surface covers the three primary operations for both raw text and local files, forming a coherent lifecycle. Minor gaps exist (e.g. no batch, URL, or stdin input variants), but agents can work around them by reading content themselves.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude to delegate coding tasks to local Ollama models, reducing API token usage by up to 98.75% while leveraging local compute resources. Supports code generation, review, refactoring, and file analysis with Claude providing oversight and quality assurance.
    383 npm
    25
    AGPL 3.0
  • A
    license
    A
    quality
    C
    maintenance
    Enables Claude Code to offload routine code generation and text processing tasks to a local Ollama LLM, saving Cloud API tokens and costs with automatic model selection and security features.
    11
    124 npm
    4
    Apache 2.0