ollama-handoff
This server lets an MCP client offload routine text tasks to a local Ollama model for summaries, code reviews, commit drafts, extraction, and chat without cloud API charges.
ask_local: Send a one-shot prompt with optional system instruction.
chat_local: Run a multi-turn conversation with explicit message history.
summarize_local: Summarize supplied text, optionally focused on a topic.
code_review_local: Get a first-pass review of code or a diff.
draft_commit_message_local: Generate a conventional-style commit message from a diff.
extract_local: Pull structured items like URLs, names, or error codes from text.
list_models: Discover installed Ollama models.
server_info: Inspect effective configuration (model, context size, etc.).
Works with configurable Ollama URL, default model, context window, keep-alive, and timeout.
Local execution avoids cloud model API charges, though results should be checked.
Enables AI agents to offload routine tasks to local Ollama models, reducing cloud costs and frontier model context usage.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ollama-handoffsummarize the errors in build.log"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ollama Handoff
Local summaries, extractions, code reviews, and commit drafts for your MCP client.
Ollama Handoff lets your agent send routine text tasks to a model running on your machine. Eight MCP tools provide focused prompts, model discovery, and configuration checks.
Local inference does not incur cloud model API charges. Your calling agent can still use paid tokens to plan tasks and read results. Speed, memory use, and quality depend on your model and hardware.
Quick start
1. Prepare Ollama
Install Ollama and uv. Python 3.11 or newer is required; uv can manage Python for you.
ollama pull llama3.1:8b
ollama listKeep Ollama running. If the desktop app or service is not already running, start ollama serve in another terminal.
These examples select llama3.1:8b. You can substitute another installed model. Without an override, the package defaults to qwen2.5-coder:14b, which must be downloaded separately.
2. Register the server
For Claude Code:
claude mcp add --transport stdio --env OLLAMA_DEFAULT_MODEL=llama3.1:8b ollama-handoff -- uvx ollama-handoff@0.1.3For a client that accepts an mcpServers JSON configuration:
{
"mcpServers": {
"ollama-handoff": {
"command": "uvx",
"args": ["ollama-handoff@0.1.3"],
"env": {
"OLLAMA_DEFAULT_MODEL": "llama3.1:8b"
}
}
}
}Add this entry using your client's MCP settings, then reconnect or restart it. If the client cannot find uvx, use the absolute path reported by where.exe uvx on Windows or command -v uvx on macOS and Linux.
Version 0.1.3 declares the MCP compatibility constraint automatically. This server uses the MCP 1 FastMCP API, which MCP 2 removed. If you remain on version 0.1.2, add --with "mcp<2" to the uvx command.
For pip, install into a virtual environment and configure your client to run that environment's ollama-handoff executable:
python -m pip install "ollama-handoff==0.1.3"3. Verify the connection
Ask your agent:
Call
server_infoandlist_modelsfrom Ollama Handoff. Confirm that the configured model is installed. Then callsummarize_localwith the following text, focusing on the failed test:10:00:01 INFO Starting build 10:00:02 INFO Compiled 12 modules 10:00:03 ERROR tests/test_checkout.py::test_total expected 42.00, got 40.00 10:00:03 ERROR Build stopped because one test failed
Check that the summary identifies test_total, the expected total of 42.00, and the actual total of 40.00. Wording varies by model. Tools accept text, not file paths: your agent must supply file contents when needed.
Related MCP server: local-llm-delegation-mcp
Run the demo without an agent
The demo launches the real MCP server over stdio, discovers its tools, checks configuration, lists installed models, and summarizes the synthetic log above. It requires local Ollama but no cloud API key.
git clone https://github.com/Michael-WhiteCapData/ollama-handoff.git
cd ollama-handoff
uv venv
uv pip install -e .
uv run --no-project python examples/demo.py --model llama3.1:8bA successful run discovers 8 tools, lists your model, and returns the summary. Each call prints its elapsed time. Verified on Windows with Python 3.14, MCP 1.30.0, and llama3.1:8b. Timing is not a benchmark; the first model load may take longer.
Tools
Tool | Use it for |
| A single prompt with an optional system instruction |
| A conversation with explicit message history |
| Summaries of supplied text, optionally focused on a topic |
| Initial review of supplied code or a diff |
| A commit message from a supplied diff |
| Extracting items such as URLs, names, or error codes |
| Discovering installed Ollama models |
| Inspecting effective server configuration |
Generated summaries and reviews need checking. The server does not read files, stage changes, or create commits for you.
Configuration
Set these variables in your MCP registration:
Variable | Default | Description |
|
| Ollama server URL |
|
| Model used when a tool call omits a model |
|
| Context window in tokens |
|
| Time to keep the model loaded |
|
| Generation and chat request timeout in seconds |
For a small task on a machine with limited memory, try OLLAMA_NUM_CTX=4096 and a smaller model.
The selected Ollama endpoint receives the text sent to these tools. A remote endpoint sends that text to another machine. Local execution does not prevent your calling client from sending prompts or results to its own cloud provider.
Troubleshooting
Symptom | What to check |
| Upgrade to version 0.1.3 and restart the MCP client |
| Restart the client after installing uv, or configure the absolute executable path |
Connection refused | Confirm Ollama is running and |
Model not found | Match the full name from |
Slow response or timeout | Allow for model loading; try a smaller model or context |
Server appears idle in a terminal | This stdio server waits for an MCP client; it is not an interactive chat CLI |
server_info checks configuration without contacting Ollama. list_models checks connectivity. A successful summarize_local call also confirms generation.
Docker
The included Dockerfile builds the source package. Keep stdin open for MCP:
docker build -t ollama-handoff .
docker run --rm -i -e OLLAMA_URL=http://host.docker.internal:11434 -e OLLAMA_DEFAULT_MODEL=llama3.1:8b ollama-handoffOn Linux without Docker Desktop, use --network=host with OLLAMA_URL=http://localhost:11434. Do not add -t when an MCP client launches the container.
Development
uv venv
uv pip install -e ".[dev]"
uv run --no-project ruff check .
uv run --no-project pytestUnit tests use httpx.MockTransport and do not need Ollama. The demo uses real inference. See CONTRIBUTING.md.
License
MIT © Michael Tierney
Available Tools
8 toolsask_localA
Send a one-shot prompt to a local Ollama model and return its text response.
Use for any handoff where the cloud model's full reasoning isn't needed: drafts, boilerplate, simple extractions, formatting, or quick lookups. Runs on the user's own GPU and consumes no cloud-LLM usage. Returns the model's raw text completion.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. | |
| prompt | Yes | The task or question to send to the model. | |
| system | No | Optional system prompt to set the model's role or behavior. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full burden. It discloses key traits: one-shot (no conversation state), local execution, no cloud usage, and raw text return. It does not mention error conditions, model availability, or potential latency, but these are secondary for a simple tool. The provided information is accurate and materially helps an agent predict behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact paragraphs. The purpose sentence is front-loaded, followed by usage context and a cost note. Every sentence serves a purpose; no fluff or repetition. Ideal length for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage, cost, and output format clearly. It omits potential failure modes (e.g., model not installed) and doesn't mention that the model must be pre-pulled, but given the tool's simplicity and the presence of an output schema, these are optional. An agent can safely invoke it with just the required prompt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — all three parameters (prompt, model, system) have descriptions. The tool description adds little beyond the schema: it frames the action as 'one-shot' and mentions 'raw text completion,' which relates to output rather than parameter semantics. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair: 'Send a one-shot prompt to a local Ollama model and return its text response.' It distinguishes from siblings by emphasizing 'one-shot' (vs chat_local) and 'raw text completion' (vs specialized extract/summarize tools). The example use cases reinforce its generic scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: 'Use for any handoff where the cloud model's full reasoning isn't needed' and lists concrete tasks (drafts, boilerplate, simple extractions, formatting, quick lookups). It also provides a cost rationale ('Runs on the user's own GPU and consumes no cloud-LLM usage'), which guides selection. While it doesn't name sibling alternatives, the categories imply exclusions (e.g., complex reasoning for cloud models, specialized extraction for extract_local).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chat_localA
Hold a multi-turn chat against a local Ollama model.
Use instead of ask_local when the handoff needs more than one turn of
context — a running conversation or a system + user + assistant history.
Runs locally at no cloud cost. Returns the model's next assistant message
as text.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. | |
| messages | Yes | Conversation as a list of {"role": "user"|"assistant"|"system", "content": str} messages, in order. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that the tool runs locally at no cloud cost and returns the next assistant message as text, which are useful behavioral traits. It does not explicitly mention potential failure modes (e.g., model unavailability, timeouts) or confirm it is read-only, but these are less critical for a chat tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. The first paragraph states the purpose and usage guidance, and the second adds contextual details. Every sentence contributes value without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description does not need to explain the return format in depth. It covers the tool's purpose, when to use it versus an alternative, and the basic return type, making it complete for an agent to correctly invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents both parameters (model and messages) with detailed descriptions, achieving 100% coverage. The tool description does not add additional parameter-level semantics, but the schema is sufficient, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool's specific function: holding a multi-turn chat against a local Ollama model. It clearly distinguishes from the sibling `ask_local` by the multi-turn requirement, so an agent can immediately understand its purpose and differentiate it from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use instead of `ask_local` when the handoff needs more than one turn of context' and provides concrete examples ('a running conversation or a system + user + assistant history'), leaving no ambiguity about when to select this tool over its alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
code_review_localA
Run a quick first-pass code review using the local coder model.
Use as a cheap pre-filter before asking the cloud model for a deeper review: it catches obvious bugs, style issues, and risky patterns. Runs locally at no cloud cost. Returns review notes as text.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. | |
| diff_or_code | Yes | A unified diff or a code block to review. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It discloses local execution ('Runs locally at no cloud cost'), the output format ('Returns review notes as text'), and the shallow/first-pass scope. It does not mention latency, determinism, or resource constraints, but the most important behavioral traits are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the operation, the second covers use context, and the third covers behavior and output. A slight redundancy exists between 'quick first-pass' and 'cheap pre-filter,' but it earns its place by adding the cost angle. No wasted or filler words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with a fully documented schema and an output schema present, the description provides the needed context: purpose, usage scenario, execution environment, and return type. It would benefit from an explicit 'when not to use' note, but nothing essential is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both `diff_or_code` and the optional `model`. The description adds no new parameter-level meaning beyond reinforcing that the input is a diff or code block. This matches the baseline for fully covered schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific action ('Run a quick first-pass code review') and a specific resource ('local coder model'), and distinguishes this tool from sibling local tools by naming what it catches: 'obvious bugs, style issues, and risky patterns.' No other sibling tool description suggests this exact code-review role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly frames the tool as a 'cheap pre-filter before asking the cloud model for a deeper review,' giving an unambiguous when-to-use context and an implied alternative. It does not name sibling tools directly or state explicit exclusions, but the pre-filter/deeper-review distinction is concrete enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_commit_message_localA
Draft a conventional-style commit message from a diff using the local model.
Use for routine commits where the cloud model's analysis isn't needed — it is cheap and fast. Runs locally at no cloud cost. Returns a single commit message (subject plus optional body) as text.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | Yes | A staged diff, e.g. the output of `git diff --staged`. | |
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses local execution, no cloud cost, and the exact return format (single commit message as text), which are important operational traits. It could be more explicit about model availability or failure modes, but for a drafting tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two paragraphs, front-loaded with the main purpose and efficiently covering usage and output. There is minor redundancy between 'cheap and fast' and 'no cloud cost', but overall it is appropriately sized and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with an existing output schema, the description covers the core purpose, usage context, and return type. It does not explicitly state when to prefer a cloud alternative, but the implicit guidance is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents both parameters (diff and model) with 100% coverage. The description adds no parameter-specific detail beyond what the schema provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action—drafting a conventional-style commit message from a diff—while explicitly distinguishing this local-model tool from the cloud-model alternative and from sibling tools like ask_local or summarize_local. The verb, resource, and style constraints are all present.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to use the tool ('routine commits where the cloud model's analysis isn't needed') and emphasizes the cost/latency benefits. However, it does not name the alternative tool directly or list explicit when-not-to-use cases beyond that implied exclusion, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_localA
Extract specific information from a text block using the local model.
Use to pull structured facts out of unstructured text — function names, URLs,
error codes, TODO comments, dependency names — without spending cloud-model
tokens. Runs locally at no cloud cost. Returns the extracted items as text,
shaped by what_to_extract.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The source text to extract from. | |
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. | |
| what_to_extract | Yes | What to pull out, e.g. 'all function definitions' or 'every URL in the file'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that the tool runs locally, incurs no cloud cost, and returns extracted items as text shaped by `what_to_extract`. This gives the agent enough behavioral context to set expectations, though it does not mention latency, size limits, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose. The only weakness is minor redundancy: "without spending cloud-model tokens" is repeated by the later sentence "Runs locally at no cloud cost." Overall it is still tight and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter tool with a full schema and output schema, the description is largely complete: it explains what the tool does, the kind of extraction to request, and the return form. It could improve by explicitly noting how it differs from sibling local tools, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds some context by noting that `what_to_extract` shapes the output and providing example values, but this is reinforcing rather than substantially expanding on the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Extract specific information from a text block." It then gives concrete target types (function names, URLs, error codes, TODO comments, dependency names), making the tool's purpose unmistakable and clearly distinct from the sibling chat/ask/summarize tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it: to pull structured facts out of unstructured text, and positions it as a cost-saving local alternative to cloud-model extraction. It does not name explicit alternatives or when-not-to-use cases, so it falls short of a 5, but the guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List the Ollama models installed locally and available to these tools.
Use to discover valid values for the model parameter before pinning a
specific model. Returns a list of model-name strings.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool lists locally installed models and returns model-name strings, but it doesn't explicitly state whether it is read-only or mention potential failure modes (e.g., no models installed). For a simple list operation, this is adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundancy. The primary action is front-loaded, followed by a clear usage hint and return type. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there are no parameters and the output schema is present (though not shown in the prompt), the description provides the essential information: what it lists, why to use it, and what it returns. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is empty and description coverage is 100% by definition. Per the rubric, 0 params earns a baseline of 4. The description adds nothing about parameters because none exist, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('Ollama models installed locally'), and clarifies its role as a discovery tool distinct from the sibling task tools (ask_local, chat_local, etc.). It also specifies the return type, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use to discover valid values for the `model` parameter before pinning a specific model,' which provides clear context for when to call it. It doesn't explicitly state when not to use it, but the intent is evident given the sibling tools are the consumers of the model names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_infoA
Report the server's effective configuration.
Returns the default model, Ollama base URL, context size, and request timeout. Use to confirm which model the tools will use by default or to debug connectivity. Returns a JSON object.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses that the tool returns a JSON object and names its payload fields, and 'Report' implies a non-mutating operation. It does not discuss error behavior or access constraints, but the behavior of a config-reporting tool is sufficiently clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four concise sentences with no filler. The core purpose is front-loaded, followed by return details and actionable use cases; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a zero-parameter tool with no output schema, the description is complete: it states what is returned, enumerates the key fields, and gives concrete usage scenarios. An agent can decide to call it and interpret the result without further information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description compensates for the absence of parameter documentation by explaining what the returned configuration report contains, which is the only semantic information an agent needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Report', with a concrete resource, 'the server's effective configuration', and enumerates the exact fields returned (default model, Ollama base URL, context size, request timeout). This clearly distinguishes it from siblings like list_models, which lists models rather than reporting overall server configuration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: to confirm the default model the tools will use or to debug connectivity. It does not mention when not to use it or name alternatives, but for a zero-parameter diagnostic tool the guidance is specific and useful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_localA
Summarize a block of text using the local model.
Use to offload long files, logs, transcripts, or docs the cloud model does
not need to fully ingest — call this instead of reading a large blob into the
cloud context. Runs on the user's GPU at no cloud cost. Returns a concise
prose summary; pass focus to bias it toward what matters.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The content to summarize; may be long (the local context window is configurable). | |
| focus | No | Optional hint to steer the summary, e.g. 'errors and stack traces' or 'API surface only'. | |
| model | No | Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It states that the tool 'Runs on the user's GPU at no cloud cost' and 'Returns a concise prose summary,' which are useful behavioral details. It does not mention potential latency, failures, or size limits beyond the schema, but the disclosure provided is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a clear opening definition, a practical use-case sentence, a cost/location statement, and a return-value sentence. Every sentence contributes useful information without redundancy or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists and all parameters are documented, the description covers the essential usage context well. It explains when to call the tool, what it returns, and why local execution is beneficial. It could be more complete by explicitly distinguishing from sibling tools like extract_local or ask_local, but nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds a small note about passing `focus` to bias the summary, but this largely repeats what the schema's focus parameter already says. A baseline 3 is appropriate because the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as 'summarize a block of text using the local model', which is a specific verb and resource. It communicates the core purpose well, but it does not explicitly differentiate from sibling tools like extract_local or ask_local, so it falls short of full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: 'Use to offload long files, logs, transcripts, or docs the cloud model does not need to fully ingest.' It also tells the agent to call this 'instead of reading a large blob into the cloud context,' giving practical usage guidance. However, it does not name specific alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.3- Changed
ask_local3 fields changed- added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model." - added
Input schema / properties / prompt / descriptionAdded value: +"The task or question to send to the model." - added
Input schema / properties / system / descriptionAdded value: +"Optional system prompt to set the model's role or behavior."
- Changed
chat_local2 fields changed- added
Input schema / properties / messages / descriptionAdded value: +"Conversation as a list of {\"role\": \"user\"|\"assistant\"|\"system\", \"content\": str} messages, in order." - added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model."
- Changed
code_review_local2 fields changed- added
Input schema / properties / diff_or_code / descriptionAdded value: +"A unified diff or a code block to review." - added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model."
- Changed
draft_commit_message_local2 fields changed- added
Input schema / properties / diff / descriptionAdded value: +"A staged diff, e.g. the output of `git diff --staged`." - added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model."
- Changed
extract_local3 fields changed- added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model." - added
Input schema / properties / text / descriptionAdded value: +"The source text to extract from." - added
Input schema / properties / what_to_extract / descriptionAdded value: +"What to pull out, e.g. 'all function definitions' or 'every URL in the file'."
- Changed
summarize_local3 fields changed- added
Input schema / properties / focus / descriptionAdded value: +"Optional hint to steer the summary, e.g. 'errors and stack traces' or 'API surface only'." - added
Input schema / properties / model / descriptionAdded value: +"Ollama model name to run, e.g. 'llama3.1' or 'qwen2.5-coder'. Omit to use the server's configured default model." - added
Input schema / properties / text / descriptionAdded value: +"The content to summarize; may be long (the local context window is configurable)."
8 tool updates
v0.1.0- First observed
ask_local - First observed
chat_local - First observed
code_review_local - First observed
draft_commit_message_local - First observed
extract_local - First observed
list_models - First observed
server_info - First observed
summarize_local
TDQS
Scored across 8 tools
Each task-specific tool (summarize, code_review, draft_commit_message, extract) has a clear purpose, and ask_local vs chat_local cleanly separate one-shot from multi-turn. The main overlap is that ask_local is framed as a catch-all ('simple extractions, formatting'), which could lead an agent to use it instead of the more specialized extract or code_review tools.
Most operational tools follow a verb_object_local pattern (ask_local, chat_local, summarize_local, extract_local), and all names use snake_case. Deviations include code_review_local and server_info, which don't fit the verb-object shape, and list_models/server_info lack the _local suffix.
Eight tools is well-scoped: six local-model task tools plus two discovery/config tools. Each tool has a distinct role with no obvious redundancy or missing foundational piece.
The set covers generic one-shot generation, multi-turn chat, summarization, extraction, code review, commit messages, model listing, and server config—strong coverage for a local handoff server. It lacks more advanced Ollama operations like streaming, model management, or embeddings, but those appear outside the declared handoff purpose.
Maintenance
Related MCP Connectors
Cost-optimized LLM model routing recommendations for autonomous AI agents
LLM Orchestration Agent
150+ vertical AI expert bots as agent tools. $1 bots run on YOUR machine - your data stays yours.
LLM Orchestration Agent 2
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables Claude to delegate coding tasks to local Ollama models, reducing API token usage by up to 98.75% while leveraging local compute resources. Supports code generation, review, refactoring, and file analysis with Claude providing oversight and quality assurance.383 npm25AGPL 3.0
- AlicenseAqualityDmaintenanceOptimizes token costs by intelligently delegating low-complexity tasks to local LLMs via LiteLLM, enabling cost-effective development workflows.31MIT
- AlicenseAqualityCmaintenanceEnables Claude Code to offload routine code generation and text processing tasks to a local Ollama LLM, saving Cloud API tokens and costs with automatic model selection and security features.1163 npm4Apache 2.0
- AlicenseAqualityBmaintenanceEvaluates task suitability for local models before cloud API calls, routing to Ollama or similar to reduce costs.157 npmMIT