ollama-handoff
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OLLAMA_URL | No | Base URL of the Ollama server | http://localhost:11434 |
| OLLAMA_NUM_CTX | No | Context window in tokens | 32768 |
| OLLAMA_TIMEOUT_S | No | Per-request timeout, seconds | 600 |
| OLLAMA_KEEP_ALIVE | No | How long to keep the model resident in VRAM | 30m |
| OLLAMA_DEFAULT_MODEL | No | Default model for handoffs | qwen2.5-coder:14b |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| ask_localA | Send a one-shot prompt to a local Ollama model and return its text response. Use for any handoff where the cloud model's full reasoning isn't needed: drafts, boilerplate, simple extractions, formatting, or quick lookups. Runs on the user's own GPU and consumes no cloud-LLM usage. Returns the model's raw text completion. |
| chat_localA | Hold a multi-turn chat against a local Ollama model. Use instead of |
| summarize_localA | Summarize a block of text using the local model. Use to offload long files, logs, transcripts, or docs the cloud model does
not need to fully ingest — call this instead of reading a large blob into the
cloud context. Runs on the user's GPU at no cloud cost. Returns a concise
prose summary; pass |
| code_review_localA | Run a quick first-pass code review using the local coder model. Use as a cheap pre-filter before asking the cloud model for a deeper review: it catches obvious bugs, style issues, and risky patterns. Runs locally at no cloud cost. Returns review notes as text. |
| draft_commit_message_localA | Draft a conventional-style commit message from a diff using the local model. Use for routine commits where the cloud model's analysis isn't needed — it is cheap and fast. Runs locally at no cloud cost. Returns a single commit message (subject plus optional body) as text. |
| extract_localA | Extract specific information from a text block using the local model. Use to pull structured facts out of unstructured text — function names, URLs,
error codes, TODO comments, dependency names — without spending cloud-model
tokens. Runs locally at no cloud cost. Returns the extracted items as text,
shaped by |
| list_modelsA | List the Ollama models installed locally and available to these tools. Use to discover valid values for the |
| server_infoA | Report the server's effective configuration. Returns the default model, Ollama base URL, context size, and request timeout. Use to confirm which model the tools will use by default or to debug connectivity. Returns a JSON object. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Each task-specific tool (summarize, code_review, draft_commit_message, extract) has a clear purpose, and ask_local vs chat_local cleanly separate one-shot from multi-turn. The main overlap is that ask_local is framed as a catch-all ('simple extractions, formatting'), which could lead an agent to use it instead of the more specialized extract or code_review tools.
Most operational tools follow a verb_object_local pattern (ask_local, chat_local, summarize_local, extract_local), and all names use snake_case. Deviations include code_review_local and server_info, which don't fit the verb-object shape, and list_models/server_info lack the _local suffix.
Eight tools is well-scoped: six local-model task tools plus two discovery/config tools. Each tool has a distinct role with no obvious redundancy or missing foundational piece.
The set covers generic one-shot generation, multi-turn chat, summarization, extraction, code review, commit messages, model listing, and server config—strong coverage for a local handoff server. It lacks more advanced Ollama operations like streaming, model management, or embeddings, but those appear outside the declared handoff purpose.