offline-mcp
Allows running local inference using Ollama models, checking Ollama status, and listing available models on the local machine.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@offline-mcprun local inference: 'hello in Swahili?'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
offline-mcp
Compatible with claude-sonnet-5 (released 2026-06-30) — Anthropic's most agentic
Sonnet yet. Runs multi-step tool chains end-to-end without stopping short.
Install: pip install offline-mcp · Use with any MCP client.
Local AI inference infrastructure — Ollama wrapper, open weights directory, degraded-mode guide for East Africa.
Why: Never assume OpenAI survives, Anthropic stays accessible, or export controls disappear. This matters more in Africa than anywhere else.
1st world equivalent: Ollama, LLaMA, Mistral local deployment
Why This Exists: Data Sovereignty
"If you take the deal, you're going to be exploited. If you don't take it, you're going to die."
— Frank Ssekamwa, Ugandan digital rights expert
Across the Global South, AI and health data from communities is being extracted, processed abroad, and used to build models whose value flows away from the communities that generated it.
offline-mcp is the sovereignty floor of the East Africa coordination stack.
When this runs on a Raspberry Pi 4 with a 50W solar panel and a 256GB SD card:
Health data stays in the clinic. Health guidance comes from local models.
Land records stay in the land office. Queries don't touch foreign servers.
Civic data stays in the county. AI assistance runs without internet.
No API key. No cloud dependency. No data leaving the community.
The models available via offline-mcp (Llama 3.2, Qwen 2.5) run entirely on device.
Community data used to generate AI outputs creates no dataset sent back to model providers.
This is not a privacy feature. It is the architectural foundation of digital independence.
Related MCP server: Ollama MCP Server
Research Foundation
This server implements patterns validated by peer-reviewed research on offline-first AI for bandwidth-constrained environments:
arXiv:2603.03339 (2026) — Offline-First LLM Architecture for Adaptive Learning in Low-Connectivity Environments — confirms that meaningful AI support is achievable with hardware-aware model selection when designed for infrastructure-limited deployment. Key finding: offline-first is a complementary paradigm, not a compromise.
Design principles applied:
Local-first: all core operations execute without internet connectivity
Graceful degradation: reduced functionality beats no functionality
JSONL queue: events accumulated offline sync when connectivity returns
Hardware-aware: tool selection adapts to available compute
East Africa context: Kenya, Tanzania, Uganda — rural areas with intermittent connectivity represent the primary deployment target. This server is not designed for ideal conditions. It is designed for real ones.
Install
pip install offline-mcpTools (6)
Tool | Description |
| Check if Ollama is running locally and list available models |
| Run a prompt through a local Ollama model |
| Best open-weight models for East Africa use cases |
| 4-level degraded mode architecture for offline operation |
| Directory of open-weight models with Africa language support |
| Deployment guide for laptop, server, Raspberry Pi, Android |
Context
Runs on a 50W solar panel + Raspberry Pi 4. Viable for rural Kenya clinics, schools, and community offices.
License
MIT © Gabriel Mahia | contact@aikungfu.dev
Part of the East Africa Coordination Stack
This MCP server is one of 32 tools in the Kenya coordination infrastructure.
Connect it to africa-coord-bus —
the coordination event bus that routes signals between domains automatically.
pip install africa-coord-busAll 32 servers: pypi.org/user/gmahia Live demo: coord-cascade-demo
IP & Collaboration
MIT licensed. Feedback via GitHub Issues only — pull requests are not accepted. Demo data is labeled DEMO and is not suitable for operational decisions. Full policy: docs/architecture/IP_POLICY.md. Security reports: see SECURITY.md.
Part of the East Africa coordination stack
Install & run:
pip install reli-cli && reli list— 33 MCP servers on the official MCP Registry underio.github.gabrielmahiaEvaluate any model on Swahili agent tasks: kipimo · dataset · leaderboard
Coordinate across servers: africa-coord-bus — offline-first event bus with a built-in Kenya routing table
Datasets: huggingface.co/gmahia · Docs hub: nairobi-stack
Model-agnostic by design: closed APIs, open-weight models, and small distilled models are all first-class citizens.
Available Tools
6 toolscheck_ollama_statusA
Check if Ollama is running locally and list available models.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description describes the tool's behavior as a read-only check and listing operation. No annotations are provided, so the description carries the burden; it is adequate but does not disclose edge cases (e.g., behavior if Ollama is not installed) or output details beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that captures the essential purpose without superfluous information. Every word adds value; no restructuring needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters) and the presence of an output schema, the description is largely complete. It covers the tool's main function. A minor gap: it could mention whether errors are thrown if Ollama is not installed, but this is not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (vacuously). The description adds no parameter details because none exist, which is acceptable. Baseline score for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Check') and the resource ('if Ollama is running locally and list available models'). It distinguishes this tool from siblings like run_local_inference or list_recommended_models by focusing on status checking rather than computation or recommendations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for verifying Ollama availability and listing models, but does not explicitly state when to use this over alternatives (e.g., before running inference) or when not to use it. Sibling tools provide context but no direct guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
degraded_mode_guideA
Guide for operating AI systems when cloud connectivity fails.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It correctly implies read-only behavior as a guide, but does not disclose any further behavioral traits such as return format or performance characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, clear sentence that is front-loaded and contains no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and presence of an output schema (not shown but indicated), the description is adequate. It could be slightly improved by hinting at the output nature, but completeness is high given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 0 parameters and schema description coverage is 100%, so baseline is 4. Description adds no parameter info, but none needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it is a guide for operating AI systems when cloud connectivity fails, specifying both the resource (guide) and the context (degraded mode). It distinguishes itself from sibling tools which are about checking status, running inference, models, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. There is no mention of when not to use it or contextual cues for selection among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recommended_modelsB
List recommended open-weight models for East Africa AI use cases.
| Name | Required | Description | Default |
|---|---|---|---|
| use_case | No | ||
| max_ram_gb | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only says 'list,' implying a read-only operation. It does not disclose aspects like authentication needs, rate limits, or whether the listing is paginated or cached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no superfluous words, making it highly concise and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description does not hint at return structure or pagination. It is minimal for a list tool with two parameters and multiple siblings, missing usage cues.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the two parameters ('use_case' and 'max_ram_gb'). The phrase 'for East Africa AI use cases' hints at 'use_case' but adds no practical semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'recommended open-weight models for East Africa AI use cases,' which is specific and distinguishes it from siblings like 'open_weights_directory' that likely lists all models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives such as 'open_weights_directory' or 'local_deployment_guide.' The description lacks context for appropriate usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
local_deployment_guideC
Guide to deploying local AI inference on modest hardware in Kenya/East Africa.
| Name | Required | Description | Default |
|---|---|---|---|
| device_type | No | laptop |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must disclose behavioral traits. It only states it is a guide, which suggests a read-only operation, but does not confirm whether it accesses external resources, requires internet, or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that is concise and easy to parse. No redundancy exists, but it could be expanded slightly to include parameter details without losing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (one optional parameter, output schema exists), the description is incomplete. It lacks context on how the guide is delivered (e.g., text, steps), what the output schema contains, and any usage hints. The presence of an output schema partially mitigates the need to describe return values, but the description remains insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has a single optional parameter (`device_type`), but the description provides no explanation of its meaning, allowed values, or how it affects the output. With 0% schema coverage, the description fails to compensate, leaving the agent without guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the subject matter (deploying local AI inference on modest hardware) and geographic scope (Kenya/East Africa), distinguishing it from sibling tools like `run_local_inference` or `check_ollama_status`. However, it lacks a verb (e.g., 'provides a guide' or 'returns deployment steps') so the action is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus siblings such as `list_recommended_models` or `degraded_mode_guide`. The description does not mention prerequisites, alternative tools, or scenarios where the guide is applicable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_weights_directoryC
Directory of open-weight AI models suitable for East Africa civic use cases.
| Name | Required | Description | Default |
|---|---|---|---|
| use_case | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and description does not disclose behavioral traits such as side effects, authentication needs, or whether it performs a read-only operation. The description is too brief to inform the agent about tool behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at one sentence, but it sacrifices necessary detail for brevity. Every word earns its place, but the content is insufficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (1 optional param, no annotations, output schema exists but not described), the description fails to provide a complete picture. It lacks return format overview, usage constraints, or how the output schema complements the directory listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and description does not explain the single parameter 'use_case' (type, acceptable values, or effect on results). The description adds no meaning beyond the schema, which itself is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description is a noun phrase 'Directory...' rather than a verb phrase indicating an action. It states the domain (open-weight AI models for East Africa civic use) but doesn't specify what the tool does (e.g., list, search, or retrieve). Compared to sibling 'list_recommended_models', the purpose is ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'list_recommended_models' or 'check_ollama_status'. Lacks any context about preferred use cases or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_local_inferenceC
Run a prompt through a local Ollama model.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | llama3.2:3b | |
| prompt | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose traits. It only says 'run a prompt', implying a blocking operation, but does not mention execution time, error handling, or any side effects. Minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the action and resource. No extraneous words, efficiently communicates the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description omits important context such as the need for Ollama to be running, model availability, or potential delays. Insufficient for a user to understand full requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, requiring the description to add meaning. The description does not explain the purpose of 'prompt' or 'model' beyond their names, failing to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the verb 'run' and the resource 'a prompt through a local Ollama model'. It is distinct from sibling tools like check_ollama_status and list_recommended_models, though it does not explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites or context provided. The description lacks any usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
6 tool updates
v0.1.4- First observed
check_ollama_status - First observed
degraded_mode_guide - First observed
list_recommended_models - First observed
local_deployment_guide - First observed
open_weights_directory - First observed
run_local_inference
TDQS
Each tool has a clearly distinct purpose: status checking, running inference, listing recommended models, providing a directory, and offering guides. No overlap in functionality.
All names use snake_case and follow a verb_noun or adjective_noun pattern, with minor variation between imperative verbs and descriptive nouns. Consistent enough for clear identification.
Six tools is well-scoped for an offline AI inference and guidance server, covering core operations and supplementary resources without bloat.
The tool set covers the full expected workflow: status check, inference, model recommendations, directory, and guides for deployment and degraded mode. No obvious gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
AI LLM with Gemini, MiniMax, Replicate, OpenRouter. Vision, search, code review. USDC on Base.
150+ vertical AI expert bots as agent tools. $1 bots run on YOUR machine - your data stays yours.
HiveCompute MCP Server — decentralized inference router for AI agents
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables interaction with locally running Ollama models through chat, generation, and model management operations. Supports listing, downloading, and deleting models while maintaining conversation history for interactive sessions.203MIT
- AlicenseBqualityDmaintenanceEnables complete local Ollama management including listing models, chatting with local LLMs, starting/stopping the server, and getting intelligent model recommendations for specific tasks through natural language commands.94MIT
- FlicenseNot gradedqualityDmaintenanceEnables text-only AI models to analyze images via local Ollama multimodal models. Supports image analysis, OCR, and multi-image comparison entirely offline.-
- FlicenseBqualityBmaintenanceEnables AI agents to interact with local Ollama models for text generation and tool calling with prompt injection protection.5-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/gabrielmahia/offline-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server