Local AI MCP
OfficialServer Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OLLAMA_HOST | No | Ollama base URL (scheme optional; added if missing). | http://localhost:11434 |
| LMSTUDIO_HOST | No | LM Studio base URL. | http://localhost:1234 |
| LOCAL_AI_DETECT_TIMEOUT_MS | No | Timeout for provider auto-detection probes. | 1500 |
| LOCAL_AI_REQUEST_TIMEOUT_MS | No | Timeout for normal requests (inference, pull progress, etc.). | 120000 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_providersA | List the configured model runtime providers (Ollama, LM Studio, llama.cpp, optional OpenAI-compat / Moonshot) with their host, whether they are detected/live, and their capabilities. Optionally filter to a single provider. |
| list_modelsA | List models installed/available on each detected provider. Without a provider arg, aggregates across all detected providers keyed by provider. |
| list_loadedA | List models currently loaded into memory on each detected provider. Without a provider arg, aggregates across all detected providers keyed by provider. |
| model_infoA | Show detailed metadata for a specific model (family, parameter size, quantization, context length). Without a provider arg, queries all detected providers. |
| pull_modelA | HEAVY: Download/pull a model onto a provider. WARNING: this may download multiple gigabytes and can take a long time. Without a provider arg, attempts the pull on every detected provider. |
| remove_modelA | DESTRUCTIVE: Permanently delete a model from a provider. Requires confirm:true; without it the action is refused. The provider argument is REQUIRED so a delete cannot fan out across every detected runtime. |
| load_modelA | Load a model into memory so it is ready for inference. Optionally set keepAlive (e.g. '5m', '1h'). Without a provider arg, loads on every detected provider. |
| unload_modelA | Unload a model from memory to free VRAM/RAM. Without a provider arg, unloads from every detected provider. |
| health_checkA | Check whether each provider's local runtime is reachable and report its version. Without a provider arg, checks all configured providers. |
| system_resourcesA | Report local hardware resources: platform, total/free RAM, CPU count, and detected GPUs with VRAM. Used to reason about which models can run locally. |
| fit_checkA | Determine whether a model fits on the local hardware. Resolves the model weight size from the provider or the static catalog (or an explicit sizeBytes), estimates KV-cache overhead for a context length (default 4096), then compares weight+KV against free GPU VRAM, falling back to system RAM. Returns fits, target (gpu/cpu/none), weight/kv/required/available bytes. |
| benchmarkA | HEAVY: Runs REAL inference. Executes one small completion against a loaded/loadable model and measures latency (ms) and throughput (tokens/sec). This consumes compute and may load the model. Without a provider arg, runs on the first detected provider. |
| search_availableA | Search the built-in catalog of well-known local models by name, family, or task. Note: this searches a curated static catalog, not the full live Ollama library (https://ollama.com/library). |
| suggest_modelA | Recommend local models for a task (chat/code/embed/vision/reasoning/general), ranked by task match and then by whether they fit the detected free GPU VRAM or system RAM. Returns a ranked list with fit flags. |
| completeA | DELEGATED INFERENCE: Offload a text/chat completion to a local model runtime for cost savings and privacy (data never leaves the machine). This is NOT a chat feature for the user; it delegates work to a local LLM. Provide either prompt or messages. Streams tokens via MCP progress notifications when the client supplies a progressToken (stream defaults to true). Without a provider arg, uses the first detected provider. |
| embedA | DELEGATED EMBEDDINGS: Offload embedding generation to a local model runtime for cost savings and privacy. Accepts a single string or an array of texts. Without a provider arg, uses the first detected provider. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 16 tools
Each tool targets a distinct operation: model management (unload, load, list, pull, remove, info), system information (list_providers, health_check, system_resources), analysis (fit_check, benchmark, suggest_model, search_available), and inference (complete, embed). No two tools have overlapping purposes.
Tool names predominantly follow a verb_noun pattern (unload_model, list_providers, pull_model). Some names like model_info, system_resources, and benchmark are noun-heavy or single-verb but still clear and readable. Consistent snake_case throughout.
With 16 tools, the server is well-scoped for its domain. It covers model lifecycle, system diagnostics, analysis, and inference without being bloated or sparse.
The tool set covers the full lifecycle: pull, load, unload, remove, list, info, and inference (complete, embed). It also includes system checks (health, resources, fit), search, recommendations, and benchmarks. No obvious gaps for the stated purpose of local AI model management and inference.