Best Ollama MCP Servers
Ollama is an open-source project that allows you to run large language models (LLMs) locally on your own hardware, providing a way to use AI capabilities privately without sending data to external services.
Why this server?
Allows managing local Ollama models: discover installed and loaded models, pull/remove models, load/unload from memory, check hardware fit, and offload inference (completion and embedding) to local Ollama instances.
AlicenseAqualityAmaintenanceUnified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.16763Creative Commons Attribution Non Commercial No Derivatives 4.0 InternationalWhy this server?
Enables local RAG question answering via Ollama, providing a fully offline and privacy-preserving LLM mode for the `/ask` endpoint when configured.
AlicenseAqualityBmaintenanceLocal-first agent memory & knowledge service: hybrid retrieval (vector + BM25/RRF) + memory governance (decay, dedup, freshness), REST + MCP dual protocol, fully offline without LLM. Python.3582Apache 2.0Why this server?
Provides tools for predicting LLM performance and optimizing local LLM inference using Ollama, including checking compatibility of specific models and getting recommendations.

Yamaru Hardware Probeofficial
AlicenseAqualityDmaintenanceExpert system hardware probe and performance diagnostic engine for AI, Gaming, and High-Performance workflows. Provides deep system insights such as real-time monitoring, thermal diagnostics, and LLM optimization.11577Apache 2.0Why this server?
Uses a local Ollama model to draft redaction configuration by analyzing private files.
AlicenseAqualityAmaintenanceProvides redacted access to a private local knowledgebase for coding agents, allowing them to inspect files while hiding sensitive names and identifiers.14131MITWhy this server?
Allows AI agents to route inference requests through the VibOps proxy to Ollama, with cost attribution and logging.

vibops-mcpofficial
AlicenseAqualityAmaintenanceVibOps MCP is the control plane between your AI agents and your GPU infrastructure. 74 tools covering: GPU fleet management (deploy, scale, monitor across NVIDIA, AMD, Intel, AWS, Google, Groq), Agent Infrastructure Control Plane (per-agent GPU cost, budget enforcement, model policies, dependency graph), governance (AI Act, SOC 2, immutable HMAC audit chain), and GPU FinOps (chargeback, waste..)..7418MITWhy this server?
Allows using Ollama as a local LLM provider for summarizing video content into notes.
AlicenseAqualityAmaintenanceMCP server that converts video links into AI-generated Markdown notes, with tools for task management, transcription engines, and LLM providers.22252MITWhy this server?
Allows querying local Ollama models via HTTP.

MCP Rubber Duckofficial
AlicenseAqualityBmaintenanceAn MCP server that bridges multiple LLMs (OpenAI-compatible APIs and CLI coding agents) for collaborative debugging and diverse AI perspectives.122945MITWhy this server?
Provides security and safety for AI agents using locally hosted Ollama models, with no API key required.
Why this server?
Provides integration with Ollama as a local LLM backend for the Gradio chat UI, allowing natural-language querying of the 3DCityDB.

3DCityDB MCP Serverofficial
AlicenseAqualityCmaintenanceEnables AI assistants to interact with 3DCityDB v5 through natural language, dynamically resolving object classes, properties, and codelists to answer spatial questions and execute SQL queries on CityGML data.1410Apache 2.0