Skip to main content
Glama
README.md
# Ollama MCP Server with Qwen3-Coder

A Model Context Protocol (MCP) server that provides web search, web fetch, and chat completion capabilities using Ollama's Qwen3-coder models. Designed to work seamlessly with Cursor IDE and other MCP-compatible clients.

## Features

- **Smart Model Selection**: Automatically uses `qwen3-coder:480b-cloud` when API key is available, falls back to local models
- **Web Search**: Powered by Ollama's hosted search API
- **Web Fetch**: Retrieve and parse content from specific URLs
- **Chat Completion**: High-quality code-focused conversations with Qwen3-coder
- **Search & Chat**: Combined tool that searches the web and generates responses based on results
- **Automatic Fallback**: Falls back to local models (`qwen3:4b`, `qwen3:7b`, etc.) when cloud is unavailable

## Installation

1. Clone or download this repository
2. Install dependencies using uv (recommended) or pip:

```bash
# Using uv (recommended)
uv sync

# Or using pip
pip install -e .
```

## Configuration

### Environment Variables

- `OLLAMA_API_KEY` (optional): Required for cloud models and web search/fetch functionality
- `OLLAMA_HOST` (optional): Ollama server URL (default: `http://localhost:11434`)

### For Cursor IDE

1. Open Cursor IDE settings
2. Go to "Extensions" → "MCP" → "Manage MCP Servers"
3. Add the configuration from `cursor-mcp-config.json`:

```json
{
  "mcpServers": {
    "ollama-qwen-mcp": {
      "type": "stdio", 
      "command": "uv",
      "args": ["run", "python", "-m", "ollamamcp.server"],
      "env": {
        "OLLAMA_API_KEY": "your_api_key_here_or_remove_for_local_only",
        "OLLAMA_HOST": "http://localhost:11434"
      }
    }
  }
}
```

### For Other MCP Clients

The server can be run directly:

```bash
# With API key for cloud features
OLLAMA_API_KEY=your_key uv run python -m ollamamcp.server

# Local only (no web search/fetch)
uv run python -m ollamamcp.server
```

## Available Tools

### 1. `web_search`
Perform web searches using Ollama's hosted API.

**Parameters:**
- `query` (str): Search query
- `max_results` (int): Maximum results (default: 3, max: 20)

**Requires:** `OLLAMA_API_KEY`

### 2. `web_fetch`
Fetch content from a specific URL.

**Parameters:**
- `url` (str): Absolute URL to fetch

**Requires:** `OLLAMA_API_KEY`

### 3. `chat_completion`
Generate responses using Qwen3-coder models.

**Parameters:**
- `messages` (list): Conversation messages
- `model` (str, optional): Override model selection
- `temperature` (float): Sampling temperature (default: 0.7)
- `max_tokens` (int, optional): Maximum tokens to generate

### 4. `search_and_chat`
Combined web search and chat completion.

**Parameters:**
- `query` (str): Search query and question
- `search_results` (int): Number of results (default: 3)
- `model` (str, optional): Override model selection  
- `temperature` (float): Sampling temperature (default: 0.7)

**Requires:** `OLLAMA_API_KEY`

### 5. `get_available_models`
Get information about available models and configuration.

**Returns:** Current model, availability status, and model lists.

## Model Fallback Strategy

1. **Cloud First**: `qwen3-coder:480b-cloud` (if API key available)
2. **Local Fallbacks** (in order):
   - `qwen3:4b`
   - `qwen3:7b` 
   - `qwen3:14b`
   - `qwen2.5-coder:7b`
   - `qwen2.5-coder:3b`
   - `qwen2.5-coder:1.5b`

The server automatically pulls local models if they're not available but Ollama is running.

## Usage Examples

### In Cursor IDE

Once configured, you can use natural language to:

- "Search for the latest Python async/await best practices"
- "Fetch the documentation from https://docs.python.org/3/library/asyncio.html"
- "What are the new features in the latest Django release?"

### Direct API Usage

```python
import json
import subprocess

# Example: Web search and chat
result = subprocess.run([
    "uv", "run", "python", "-m", "ollamamcp.server"
], input=json.dumps({
    "jsonrpc": "2.0",
    "id": 1,
    "method": "tools/call",
    "params": {
        "name": "search_and_chat",
        "arguments": {
            "query": "latest Python asyncio patterns",
            "search_results": 5
        }
    }
}), text=True, capture_output=True)

print(result.stdout)
```

## Requirements

- Python 3.12+
- Ollama (for local models)
- Internet connection (for cloud models and web search)

## Troubleshooting

### No Models Available
- Ensure Ollama is running: `ollama serve`
- Pull a local model: `ollama pull qwen3:4b`

### Web Search/Fetch Not Working
- Verify `OLLAMA_API_KEY` is set and valid
- Check internet connection

### Cloud Model Not Available  
- Verify API key has access to cloud models
- Server will automatically fall back to local models

## License

This project follows the same license as the Ollama Python library.

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation4/5

Most tools have clear boundaries: web_search finds pages, web_fetch retrieves a specific URL, chat_completion generates responses, and search_and_chat explicitly combines search and chat. The main ambiguity is between chat_completion and search_and_chat, but their descriptions make the distinction clear.

Naming Consistency3/5

Naming is readable and consistent in style (snake_case), but lacks a unifed verb_noun pattern: web_search/web_fetch share a prefixed verb, get_available_models uses get_, chat_completion is a bare noun phrase, and search_and_chat is a compound verb. This is mixed but not chaotic.

Tool Count5/5

Five tools is a well-scoped size for a search-and-chat helper. Each tool has a discernible purpose and none feels redundant.

Completeness4/5

The core workflows—web search, page fetching, listing models, plain chat, and search-grounded chat—are all covered. Missing Ollama-specific operations like embedding generation or model management, but those feel like minor gaps given the server's apparent focus.

Maintenance

ActivityInactive
ResponsivenessNo issues