LocalLLM-MCP
by amit11-ibm
README.md
# LocalLLM-MCP
A production-ready Python MCP (Model Context Protocol) server that exposes
local LLM models running under **llama.cpp** to **IBM Bob** in VS Code via
the STDIO transport.
All model configuration lives in `models.json`. Adding a new model requires
editing that file only — no Python changes needed.
---
## Table of Contents
1. [Architecture](#architecture)
2. [Prerequisites](#prerequisites)
3. [Installation](#installation)
4. [Configuration](#configuration)
- [models.json reference](#modelsjson-field-reference)
- [.env variable reference](#env-variable-reference)
5. [Adding a New Model](#adding-a-new-model)
6. [IBM Bob Integration](#ibm-bob-integration)
7. [MCP Tools Reference](#mcp-tools-reference)
8. [Supported Runtimes](#supported-runtimes)
9. [Troubleshooting](#troubleshooting)
10. [Testing](#testing)
---
## Architecture
```
<project-root>\
│
├── server.py MCP entry point — STDIO transport, tool registration
├── router.py Orchestration — validates requests, calls HTTP client
├── config.py Pydantic loader — reads models.json + .env
├── models.json All model configuration (single source of truth)
├── requirements.txt Pinned Python dependencies
├── .env.example Environment variable template
├── install.bat One-click setup (creates .venv, installs deps)
└── start_server.bat Manual launch for testing
│
└── src/
├── clients/
│ └── openai_client.py Async HTTP client — POST /v1/chat/completions
├── logging_setup/
│ └── logger.py Structured JSON logging to stderr
└── utils/
└── helpers.py retry_async, with_timeout, sanitize_model_name
```
### Data Flow
```
IBM Bob (VS Code)
│ JSON-RPC over stdin/stdout
▼
server.py ── @server.tool() handlers
│
▼
router.py ── validate model → get config → call client
│
▼
src/clients/openai_client.py
│ POST /v1/chat/completions
▼
llama-server instance (one per model, unique port)
│
▼
*.gguf model file on D:\Local-LLM\
```
**Transport:** STDIO — Bob spawns `server.py` as a child process. MCP
JSON-RPC messages travel over stdin/stdout. All application logs travel
over stderr and never corrupt the MCP stream.
---
## Prerequisites
| Requirement | Version | Notes |
|---|---|---|
| Python | 3.14+ | Must be on `PATH` |
| llama.cpp | latest | `llama-server` binary must be on `PATH` or full path used |
| IBM Bob | latest | VS Code extension installed |
| VS Code | latest | |
| Internet access | — | Required during `install.bat` only (pip download) |
---
## Installation
### Step 1 — Clone / copy the project
Place the project folder anywhere on your machine (e.g. `C:\LocalLLM-MCP\`) and
update the paths in `mcp.json` accordingly.
### Step 2 — Run the installer
Open a terminal in the project root and run:
```bat
cd C:\path\to\LocalLLM-MCP
install.bat
```
This will:
- Create `.venv\` (Python virtual environment)
- Upgrade pip
- Install all dependencies from `requirements.txt`
- Copy `.env.example` → `.env` (if `.env` does not already exist)
### Step 3 — Start llama-server instances
Each model requires its own `llama-server` process running on a unique port.
Open **five separate terminals** and run one command per terminal:
```bat
REM Granite (port 8080)
llama-server --model "<YOUR_MODELS_DIR>\IBM-Models\granite-4.1-3b\<model>.gguf" --port 8080
REM Nemotron (port 8081)
llama-server --model "<YOUR_MODELS_DIR>\Nvidia-Models\NVIDIA-Nemotron-3-Nano-4B-GGUF\<model>.gguf" --port 8081
REM Gemma (port 8082)
llama-server --model "<YOUR_MODELS_DIR>\Google-Models\googlegemma-4-E4B-it-qat-q4_0-gguf\<model>.gguf" --port 8082
REM Qwen (port 8083)
llama-server --model "<YOUR_MODELS_DIR>\Alibaba-Models\Qwen2.5-Coder-3B-Instruct-GGUF\<model>.gguf" --port 8083
REM Llama (port 8084)
llama-server --model "<YOUR_MODELS_DIR>\Meta-Models\Llama-3.2-3B-Instruct-Q4_K_M-GGUF\<model>.gguf" --port 8084
```
> Replace `<model>.gguf` with the actual filename inside each folder.
### Step 4 — Register with IBM Bob
Add the snippet from the [IBM Bob Integration](#ibm-bob-integration) section
to your `mcp.json` file, then restart Bob.
---
## Configuration
### models.json field reference
`models.json` is the single source of truth for all model configuration.
It lives in the project root and is loaded at server startup.
```json
{
"models": {
"<key>": {
"display_name": "Human-readable name shown in list_models()",
"vendor": "Model vendor / creator",
"runtime": "llama.cpp",
"model_path": "Absolute path to the model FOLDER on disk",
"endpoint": "Base URL of the running llama-server, e.g. http://localhost:8080",
"context_length": 8192,
"temperature": 0.7,
"enabled": true
}
}
}
```
| Field | Type | Required | Description |
|---|---|---|---|
| `display_name` | string | yes | Shown in `list_models()` responses |
| `vendor` | string | yes | Model creator (IBM, NVIDIA, Google, …) |
| `runtime` | string | yes | Always `"llama.cpp"` in this deployment |
| `model_path` | string | yes | Absolute path to the model folder (not the .gguf file) |
| `endpoint` | string | yes | Full base URL of the llama-server for this model |
| `context_length` | integer | yes | Max context window in tokens |
| `temperature` | float | yes | Sampling temperature (0.0 – 2.0) |
| `enabled` | boolean | yes | `false` hides the model from Bob and blocks calls to it |
### .env variable reference
Copy `.env.example` to `.env` and edit as needed. Variables set in `.env`
override `models.json` endpoint defaults. OS-level environment variables
take precedence over `.env`.
| Variable | Default | Description |
|---|---|---|
| `LOG_LEVEL` | `INFO` | Logging verbosity: `DEBUG`, `INFO`, `WARNING`, `ERROR` |
| `REQUEST_TIMEOUT_SECONDS` | `120` | Seconds before a request is abandoned |
| `MAX_RETRIES` | `3` | Retry attempts on connection errors and 5xx responses |
| `GRANITE_ENDPOINT` | *(from models.json)* | Override endpoint for the `granite` model |
| `NEMOTRON_ENDPOINT` | *(from models.json)* | Override endpoint for the `nemotron` model |
| `GEMMA_ENDPOINT` | *(from models.json)* | Override endpoint for the `gemma` model |
| `QWEN_ENDPOINT` | *(from models.json)* | Override endpoint for the `qwen` model |
| `LLAMA_ENDPOINT` | *(from models.json)* | Override endpoint for the `llama` model |
> The endpoint override pattern works for any model key:
> `<MODEL_KEY_UPPER>_ENDPOINT=http://...`
---
## Adding a New Model
No Python code changes are needed. Edit `models.json` only.
### Example — adding a new Mistral model
**1.** Add an entry to `models.json`:
```json
"mistral": {
"display_name": "Mistral 7B Instruct",
"vendor": "Mistral AI",
"runtime": "llama.cpp",
"model_path": "D:\\Local-LLM\\Mistral-Models\\Mistral-7B-Instruct-GGUF",
"endpoint": "http://localhost:8085",
"context_length": 32768,
"temperature": 0.7,
"enabled": true
}
```
**2.** Start a new llama-server instance on port 8085:
```bat
llama-server --model "D:\Local-LLM\Mistral-Models\Mistral-7B-Instruct-GGUF\mistral-7b-instruct.gguf" --port 8085
```
**3.** Restart the MCP server (Bob will restart it automatically on next use,
or restart VS Code).
**4.** Verify: ask Bob to call `list_models()` — the new model should appear.
### To disable a model temporarily
Set `"enabled": false` in `models.json` and restart the server.
The model will no longer appear in `list_models()` and calls to it
will return a clear error message.
---
## IBM Bob Integration
The server communicates with Bob via **STDIO transport** — Bob spawns
`server.py` as a child process; no port or HTTP server is needed.
### Step 1 — Locate your mcp.json
Bob supports two configuration levels:
| Level | File location | Scope |
|---|---|---|
| Global | `C:\Users\<YOUR_USERNAME>\.bob\mcp.json` | All workspaces |
| Project | `C:\path\to\LocalLLM-MCP\.bob\mcp.json` | This project only |
If the file does not exist, create it.
### Step 2 — Add the server entry
Add the following JSON to your chosen `mcp.json`:
```json
{
"mcpServers": {
"localllm-mcp": {
"command": "C:\\path\\to\\LocalLLM-MCP\\.venv\\Scripts\\python.exe",
"args": ["C:\\path\\to\\LocalLLM-MCP\\server.py"],
"cwd": "C:\\path\\to\\LocalLLM-MCP",
"env": {
"LOG_LEVEL": "INFO"
},
"alwaysAllow": ["list_models", "health_check"],
"disabled": false
}
}
}
```
> If you placed the project at a different path, update `command`, `args`,
> and `cwd` accordingly.
### Step 3 — Reload Bob
In VS Code, open the Bob panel → Settings → MCP tab, or restart VS Code.
Bob will spawn the server process automatically when a tool is first called.
### Step 4 — Verify
Ask Bob:
```
Call the health_check tool on the localllm-mcp server.
```
Expected response:
```json
{
"status": "ok",
"models_total": 5,
"models_enabled": 5,
"model_keys": ["gemma", "granite", "llama", "nemotron", "qwen"],
"runtime": "llama.cpp"
}
```
---
## MCP Tools Reference
### `ask_model`
Send a prompt to a model and receive the **complete response**.
| Parameter | Type | Description |
|---|---|---|
| `model_name` | string | Model key (e.g. `granite`, `llama`, `gemma`, `nemotron`, `qwen`) |
| `prompt` | string | The text prompt to send |
**Example:**
```
Ask granite: "Explain what the MCP protocol is in two sentences."
```
---
### `ask_model_stream`
Send a prompt and receive the response **assembled from streaming token
chunks**. Functionally identical to `ask_model` from Bob's perspective;
uses less peak memory on the server for long responses.
| Parameter | Type | Description |
|---|---|---|
| `model_name` | string | Model key |
| `prompt` | string | The text prompt to send |
---
### `list_models`
List all enabled models and their metadata. No parameters.
**Example response:**
```json
{
"models": [
{
"key": "granite",
"display_name": "IBM Granite 4.1 3B",
"vendor": "IBM",
"runtime": "llama.cpp",
"endpoint": "http://localhost:8080",
"context_length": 8192,
"temperature": 0.7
},
...
]
}
```
> `model_path` is intentionally omitted from responses.
---
### `health_check`
Report server and configuration status. No parameters. Does **not** contact
inference servers — fast in-process check only.
**Example response:**
```json
{
"status": "ok",
"models_total": 5,
"models_enabled": 5,
"model_keys": ["gemma", "granite", "llama", "nemotron", "qwen"],
"runtime": "llama.cpp"
}
```
---
## Supported Runtimes
This deployment uses **llama.cpp** for all models. The server is designed to
work with any OpenAI-compatible inference API. To switch a model to a
different runtime, update its `endpoint` in `models.json` — no Python
changes needed.
| Runtime | Default Port | OpenAI-compatible endpoint | Notes |
|---|---|---|---|
| llama.cpp | 8080–8084 | `/v1/chat/completions` | Used in this deployment |
| Ollama | 11434 | `/v1/chat/completions` | Requires `OLLAMA_ORIGINS=*` |
| LM Studio | 1234 | `/v1/chat/completions` | Enable local server in UI |
| vLLM | 8000 | `/v1/chat/completions` | `python -m vllm.entrypoints.openai.api_server` |
| NVIDIA NIM | 8000 | `/v1/chat/completions` | Docker container |
---
## Troubleshooting
### `[Error] Could not reach 'IBM Granite 4.1 3B' at http://localhost:8080`
The llama-server for that model is not running. Start it:
```bat
llama-server --model "D:\Local-LLM\IBM-Models\granite-4.1-3b\<model>.gguf" --port 8080
```
### `[Error] Unknown model 'xyz'. Available model keys: ...`
The model key you used does not match any key in `models.json`.
Use one of the listed keys. Keys are **case-insensitive**.
### `[Error] Model 'xyz' is disabled.`
The model has `"enabled": false` in `models.json`. Set it to `true` and
restart the MCP server.
### `models.json failed validation: ...`
`models.json` contains a syntax error or an invalid field value. Run the
config smoke test to see the exact error:
```bat
cd C:\LocalLLM-MCP
.venv\Scripts\python config.py
```
### Bob does not see the server / tools are not listed
1. Check that `mcp.json` exists and has valid JSON (no trailing commas).
2. Check that `disabled` is `false` in the server entry.
3. Check that the path in `command` points to the correct `.venv\Scripts\python.exe`.
4. Restart VS Code.
5. Look at the Bob output panel for MCP connection errors.
### Server starts but requests time out
- Increase `REQUEST_TIMEOUT_SECONDS` in `.env` (default: 120 seconds).
- Check that the llama-server for that model is fully loaded (watch its
terminal — it prints `llama server listening` when ready).
- Reduce `context_length` in `models.json` for the slow model.
### Port conflict — address already in use
Another process is using that port. Either:
- Stop the conflicting process: `netstat -ano | findstr :8080`
- Change the port in `models.json` and `.env.example` for that model, and
restart llama-server on the new port.
---
## Testing
### 1. Config smoke test
Verifies `models.json` loads and validates correctly:
```bat
cd C:\LocalLLM-MCP
.venv\Scripts\python config.py
```
Expected: prints all 5 models with their endpoints and exits with code 0.
### 2. Logger smoke test
Verifies structured JSON logging to stderr:
```bat
.venv\Scripts\python -m src.logging_setup.logger
```
### 3. Utilities smoke test
Verifies retry decorator, timeout, and name sanitisation:
```bat
.venv\Scripts\python -m src.utils.helpers
```
### 4. HTTP client smoke test (offline)
Verifies the client raises `ConnectionError` correctly when no server is running:
```bat
.venv\Scripts\python -m src.clients.openai_client
```
### 5. Router smoke test
Verifies validation, error handling, and config integration:
```bat
.venv\Scripts\python router.py
```
### 6. Full server tools test (manual)
Start a llama-server on port 8080, then run:
```bat
.venv\Scripts\python -c "
import asyncio
import server
async def t():
print(await server.health_check())
print(await server.list_models())
print(await server.ask_model('granite', 'Say hello in one sentence.'))
asyncio.run(t())
"
```
### 7. End-to-end via Bob
With at least one llama-server running, open Bob in VS Code and ask:
```
Use the localllm-mcp server to ask granite: "What is 2 + 2?"
```
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues