MCP Dispatch Agent
README.md
# MCP Dispatch Agent
An AI assistant for a freight dispatcher, built on the Model Context Protocol.
It takes an order in natural language, picks a vehicle from the fleet, checks the
relevant regulations and prepares documents for the client and the driver.
The project has two independent halves:
- **MCP server** - tools, resources and prompts. Runs on its own; any MCP client
can connect to it.
- **Agent** - a model-to-tool loop on the Claude Messages API. Connects to the
server as an ordinary client, using the same Bearer token.
---
## Quick start
```bash
git clone https://github.com/yurii-sheremeta/mcp-dispatch-agent.git
cd mcp-dispatch-agent
cp .env.example .env # fill in your tokens and ANTHROPIC_API_KEY
```
**Local:**
```bash
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
python -m src.rag.ingest # build the document index
python -m src.server.main # server on http://localhost:8000
```
**Docker:**
```bash
docker compose up --build
```
**Check that it works:**
```bash
curl http://localhost:8000/health
```
**Talk to the agent** (separate terminal, server must be running):
```bash
python -m src.agent
```
---
## Architecture
```mermaid
flowchart TB
U([Dispatcher]) --> A
subgraph AG["Agent - model to tool loop"]
A[loop.py<br/>Claude Messages API]
G[guardrails.py<br/>context boundaries]
M[memory.py<br/>conversation history]
A --- G
A --- M
end
A -->|Bearer token, HTTP| S
subgraph SV["MCP server :8000"]
AU[auth.py<br/>Bearer and roles]
S[main.py<br/>streamable-http]
MO[monitoring.py<br/>/health /metrics]
AU --> S
S --- MO
end
S --> T1[fleet.py]
S --> T2[docs.py]
S --> T3[external.py]
T1 --> DB[(SQLite<br/>fleet, drivers,<br/>orders)]
T2 --> RAG[(BM25 index<br/>5 documents)]
T3 --> API[[ECB rates - Open-Meteo]]
S -.-> LOG[/logs/server.log/]
```
**Routing.** The model picks the source itself, based on tool descriptions and
the rules in `prompts/routing.md`:
| Question | Source |
|---|---|
| "which reefers are free" | SQLite - `list_vehicles` |
| "how long may a driver drive" | RAG - `search_regulations` |
| "what is that in dollars" | ECB rates - `currency_rate` |
| "find a truck and check the driver's hours" | SQLite plus RAG in sequence |
---
## Data sources
| Source | Contents | Module |
|---|---|---|
| **SQLite** | 10 vehicles, 8 drivers, 3 clients, orders | `src/db/` |
| **Documents (RAG)** | 5 markdown files: driving time, tachograph, contract, claims, cargo types | `src/rag/`, `data/docs/` |
| **External APIs** | ECB exchange rates, Open-Meteo weather | `src/server/tools/external.py` |
Document search uses **BM25**, not embeddings. This is a deliberate decision:
Anthropic has no embeddings endpoint, and a second API key or a ~400 MB local
model is not justified for five short documents. `src/rag/retriever.py` exposes a
narrow `search(query, k)` interface, so moving to semantic search means replacing
a single file.
---
## Tools
| Tool | Source | Access | Context mechanism demonstrated |
|---|---|---|---|
| `list_vehicles` | SQLite | guest, dispatcher | `ctx.info` |
| `plan_assignment` | SQLite | dispatcher | `report_progress`, `debug` with `extra`, `set_state` |
| `book_vehicle` | SQLite | dispatcher | `send_notification`, `disable_components`, error paths |
| `get_stats` | SQLite | guest, dispatcher | - |
| `search_regulations` | RAG | guest, dispatcher | `ctx.info` |
| `get_contract_clause` | RAG | guest, dispatcher | - |
| `list_documents` | RAG | guest, dispatcher | - |
| `currency_rate` | ECB | guest, dispatcher | external failure handling |
| `route_weather` | Open-Meteo | guest, dispatcher | external failure handling |
**Resources:** `resource://fleet`, `resource://tariffs`,
`resource://client/{registry_id}/profile`
**Prompts:** `quote_letter`, `driver_brief`
---
## Security
**Authorization.** Bearer token in the `Authorization` header. Two tokens, two
roles:
- `dispatcher` - every tool;
- `guest` - read-only; `plan_assignment` and `book_vehicle` are unavailable.
`/health` and `/metrics` are open on purpose: otherwise external monitoring could
not reach them.
**Safe context boundaries** (`src/agent/guardrails.py`, `prompts/guardrails.md`):
- Tool results are wrapped in `<tool_output trust="data">`. Text inside a
document is data, not instructions to the model.
- Instruction-override attempts ("ignore previous instructions", `<system>`) are
detected and flagged; the agent continues with the original request.
- Personal data never reaches the logs: `redact()` masks names, phone numbers and
tokens. Logs carry identifiers only (`DRV-03`, a plate number).
- Tokens in logs are masked down to the last four characters.
**Not committed to the repository:** `.env`, the database, the index and the logs
- see `.gitignore`.
---
## Logging and monitoring
Logs go to both the console and `logs/server.log` in a single format. Structured
data is passed in a separate field and rendered as JSON - readable for a human,
parsable by a machine:
```
2026-08-13 13:11:40 [INFO ] server | === MCP Dispatch Server starting ===
2026-08-13 13:11:40 [INFO ] server | Database ready | {"booked": 2, "free": 7, "service": 1}
2026-08-13 13:11:48 [WARNING] auth | AUTH FAIL | {"client": "172.17.0.1", "reason": "no Authorization header"}
2026-08-13 13:11:48 [INFO ] auth | AUTH OK | {"client": "172.17.0.1", "role": "dispatcher", "token": "***"}
2026-08-13 13:12:03 [DEBUG] fleet | Vehicle rejected | {"plate": "TR-9911", "reason": "tachograph: 1.5 h left, 4.9 h needed"}
2026-08-13 13:12:05 [INFO ] fleet | Assignment scan finished | {"chosen": "TR-4471", "price": 330, "checked": 10}
2026-08-13 13:12:19 [ERROR] fleet | Vehicle unavailable | {"plate": "TR-2280", "status": "booked"}
```
A token value is never visible in the logs: `mask()` keeps the last four
characters and `redact()` in the formatter additionally strips the `token` key
entirely.
**Endpoints:**
```bash
curl http://localhost:8000/health # {"status":"ok","uptime_seconds":312,...}
curl http://localhost:8000/metrics # request, error, 401 and tool-call counters
```
`/metrics` reports uptime, request count, rejected-authorization count, errors,
per-tool call counts, average assignment duration, memory usage and fleet state.
---
## Demonstration
```bash
python -m scripts.reset_db # restore the fleet to its initial state
python -m src.server.main # terminal 1: server
./scripts/demo.ps1 # terminal 2: curl scenario (401 / 200 / metrics / logs)
python -m scripts.client_demo # full scenario: progress, logs, RAG, booking
python -m scripts.check_roles # guest vs dispatcher permissions
python -m src.agent # interactive chat with the agent
```
`python -m scripts.capture_demo` runs the whole scenario in one pass and writes
`logs/demo_run.md` — verbatim output of every step against a live server: tests,
authorization, role separation, tool calls with progress and rejection reasons,
agent routing, metrics, and one example of each required log record type.
---
## Tests
```bash
pytest -q # 42 tests
ruff check src tests scripts # lint
```
Coverage: authorization (401 without a token, 401 with a wrong one, 200 with a
valid one, `/health` staying public), role separation, token masking, vehicle
selection logic, RAG relevance, prompt-injection detection, PII masking, and
source routing.
CI (`.github/workflows/ci.yml`) runs lint and tests, builds the Docker image,
starts the container and verifies that `/health` responds while an unauthorized
request receives a 401.
---
## Layout
```
src/
server/ MCP server: main, auth, monitoring, logging_conf
tools/ fleet - docs (RAG) - external (APIs)
agent/ loop, guardrails, memory, router
rag/ store (chunking) - retriever (BM25) - ingest (CLI)
db/ schema.sql - seed.sql - repo.py
prompts/ Prompt Book: system - routing - guardrails - templates
data/docs/ 5 documents for RAG
tests/ 42 tests
scripts/ demo.ps1 - client_demo - check_roles - reset_db
```
---
## Limitations
This is a training project, not a production system.
1. **The documents in `data/docs/` are simplified training extracts**, not
current legal instruments. Each file carries a notice in its header. Wording
must be verified against the applicable regulation before operational use.
2. **The data is fictitious:** plate numbers, registry IDs and driver names do
not correspond to real ones.
3. **Tariffs and distances are a simplified model:** a lookup table of distances
between seven cities and a linear per-kilometre rate instead of a routing
service and real pricing.
4. **Counters live in process memory.** Several replicas would need a Prometheus
exporter; the `monitoring.py` interface anticipates that.
5. **Session state is not a database.** The assignment draft lives as long as the
connection does; a confirmed booking is written to SQLite immediately.
6. **BM25 instead of embeddings** - see the Data sources section.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues