yerevan-landuse-mcp
by hovogrig
README.md
# Yerevan Land Use MCP Server
A streamable-HTTP MCP server over the Yerevan land-use classification: **28,832 H3 resolution-10 hexagons**
(429.6 km², ~0.0149 km² each) classified into 8 land-use classes for **2015, 2020 and 2024**.
A single shared classifier was applied to all three years, so year-to-year differences reflect
the imagery rather than a change of model. The panel is balanced — every hex exists in every
year, which is what makes the transition and change tools meaningful.
## Install
```bash
python3 -m venv --system-site-packages .venv
.venv/bin/pip install -e .
```
## Run
The server speaks **streamable HTTP**. It is a long-running process — start it, then point
clients at the URL:
```bash
./run_server.sh # http://127.0.0.1:8000/mcp
```
```
yerevan-landuse 1.0.0 | 28832 hexes, 429.552 km2, years 2015, 2020, 2024
MCP endpoint http://127.0.0.1:8000/mcp
health check http://127.0.0.1:8000/health
mode session, SSE responses
```
The dataset is loaded **before** the port opens, so a broken or missing CSV fails at startup
instead of on someone's first tool call. Startup takes ~0.7 s; tool calls are then a few
milliseconds each, and the loaded table is shared across all connected clients.
| Flag | Default | Purpose |
|---|---|---|
| `--host` | `127.0.0.1` | Bind address. `0.0.0.0` to accept remote connections — read the security note below. |
| `--port` | `8000` | Bind port |
| `--path` | `/mcp` | URL path of the MCP endpoint |
| `--json-response` | off | Reply with one JSON body instead of an SSE stream |
| `--stateless` | off | No server-side session; each request stands alone |
| `--allowed-host` / `--allowed-origin` | — | Turn DNS-rebinding protection back on for a non-loopback bind (repeatable) |
| `--transport stdio` | — | Fall back to the local subprocess transport |
Two extra HTTP endpoints come along for free: `GET /health` (liveness + hex count, for a
process supervisor or container probe) and `GET /` (points stray browsers at `/mcp`).
## Connect
```json
{
"mcpServers": {
"yerevan-landuse": {
"type": "http",
"url": "http://127.0.0.1:8000/mcp"
}
}
}
```
- **Claude Code** — `claude mcp add --transport http yerevan-landuse http://127.0.0.1:8000/mcp`, or keep the `.mcp.json` in this folder and open the project.
- **Claude Desktop** — paste the block above into `~/.config/Claude/claude_desktop_config.json`.
**The server must already be running** before a client connects — unlike stdio, nothing
launches it for you. If tools do not appear, check `curl localhost:8000/health` first.
### Calling it without an MCP client
In `--stateless --json-response` mode a tool call is a single POST, no handshake:
```bash
./run_server.sh --port 8000 --stateless --json-response &
curl -s -X POST http://127.0.0.1:8000/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-H 'MCP-Protocol-Version: 2025-06-18' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
"params":{"name":"change_summary","arguments":{}}}'
```
That mode is also the right choice behind a load balancer, or for any client that
reconnects often — there is no session to lose. The default session mode is better for a
single interactive client, since it supports streaming and resumable connections.
### Exposing it beyond localhost
Binding to `0.0.0.0` serves **every tool to anyone who can reach the port** — there is no
authentication, and Host validation is off by default for non-loopback binds (the server
prints a warning saying so). Before exposing it:
```bash
./run_server.sh --host 0.0.0.0 --allowed-host yerevan.internal:8000
```
and put a reverse proxy with TLS and auth in front. The tools are read-only — nothing can
modify the CSVs — but `export_geojson` does write files into `outputs/`, and `run_query`
will happily burn CPU on a large scan.
## Verify
```bash
.venv/bin/python tests/validate_data.py # data integrity, all 4 CSVs cross-checked
.venv/bin/python tests/smoke_test.py # every tool in-process
./run_server.sh --port 8765 &
cd tests && ../.venv/bin/python http_smoke_test.py http://127.0.0.1:8765/mcp
```
The HTTP test drives a real client over the network: it checks `/health` and `/`, lists the
tools, runs all 23 tool calls, confirms the 7 bad-input cases come back as protocol errors,
reads every resource and prompt, re-checks the arithmetic invariants over the wire, and
reconnects a second session to prove the server survives a client going away.
## Data
| File | Role |
|---|---|
| `data/yerevan_shared_model_2015_2020_2024.csv` | master wide table — `p_<Class>_<year>` and `is_<Class>_<year>` for all 3 years |
| `data/yerevan_<year>_binary_classified_shared.csv` | single-year table; identical values plus `n_labels` |
The server reads the master file; the single-year files are cross-checked by
`tests/validate_data.py` (verified identical, same hexes in the same order). Point
`YEREVAN_DATA_DIR` elsewhere to swap the data out, `YEREVAN_OUTPUT_DIR` to redirect exports.
**Classes:** Industrial, OpenZone, Parking, Residential, Road, Sport, Vegetation, Water.
**Built-up** = Residential + Industrial + Parking + Road.
### Two bases — this matters
The classifier is multi-label: each class gets its own probability, and a hex can carry
several labels or none (1,981 hexes carry no label in 2024). Every tool therefore accepts
a `basis`:
- **`dominant`** (default) — one class per hex, the highest probability. Mutually exclusive,
so counts and areas sum to the study area. Use for area statistics.
- **`labels`** — the raw binary flags. Counts overlap and do **not** sum to the study area.
Use for co-occurrence and coverage questions.
Distances, rings and compass sectors are measured from Republic Square (40.1776, 44.5126)
using a local equirectangular projection — sub-metre error at this extent, and consistent
with the existing Yerevan analysis.
## Tools
**Orientation**
| Tool | Purpose |
|---|---|
| `dataset_info` | coverage, grid, classes, extent, vocabulary every other tool expects |
| `class_composition` | hexes, km² and share per class, per year, either basis |
| `class_cooccurrence` | which labels share a hex; label-count distribution |
| `probability_stats` | probability distribution + histogram for one class/year |
**Change over time**
| Tool | Purpose |
|---|---|
| `transition_matrix` | class-to-class flows between two years, plus stability |
| `change_summary` | built-up gain/loss, vegetation lost, net change per class, core vs edge |
| `built_up_by_ring` | the centre-to-periphery gradient, per year |
| `directional_growth` | which compass sector grew, for the outer city |
**Places**
| Tool | Purpose |
|---|---|
| `hex_at_location` | lat/lng → the hexagon there, with everything known about it |
| `hex_detail` | one hexagon: geometry, per-year probabilities, labels, change path |
| `hex_neighbors` | the k-ring around a hexagon, with a composition profile |
| `area_profile` | composition and change for a circular area around a point |
**Retrieval and export**
| Tool | Purpose |
|---|---|
| `query_hexes` | filter by class, probability, location, sector and change; paged |
| `run_query` | ad-hoc pandas expression / group-by over the full table |
| `export_geojson` | write matching hexagons as real polygons to `outputs/` |
`query_hexes` always reports the total match count and area even when returning one page,
so it doubles as a counting tool — set `limit=1` when you only want the number.
**Resources:** `yerevan://dataset/info`, `yerevan://dataset/schema`, `yerevan://dataset/files`,
`yerevan://summary/{year}`.
**Prompts:** `land_use_report`, `change_analysis`, `site_briefing`.
## What the data says
Figures from `change_summary` and `class_composition`, dominant-class basis, 2015 → 2024:
- Built-up share **39.3% → 43.4%** (168.7 → 186.5 km²); 33.5 km² densified, 15.6 km²
opened up, **+17.9 km² net**.
- Residential is the biggest gainer (**+15.1 km²**); OpenZone (−13.8 km²) and Vegetation
(−8.1 km²) give up the most. 9.5 km² went directly from vegetation to built-up.
- **77.9%** of hexes kept the same dominant class across the nine years.
- Growth is at the edge, not the core: the 0–2 km core moved 78.3% → 80.9% built,
while the 6–10 km band moved 36.4% → 42.6%.
## Notes and caveats
- These are model outputs, not ground truth. Small class-to-class flows can be
classification noise rather than real change — check a flow against the stable share
before reading anything into it, and use `probability_stats` to see how confident the
model is for a given class.
- `Water` gaining 4.0 km² between 2015 and 2024 is a good example: it is more likely
seasonal imagery or spectral confusion than new water bodies.
- `run_query` evaluates pandas expressions against a read-only in-memory copy of the
table. It cannot modify the CSVs, but it is a power tool — the other tools are safer
and self-documenting.
- `export_geojson` writes only into `outputs/`; the filename is sanitised, so it cannot
escape that folder.
## Layout
```
src/yerevan_mcp/
dataset.py loading, derived geometry, analytics
query.py selection, spatial lookups, GeoJSON export
server.py MCP tools, resources, prompts, HTTP routes, CLI
tests/
validate_data.py cross-checks all 4 CSVs
smoke_test.py exercises every tool in-process
http_smoke_test.py the same suite over a real HTTP connection
data/ the 4 source CSVs
outputs/ GeoJSON exports land here
run_server.sh launcher; flags pass straight through
```
### Keeping it running
`run_server.sh` runs in the foreground. For a persistent service, wrap it in a systemd
user unit (`~/.config/systemd/user/yerevan-mcp.service`):
```ini
[Service]
ExecStart=/home/hovhannes/Desktop/Yerevan-MCP-Server/run_server.sh
Restart=on-failure
[Install]
WantedBy=default.target
```
Then `systemctl --user enable --now yerevan-mcp`.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessSyncing