Skip to main content
Glama
masterries

SIP News MCP

by masterries
README.md
# SIP News MCP Server for Luxembourg Government News

An [MCP](https://modelcontextprotocol.io) connector that lets an AI assistant
**search and read the official news of the Luxembourg government press service**
(Service information et presse, SIP) at
[`sip.gouvernement.lu`](https://sip.gouvernement.lu/de/actualites.html).

It exposes the full SIP news archive (press releases, communiqués, speeches,
state visits, ministerial news, ~2300 items back to February 2000) as clean,
structured, full-text-searchable data, in German, French or English.

## Why it exists

The SIP news page loads its list through a JavaScript component, so a naive
HTTP fetch returns an empty shell. This connector instead uses the site's own
**RSS endpoints**, which return clean, paginated, newest-first data:

| Purpose            | Endpoint                                              |
| ------------------ | ----------------------------------------------------- |
| Browse all news    | `GET /{lang}/actualites.rss?page={n}`                 |
| Full-text search   | `GET /{lang}/support/recherche.rss?q={query}&page={n}`|

Each page returns 50 items as a sliding window over the whole archive. The
connector pages through automatically, de-duplicates, decodes the
(double-encoded) entities, recovers each item's category/language/date from its
URL, and can fetch the full article body on demand.

## Tools

| Tool                   | What it does                                                       |
| ---------------------- | ----------------------------------------------------------------- |
| `search_news`          | Keyword full-text search across the whole archive (newest first). |
| `semantic_search`      | Meaning-based (vector) search; handles natural-language questions. |
| `browse_latest_news`   | The chronological news feed, newest first.                        |
| `get_article`          | Fetch the full, cleaned text of one article by URL.               |
| `list_categories`      | Reference list of category keys and interface languages.          |
| `build_semantic_index` | (Re)build the local vector index used by `semantic_search`.       |
| `semantic_index_status`| Size and coverage of the vector index.                            |

Common parameters:

- `language` — `de`, `fr` or `en` (interface language; many items are in French
  regardless, so try more than one if needed).
- `limit` — how many results to return; larger values page deeper into history.
- `since` / `until` — `YYYY-MM-DD` date bounds.
- `category` — e.g. `communiques` (press releases), `articles`, `discours`
  (speeches); see `list_categories`.

Returned fields per item: `title`, `summary`, `url`, `published` (ISO),
`published_human`, `category`, `category_label`, `content_language`, `source`.

`search_news` / `browse_latest_news` also return `count`, `complete` and an
optional `note`. `complete: false` means the scan stopped at the page cap or
time budget rather than covering the whole archive/range, so a short or empty
list there is **not** an authoritative "nothing exists" (the `note` explains how
to narrow the query). `get_article` returns the body and a `truncated` flag.

## Semantic search

`semantic_search` matches news by **meaning** rather than exact keywords, so it
answers natural-language questions ("military cooperation with Belgium",
"attacks on state IT systems") even when the wording differs from the article.

How it works:

- `build_semantic_index` pages through the archive, embeds each item
  (title + summary) with an embedding model, and stores the vectors in the
  configured vector store (**ChromaDB** by default). For NVIDIA retrieval models
  on OpenRouter, documents are embedded as `passage` and queries as `query`
  (asymmetric retrieval), which sharply improves ranking.
- `semantic_search` embeds the question and returns the nearest items by cosine
  similarity, with the same `language` / `category` / `since` / `until` filters.
  Each result carries a `similarity` score in `[0, 1]`.

Setup:

1. Get an [OpenRouter API key](https://openrouter.ai/keys) and put it in the
   server's environment as `OPENROUTER_API_KEY` (see the config below). The
   default model `nvidia/llama-nemotron-embed-vl-1b-v2:free` is free. (Or use a
   local embedder instead, see below.)
2. Build the index once (from Claude, call `build_semantic_index`, or run the
   one-liner below). A full build covers the whole archive (~2300 items, back to
   2000) in about a minute. Re-run it periodically to pick up new news.

```powershell
uv run --extra semantic python -c "import asyncio; from sip_news_mcp.semantic import SemanticIndex; print(asyncio.run(SemanticIndex().build(language='fr', max_items=3000)))"
```

Relevant environment variables:

| Variable                  | Purpose                                                |
| ------------------------- | ------------------------------------------------------ |
| `SIP_NEWS_SEMANTIC`       | `on` / `off` / `auto` (default `auto`). Turn semantic search off for a keyword-only server. |
| `OPENROUTER_API_KEY`      | Embedding API key (only for the default OpenRouter provider). |
| `SIP_NEWS_EMBED_MODEL`    | Embedding model id (default: the free Nemotron model). |
| `SIP_NEWS_EMBED_BASE_URL` | Embedding API base; point at a local server for self-hosted embeddings. |
| `SIP_NEWS_EMBED_API_KEY`  | Embedding key (falls back to `OPENROUTER_API_KEY`).    |
| `SIP_NEWS_EMBED_INPUT_TYPE` | `on` / `off` / `auto` asymmetric retrieval (default `auto`). |
| `SIP_NEWS_VECTOR_BACKEND` | `chroma` (default) or `qdrant`.                        |
| `SIP_NEWS_CHROMA_DIR`     | Chroma index location; **local disk only** (default under `%LOCALAPPDATA%`). |

See [`.env.example`](.env.example).

#### Local / self-hosted embeddings (Ollama, OpenAI-compatible)

Point `SIP_NEWS_EMBED_BASE_URL` at any OpenAI-compatible `/embeddings` server
(Ollama, vLLM, LocalAI, text-embeddings-inference, ...). No OpenRouter key is
needed. For example, with [Ollama](https://ollama.com) (`ollama pull
nomic-embed-text`):

```bash
SIP_NEWS_EMBED_BASE_URL=http://localhost:11434/v1
SIP_NEWS_EMBED_MODEL=nomic-embed-text
# no key; input_type is auto-disabled for non-OpenRouter endpoints
```

Then build the index. (Changing the embedding model/endpoint changes the vector
space, so rebuild the index with `refresh=true` when you switch.)

### Keyword-only mode (no RAG, easiest to deploy)

Semantic search is optional. With `SIP_NEWS_SEMANTIC=off` (or simply by not
installing a vector backend), the server exposes only the four keyword tools
(`search_news`, `browse_latest_news`, `get_article`, `list_categories`). No
vector database, no embeddings, no OpenRouter key, and none of the heavy
optional dependencies are needed, which makes it the simplest thing to deploy.

The vector-DB dependencies are **optional extras**, so the base install is
lightweight:

| Install                      | What you get                          |
| ---------------------------- | ------------------------------------- |
| `pip install .`              | Keyword-only (no vector deps).        |
| `pip install ".[chroma]"`    | + embedded ChromaDB backend.          |
| `pip install ".[qdrant]"`    | + Qdrant backend.                     |
| `pip install ".[semantic]"`  | + both backends.                      |

With `SIP_NEWS_SEMANTIC=auto` (the default) the server enables semantic search
only when the configured backend's library is actually installed.

## Requirements

- [uv](https://docs.astral.sh/uv/) (recommended), or Python 3.10+ with `pip`.

`uv` will download a suitable Python automatically; you do not need one
installed system-wide.

> **Note (Windows + network drives).** If you keep this project on a network /
> UNC drive, Python's Windows extensions (`pywin32`, pulled in by `mcp`)
> **cannot load their DLLs from a UNC path**, so the virtual environment must
> sit on a **local disk**. Point `UV_PROJECT_ENVIRONMENT` at a local folder (the
> project *code* can stay on the network drive; only the installed environment
> needs to be local). The commands and MCP config below set it.

## Quick start

```powershell
cd C:\path\to\SIP-MCP-Connector
$env:UV_PROJECT_ENVIRONMENT = "$env:LOCALAPPDATA\sip-news-mcp\venv"
uv sync --extra semantic --extra dev    # LOCAL env with semantic + test deps
uv run --extra dev pytest               # run the offline test suite
uv run --extra semantic sip-news-mcp    # start the server (stdio, with semantic)
# For a keyword-only server, drop the extras:  uv run sip-news-mcp
```

A quick live check without an MCP client (run in the same shell, so it reuses
the local environment set above):

```powershell
uv run python -c "import asyncio; from sip_news_mcp.client import SipNewsClient; print(asyncio.run(SipNewsClient().search('cyber', language='de', limit=3)))"
```

## Use it from Claude

### Claude Desktop

Add this to `claude_desktop_config.json`
(`%APPDATA%\Claude\claude_desktop_config.json` on Windows), then restart Claude
Desktop. See [`examples/claude_desktop_config.json`](examples/claude_desktop_config.json):

```json
{
  "mcpServers": {
    "sip-news": {
      "command": "uv",
      "args": ["--directory", "C:\\path\\to\\SIP-MCP-Connector", "run", "--extra", "semantic", "sip-news-mcp"],
      "env": {
        "UV_PROJECT_ENVIRONMENT": "C:\\sip-news-mcp\\venv",
        "OPENROUTER_API_KEY": "sk-or-v1-...your key...",
        "SIP_NEWS_CHROMA_DIR": "C:\\sip-news-mcp\\chroma"
      }
    }
  }
}
```

(`--extra semantic` and the `OPENROUTER_API_KEY` / `SIP_NEWS_CHROMA_DIR` env are
only needed for semantic search; for a keyword-only server drop them. If Claude
Desktop cannot find `uv`, use the absolute path to `uv.exe`, since it may not
inherit your shell `PATH`.)

### Claude Code (CLI)

```powershell
claude mcp add sip-news `
  --env UV_PROJECT_ENVIRONMENT="$env:LOCALAPPDATA\sip-news-mcp\venv" `
  --env OPENROUTER_API_KEY="sk-or-v1-...your key..." `
  -- uv --directory "C:\path\to\SIP-MCP-Connector" run --extra semantic sip-news-mcp
```

## Example prompts

- "Search SIP for news about *cybersécurité* in 2025 and summarise the top 5."
- "List all SIP press releases (`communiques`) since 2026-01-01."
- "Find SIP articles mentioning *armée* and open the most recent one in full."

## Notes and limits

- Search uses the SIP site's own full-text engine, which is a broad match: a
  hit may mention the term only in its body, and ranking is the site's, not
  ours. The connector returns those results faithfully. Use `get_article` to
  confirm relevance before quoting.
- Content language varies per item; the connector reports `content_language`
  per result so you can tell French items from German ones.
- Date filtering is most efficient for recent ranges (the feed is newest-first
  and stops early once it passes `since`); very old ranges page deeper and may
  hit the page cap or the ~45s time budget, in which case the result is marked
  `complete: false` with an explanatory `note`.
- Identical requests are cached in-process for 5 minutes (bounded LRU), so
  repeating the same query does not re-hit the server. A single search/browse
  call still issues up to ~40 sequential page requests, but they are made one at
  a time (each awaited before the next) under a descriptive User-Agent, which
  keeps load on the public government server modest.
- `get_article` only fetches `https` URLs on `sip.gouvernement.lu` (host is
  parsed and checked, not substring-matched), so it cannot be turned into a
  request to other hosts.
- This connector only reads public pages; it performs no writes and needs no
  credentials.

## Deployment (Docker / Kubernetes)

The server speaks two transports, chosen by `MCP_TRANSPORT`:

- `stdio` (default) for Claude Desktop / Code (local subprocess).
- `http` (streamable-http) for containers and Kubernetes, listening on
  `MCP_HOST:MCP_PORT` (default `0.0.0.0:8000`, path `/mcp`).

It also supports two vector backends via `SIP_NEWS_VECTOR_BACKEND`:

- `chroma` (embedded, default) for local use.
- `qdrant` (a shared, network Qdrant) for containers / multiple replicas, set
  with `QDRANT_URL` (and optional `QDRANT_API_KEY`).

### Docker Compose

Brings up Qdrant + the MCP server (HTTP) together:

```bash
echo "OPENROUTER_API_KEY=sk-or-v1-...your key..." > .env
docker compose up -d --build
docker compose run --rm index-build      # populate the vector index once
# MCP server: http://localhost:8000/mcp   (streamable-http)
```

### Minimal / keyword-only (no RAG)

The simplest deployment: no Qdrant, no embeddings, no API key. Just the four
keyword tools.

```bash
docker compose -f docker-compose.minimal.yml up -d --build
# MCP server on http://localhost:8000/mcp
```

The minimal image is built with no vector-DB dependencies
(`docker build --build-arg EXTRAS="" -t sip-news-mcp:minimal .`) and runs with
`SIP_NEWS_SEMANTIC=off`.

### Docker (image only)

```bash
docker build -t sip-news-mcp:latest .                 # full (RAG) image
docker run --rm -p 8000:8000 \
  -e OPENROUTER_API_KEY=sk-or-v1-... \
  -e SIP_NEWS_VECTOR_BACKEND=qdrant -e QDRANT_URL=http://host.docker.internal:6333 \
  sip-news-mcp:latest
```

### Kubernetes (Helm)

The chart in [`deploy/helm/sip-news-mcp`](deploy/helm/sip-news-mcp) deploys the
server, a bundled Qdrant (StatefulSet + PVC), a Secret for the API key, and an
optional one-shot Job that builds the index after install.

```bash
helm install sip-news deploy/helm/sip-news-mcp \
  --set openrouter.apiKey=sk-or-v1-...your key...
# or reference an existing Secret:  --set openrouter.existingSecret=my-secret
```

For a keyword-only deployment (just a Deployment + Service, no Qdrant, Secret or
Job), set `semantic.enabled=false`:

```bash
helm install sip-news deploy/helm/sip-news-mcp --set semantic.enabled=false
```

Key values (see [`values.yaml`](deploy/helm/sip-news-mcp/values.yaml)):

| Value                    | Default  | Purpose                                            |
| ------------------------ | -------- | -------------------------------------------------- |
| `semantic.enabled`       | `true`   | Set `false` for a keyword-only server (no Qdrant/key/Job). |
| `vector.backend`         | `qdrant` | `qdrant` or `chroma`.                              |
| `qdrant.enabled`         | `true`   | Deploy a bundled Qdrant; set `false` + `qdrant.url` for an external one. |
| `openrouter.apiKey`      | `""`     | API key (creates a Secret), or use `existingSecret`. |
| `indexBuild.enabled`     | `true`   | Run a post-install Job to populate the index (qdrant backend). |
| `ingress.enabled`        | `false`  | Expose via an Ingress.                             |
| `replicaCount`           | `1`      | Scale out (qdrant backend; keep `1` for chroma).   |

Reach it with `kubectl port-forward svc/sip-news-... 8000:8000`, then point an
MCP client at `http://localhost:8000/mcp`.

### Connecting an MCP client over HTTP

Claude Desktop's config is stdio-only; to use the HTTP server, configure an
MCP client that supports the streamable-http transport with URL
`http://<host>:8000/mcp`.

## Project layout

```
src/sip_news_mcp/
  client.py       HTTP client, RSS/article parsers, filtering, pagination
  semantic.py     OpenRouter embeddings + the SemanticIndex
  vectorstore.py  pluggable vector backends (ChromaDB / Qdrant)
  build_index.py  `sip-news-index` CLI (one-shot index build, used by jobs)
  server.py       FastMCP server, tool definitions, stdio/http transport
  __main__.py     `python -m sip_news_mcp`
tests/
  test_urls.py          URL / host / selector-metadata tests
  test_parsing.py       RSS feed + article parser tests
  test_client.py        SipNewsClient pagination / filtering tests
  test_semantic.py      vector layer tests (fake embedder)
  test_server_config.py semantic on/off/auto config logic
  test_server_tools.py  MCP tool registration / integration (each tool >=2x)
  conftest.py           shared fixtures; fixtures/ captured live responses
Dockerfile          container image (HTTP transport)
docker-compose.yml  Qdrant + MCP server + one-shot index build
deploy/helm/sip-news-mcp/   Kubernetes Helm chart
examples/
  claude_desktop_config.json
```

TDQS

A4.2/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: browsing chronological feed, keyword search, semantic search, fetching articles, listing categories, and index management are all well-separated. Descriptions explicitly clarify differences (e.g., search vs. semantic_search).

Naming Consistency3/5

Most tool names follow a verb_noun snake_case pattern (browse_latest_news, get_article, list_categories, search_news, build_semantic_index). However, 'semantic_index_status' and 'semantic_search' are noun phrases without a verb, breaking the pattern and introducing inconsistency.

Tool Count5/5

With 7 tools, the server is well-scoped. It covers all core functionalities for a news archive (browsing, searching, retrieving, and managing the semantic index) without unnecessary bloat.

Completeness4/5

The tool surface is nearly complete for a read-only news server: both chronological and full-text search, plus semantic search, article retrieval, and index management. A minor gap is the lack of a dedicated 'get_article_by_id' tool, but URLs serve as identifiers.