Skip to main content
Glama
README.md
# DDB MCP Server

MCP server and CLI for the [Deutsches Zeitungsportal](https://www.deutsche-digitale-bibliothek.de/newspaper), the newspaper collection of the Deutsche Digitale Bibliothek (DDB). Search the full text of ~33.8 million digitised German newspaper pages, of which roughly 27.9 million fall between 1850 and 1949.

- **search**: full Solr syntax over page OCR — exact phrases, boolean operators, wildcards, fuzzy matching and proximity — with filters for date, title, place, language and holding institution. Every hit is an individual **page**, and comes with highlighted snippets showing where the query matched.
- **facets**: list the values a filter can take, with page counts — the places, holding institutions, languages and newspaper titles the index actually holds, either across the whole corpus or within one query's results.
- **snippets**: locate a query inside one page or across every page of an issue you already have in hand.
- **get**: download the OCR text of a page or a whole issue, cached locally.

There are two ways to use it: an **MCP server** for clients that speak MCP, and a **`ddb` CLI** for agents driven through a shell. Both share one client, one cache and one set of behaviours. The CLI is what the bundled `ddb-search` skill uses.

## Installation

### Install the code

```bash
uv sync
```

### Install the CLI

```bash
uv tool install .        # puts `ddb` on your PATH
```

### Install to MCP CLIs

Installs to Claude Code, Codex CLI, and Gemini CLI:

```bash
uv run ddb-mcp-install
```

Verify the installation:

```bash
claude mcp list   # For Claude Code
codex mcp list    # For Codex CLI
gemini mcp list   # For Gemini CLI
```

## Usage

```bash
ddb search '"Luftschiffhafen Friedrichshafen"'                # 29 pages, 1909-1935
ddb search '"Luftschiffhafen Friedrichshafen"' --pages all    # sweep a bounded query
ddb search 'Zeppelin' --from-year 1900 --to-year 1910 --rows 50
ddb search 'Straßenbahn AND Unfall' --place Berlin
ddb facets place                                              # places the index holds, by page count
ddb facets place 'Zeppelin' --limit 5                         # where a term appears
ddb facets title 'Zeppelin' --from-year 1900 --to-year 1910   # which papers carried it
ddb snippets BGQICR4U4JYQ35MJ35MNBWQSJ6LA7KDR 'Zeppelin'      # where it appears in an issue
ddb get BGQICR4U4JYQ35MJ35MNBWQSJ6LA7KDR-ALTO10268886_DDB_FULLTEXT   # cached OCR text path
```

Add `--json` for machine-readable output.

**`--place` and `--provider` match a whole string, which is what `facets` is for.** They are Solr string fields, so a near-miss is not a near-miss: `--place Halle` reports `0 pages matched` with no error at all, while `--place 'Halle (Saale)'` matches 1.1 million pages. Institution names are worse, being long and official — `Bayerische Staatsbibliothek` matches 8.3 million pages and `Bayerische` matches none. `ddb facets place` and `ddb facets provider` print the values verbatim, so they can be copied across. `--title` is the forgiving one: it is analysed text with German stemming, so a fragment of a title is enough.

Given a query, `facets` describes that result set rather than the corpus, which makes it a characterisation tool as much as a lookup: the geography of a term, or the newspapers that carried it. Counts are computed by the index over the whole result set, not over the results already fetched. `facets title` lists ZDB identifiers with one recorded title form each — the title field is stemmed text and facets into word stems (`zeitung`, `nachricht`) rather than titles, so titles are counted through `zdb_id` and labelled afterwards. The label is a signpost, not a date-accurate title: the recorded string describes a paper's whole run, so a date-bounded listing can show a subtitle the paper only acquired later. The identifier beside it is the exact thing, and goes straight into `--zdb-id`.

The filters the portal itself offers are year, title, place, provider and language, and all five are here. There is no facet for publication frequency, region or state, subject, format, contributor or material type: those fields are absent from this index, not merely unexposed.

**Search already includes snippets**, which is the important workflow difference from the sibling clients. DDB returns highlighted excerpts in the search response itself, so judging a hit costs nothing beyond the search that found it. Reach for `get` only when a page or issue is worth reading at length. Use `--no-snippets` when you want a compact listing.

**There is no date ordering, deliberately.** `publication_date` is a Solr `DateRangeField` and the server refuses to sort on it, so results always come back by relevance. Rather than offer a flag that quietly reordered only the handful of results already fetched — a chronology in name while the *selection* stayed relevance-ranked — the client omits it. Chronological work means bounding the query with `--from-year`/`--to-year` and sweeping the range with `--pages all`; the ordering then falls out of the sweep.

**The result total is a true count.** Solr reports `numFoundExact`, and it survives checking: `"Luftschiffhafen Friedrichshafen"` reports 29 results and returns exactly 29 documents, spanning 1909 to 1935. This is unlike Gallica, whose totals are a ranking depth — here a total can be quoted, and a swept query really has been swept.

Downloads are cached in `$XDG_CACHE_HOME/ddb-mcp` (override with `--cache-dir` or `DDB_CACHE_DIR`). The cache location does not depend on the working directory, so the CLI can be run from anywhere.

Requests are paced one per second by default, overridable with `DDB_MIN_REQUEST_INTERVAL`. DDB publishes no rate limit for this endpoint, sends no rate-limit headers, and serves no `robots.txt` on the API host; none was observed across roughly eighty requests including a deliberate burst. **One second is therefore a conservative choice, not a measured ceiling** — there is no evidence about where the real limit sits, only that ordinary use does not come near it.

### API key

None is needed: the newspaper search index answers unauthenticated. That may be an unenforced gate rather than deliberate policy, so if you hold a DDB API key, set `DDB_API_KEY` and the client will send it — the CLI keeps working if enforcement is ever switched on.

**Where to get one:** <https://www.deutsche-digitale-bibliothek.de/user/apikey> — the key page inside your DDB account. It needs a free DDB account first, registered at <https://www.deutsche-digitale-bibliothek.de/user/register>; once logged in, the key is generated on that page and shown immediately.

It is free and needs no institutional affiliation. Per [DDB's own documentation](https://pro.deutsche-digitale-bibliothek.de/daten-nutzen/schnittstellen), *"Alle registrierten Nutzer\*innen der Deutschen Digitalen Bibliothek können sich einen Authentifikationsschlüssel für die Verwendung der APIs erzeugen lassen"* — any registered DDB user can have a key generated, from the *Meine DDB* area of their own account. There is no paid tier and no approval step.

Then:

```bash
export DDB_API_KEY=your_key_here
```

Both links live on the `www` host, which serves an anti-bot challenge to scripted clients but passes a real browser transparently — so open them in a browser, and expect `curl` to get a challenge page instead. The `apikey` URL is DDB's own, taken from the documentation page linked above; the registration path has not been walked through here.

### MCP server

Run the server directly:

```bash
uv run ddb-mcp
```

Test with MCP Inspector:

```bash
uv run fastmcp dev src/ddb_mcp/server.py
```

TDQS

A4.5/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: search_ddb for searching across the database, get_ddb_snippets for finding query matches within a specific page or issue, and download_ddb_text for downloading OCR text. No overlap in functionality.

Naming Consistency4/5

All tool names follow a consistent verb_noun pattern in snake_case. However, 'get' and 'download' are slightly different verbs for similar retrieval actions, causing minor inconsistency.

Tool Count5/5

With 3 tools covering search, snippet retrieval, and text download, the count is well-scoped for the domain of accessing digitized newspaper pages. No excess or shortage.

Completeness4/5

The tool set supports a complete workflow: search across pages, examine snippets in context, and download full text. Minor gap: no dedicated metadata retrieval tool, but metadata is included in search results.

Maintenance

ActivitySlowing
ResponsivenessNo issues