Skip to main content
Glama
cyanheads

@cyanheads/libofcongress-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://libofcongress.caseyjhand.com/mcp


Overview

Library of Congress digital collections, Chronicling America newspaper archives, and LC Subject Headings (LCSH) authority data. Search items and newspaper pages, retrieve full item metadata and OCR text, resolve LCSH subject terms, and browse curated collections from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool

Description

libofcongress_search

Search LOC digital collections by keyword with format, date range, subject, location, and collection filters.

libofcongress_get_item

Retrieve full metadata for a specific LOC digital item — contributors, subjects, rights, formats, and resource links.

libofcongress_search_newspapers

Search historical newspaper pages in the Chronicling America corpus with OCR excerpts.

libofcongress_get_newspaper_page

Retrieve the full OCR text and metadata for a specific newspaper page.

libofcongress_search_subjects

Search Library of Congress Subject Headings (LCSH) by keyword.

libofcongress_browse_collections

List and browse LOC curated digital collections, optionally filtered by keyword.

Resources

Resource

Description

libofcongress://item/{+item_id}

LOC digital item metadata by ID — stable URI for injecting item context into agent conversations.

All resource data is also reachable via libofcongress_get_item. Use libofcongress_search to discover item IDs first.

Related MCP server: historical-investigator-mcp

Capability reference

  • Filters: eight material formats (photo, map, newspaper, manuscript, audio, film, book, notated-music), inclusive year range (date_start/date_end), subject heading (use libofcongress_search_subjects for the exact LCSH spelling), and geographic location

  • collection_slug scopes the search to one curated collection (slug from libofcongress_browse_collections) — mutually exclusive with format; an unrecognized slug returns collection_not_found on page 1

  • Up to 100 results per page, capped at LOC's ~100,000-item retrieval ceiling — a notice discloses how to partition by date, subject, or location to reach the rest; real results on a page beyond the reported total are always returned, never discarded

  • Empty results carry a notice field with recovery hints, echoing the applied filters

  • Each result carries is_item — true for catalog items whose id resolves via libofcongress_get_item, false for non-item results (collections, exhibit/guide pages, newspaper pages), whose url should be opened instead


libofcongress_get_item tool

  • Returns full metadata in one call: contributors (with their roles, e.g. Washington, George, 1732-1799 (Author)), LCSH subject headings, cataloger notes, summary, languages, locations, rights information, physical description, call number, former IDs, original/online formats, and access_restricted

  • resource_links (deduplicated from nested upstream files[] arrays) carries downloadable digital file URLs (TIFF/JPEG/PDF); related_items lists related LOC record IDs or URLs, normalized to strings whichever form LOC sends — both render in full on structuredContent and content[], never truncated

  • Accepts multi-segment item IDs verbatim (e.g. newspaper pages sn95047246/1935-09-05/ed-1); the returned url is always an absolute https:// URL

  • Fields absent upstream are omitted rather than filled — a sparse record stays sparse


libofcongress_search_newspapers tool

  • OCR text excerpts (~500 chars) returned inline for relevance assessment without a second hop

  • Filters: keyword, inclusive date range, US state (full name), and newspaper title (partial match)

  • Each result carries states — every state LOC indexes the title under, in LOC order; a title indexed against its circulation area lists several

  • Up to 100 results per page, capped at LOC's ~100,000-page retrieval ceiling — a notice discloses how to partition by date or state to reach the rest

  • Returns the url field needed by libofcongress_get_newspaper_page — do not construct these URLs manually

  • OCR quality varies by digitization batch and era; 19th-century and degraded materials may contain garbled text

  • Empty results carry a notice with recovery suggestions (broaden the date range, drop the state filter, historical-OCR caveat)


libofcongress_get_newspaper_page tool

  • Accepts the url field from a libofcongress_search_newspapers result — validates the URL prefix before any outbound request and rejects anything else as invalid_page_url

  • Returns the issue's publication metadata alongside the text — newspaper_title, date, place_of_publication, states, edition, the page's sequence, and the issue's page count (segment_count) — from the same single request; fields LOC doesn't send are omitted

  • Fetches JSON from the LOC text-services endpoint (tile.loc.gov) and reads plain text from the full_text field

  • ocr_available: false when the page has no digitized text (image-only batch) — a data property, not an error

  • When ocr_available is true but the text service returns nothing, a notice discloses the retrieval miss, distinct from a genuinely image-only page

  • Strips echoed q= params from fulltext URLs to avoid tile.loc.gov 404s (a known LOC API quirk)


libofcongress_search_subjects tool

  • Returns standardized LCSH labels and stable LOC URIs; use the returned label verbatim in libofcongress_search's subject filter — LCSH uses inverted forms ("Photography, Aerial", "World War, 1939-1945") that differ from natural language

  • Up to 50 results per call (default 10); for how many LOC items carry a heading, run libofcongress_search with it as the subject filter and read total

  • Draws from the id.loc.gov suggest endpoint's full 50-candidate pool (not scaled to limit) and filters to true LCSH headings, so a heading ranked below name-authority records isn't reported as a false empty

  • When the ranked pool — rather than a lack of coverage — yields an empty or short result, the response discloses it with a recovery hint


libofcongress_browse_collections tool

  • Returns collection slug — pass it to libofcongress_search as collection_slug to search inside that collection

  • Slugs come from the collection's loc.gov route, not its title — not guessable from the display name

  • Optional keyword filter by collection name/description; up to 100 collections per page

  • Item counts are approximate and omitted when the API doesn't provide them


libofcongress://item/{+item_id} resource

  • Returns the same full record as libofcongress_get_item, as application/json

  • item_id comes from a libofcongress_search result's id field, or from libofcongress_get_item

  • Multi-segment newspaper IDs keep their slashes intact (e.g. libofcongress://item/sn95047246/1935-09-05/ed-1); percent-encoded slashes (%2F) also resolve

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Library of Congress-specific:

  • Module-level rate-limit enforcement: 20 req/min limit; 429 responses trigger a 1-hour block, and the error's recovery hint counts down the minutes left and names the time it lifts

  • Configurable pacing delay (default 3100ms, ~19 req/min) applied before every outbound LOC API request

  • HTML-response detection guards against silent rate-limit proxy pages that return 200 with HTML

  • Out-of-range page handling: a page past the end of the results (HTTP 404 on any page after the first) or past the retrieval ceiling (HTTP 400) returns an empty result with a notice, not an error — both point back to page 1 for the real page count, and the ceiling notice also explains how to partition a search that matches more than LOC will page through

  • Transient-fault resilience: network drops and timeouts retry with backoff behind a 30s per-request timeout ceiling; the 429 rate-limit path is never retried, since a retry would deepen LOC's 1-hour block

Agent-friendly output:

  • Empty results always include a notice field with recovery hints — echoes the applied filters and suggests how to broaden

  • Pagination status on every search response (total, page, pages, has_next), capped at LOC's ~100,000-item retrieval ceiling, with a notice disclosing how to page past it

  • ocr_available and is_item discriminator fields let callers branch on data availability without parsing text

  • Recovery hints on every typed error contract — actionable next steps for the agent on every failure mode

Getting started

Public Hosted Instance

A public instance is available at https://libofcongress.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "libofcongress-mcp-server": {
      "type": "streamable-http",
      "url": "https://libofcongress.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file.

{
  "mcpServers": {
    "libofcongress-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/libofcongress-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_SESSION_MODE": "stateless",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "libofcongress-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/libofcongress-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_SESSION_MODE": "stateless",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "libofcongress-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "-e", "MCP_SESSION_MODE=stateless",
        "ghcr.io/cyanheads/libofcongress-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).

  • No API key required — the LOC JSON API and LC Linked Data endpoints are open. LOC recommends a descriptive LOC_USER_AGENT for polite access.

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/libofcongress-mcp-server.git
  1. Navigate into the directory:

cd libofcongress-mcp-server
  1. Install dependencies:

bun install
  1. Configure environment:

cp .env.example .env
# edit .env if you want to set LOC_USER_AGENT or LOC_REQUEST_DELAY_MS

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts.

Variable

Description

Default

LOC_USER_AGENT

User-Agent header sent with LOC API requests. LOC recommends a descriptive value for polite access.

libofcongress-mcp-server/0.3.0

LOC_REQUEST_DELAY_MS

Delay in milliseconds between LOC API requests to stay under the 20 req/min rate limit.

3100

MCP_TRANSPORT_TYPE

Transport: stdio or http.

stdio

MCP_HTTP_PORT

Port for HTTP server.

3010

MCP_AUTH_MODE

Auth mode: none, jwt, or oauth.

none

MCP_SESSION_MODE

HTTP session mode. This server is explicitly stateless.

stateless

MCP_LOG_LEVEL

Log level (RFC 5424).

info

LOGS_DIR

Directory for log files (Node.js only).

<project-root>/logs

STORAGE_PROVIDER_TYPE

Storage backend.

in-memory

OTEL_ENABLED

Enable OpenTelemetry instrumentation (spans, metrics, completion logs).

false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t libofcongress-mcp-server .
docker run --rm -p 3010:3010 libofcongress-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/libofcongress-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory

Purpose

src/index.ts

createApp() entry point — registers tools, resource, and initializes services.

src/config

Server-specific environment variable parsing (LOC_USER_AGENT, LOC_REQUEST_DELAY_MS).

src/mcp-server/tools

Tool definitions (*.tool.ts) — six LOC tools.

src/mcp-server/resources

Resource definitions — libofcongress://item/{+item_id}.

src/services/loc-api

LocApiService wrapping www.loc.gov — search, item fetch, newspaper page, collection browser.

src/services/lc-linked-data

LcLinkedDataService wrapping id.loc.gov — LCSH subject heading suggest.

tests/

Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools and resources via the arrays in src/index.ts

  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    MCP server for searching and retrieving full-text pages from the Library of Congress, including newspapers, books, and manuscripts, via the loc.gov API.
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Search and trace US federal rules across the Federal Register (proposed/final rules and notices), the eCFR (codified, point-in-time CFR full text, locally mirrored), and Regulations.gov (rulemaking dockets and public comments) via MCP.
    708 npm
    1
    Apache 2.0