@cyanheads/libofcongress-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@cyanheads/libofcongress-mcp-serverSearch Chronicling America for newspapers about the 1918 flu"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Public Hosted Server: https://libofcongress.caseyjhand.com/mcp
Overview
Library of Congress digital collections, Chronicling America newspaper archives, and LC Subject Headings (LCSH) authority data. Search items and newspaper pages, retrieve full item metadata and OCR text, resolve LCSH subject terms, and browse curated collections from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
Tool | Description |
| Search LOC digital collections by keyword with format, date range, subject, location, and collection filters. |
| Retrieve full metadata for a specific LOC digital item — contributors, subjects, rights, formats, and resource links. |
| Search historical newspaper pages in the Chronicling America corpus with OCR excerpts. |
| Retrieve the full OCR text and metadata for a specific newspaper page. |
| Search Library of Congress Subject Headings (LCSH) by keyword. |
| List and browse LOC curated digital collections, optionally filtered by keyword. |
Resources
Resource | Description |
| LOC digital item metadata by ID — stable URI for injecting item context into agent conversations. |
All resource data is also reachable via libofcongress_get_item. Use libofcongress_search to discover item IDs first.
Related MCP server: historical-investigator-mcp
Capability reference
libofcongress_search tool
Filters: eight material formats (
photo,map,newspaper,manuscript,audio,film,book,notated-music), inclusive year range (date_start/date_end), subject heading (uselibofcongress_search_subjectsfor the exact LCSH spelling), and geographic locationcollection_slugscopes the search to one curated collection (slug fromlibofcongress_browse_collections) — mutually exclusive withformat; an unrecognized slug returnscollection_not_foundon page 1Up to 100 results per page, capped at LOC's ~100,000-item retrieval ceiling — a notice discloses how to partition by date, subject, or location to reach the rest; real results on a page beyond the reported total are always returned, never discarded
Empty results carry a
noticefield with recovery hints, echoing the applied filtersEach result carries
is_item—truefor catalog items whoseidresolves vialibofcongress_get_item,falsefor non-item results (collections, exhibit/guide pages, newspaper pages), whoseurlshould be opened instead
libofcongress_get_item tool
Returns full metadata in one call: contributors (with their roles, e.g.
Washington, George, 1732-1799 (Author)), LCSH subject headings, cataloger notes, summary, languages, locations, rights information, physical description, call number, former IDs, original/online formats, andaccess_restrictedresource_links(deduplicated from nested upstreamfiles[]arrays) carries downloadable digital file URLs (TIFF/JPEG/PDF);related_itemslists related LOC record IDs or URLs, normalized to strings whichever form LOC sends — both render in full onstructuredContentandcontent[], never truncatedAccepts multi-segment item IDs verbatim (e.g. newspaper pages
sn95047246/1935-09-05/ed-1); the returnedurlis always an absolutehttps://URLFields absent upstream are omitted rather than filled — a sparse record stays sparse
libofcongress_search_newspapers tool
OCR text excerpts (~500 chars) returned inline for relevance assessment without a second hop
Filters: keyword, inclusive date range, US state (full name), and newspaper title (partial match)
Each result carries
states— every state LOC indexes the title under, in LOC order; a title indexed against its circulation area lists severalUp to 100 results per page, capped at LOC's ~100,000-page retrieval ceiling — a notice discloses how to partition by date or state to reach the rest
Returns the
urlfield needed bylibofcongress_get_newspaper_page— do not construct these URLs manuallyOCR quality varies by digitization batch and era; 19th-century and degraded materials may contain garbled text
Empty results carry a
noticewith recovery suggestions (broaden the date range, drop the state filter, historical-OCR caveat)
libofcongress_get_newspaper_page tool
Accepts the
urlfield from alibofcongress_search_newspapersresult — validates the URL prefix before any outbound request and rejects anything else asinvalid_page_urlReturns the issue's publication metadata alongside the text —
newspaper_title,date,place_of_publication,states,edition, the page'ssequence, and the issue's page count (segment_count) — from the same single request; fields LOC doesn't send are omittedFetches JSON from the LOC text-services endpoint (
tile.loc.gov) and reads plain text from thefull_textfieldocr_available: falsewhen the page has no digitized text (image-only batch) — a data property, not an errorWhen
ocr_availableistruebut the text service returns nothing, anoticediscloses the retrieval miss, distinct from a genuinely image-only pageStrips echoed
q=params from fulltext URLs to avoidtile.loc.gov404s (a known LOC API quirk)
libofcongress_search_subjects tool
Returns standardized LCSH labels and stable LOC URIs; use the returned
labelverbatim inlibofcongress_search'ssubjectfilter — LCSH uses inverted forms ("Photography, Aerial", "World War, 1939-1945") that differ from natural languageUp to 50 results per call (default 10); for how many LOC items carry a heading, run
libofcongress_searchwith it as thesubjectfilter and readtotalDraws from the id.loc.gov suggest endpoint's full 50-candidate pool (not scaled to
limit) and filters to true LCSH headings, so a heading ranked below name-authority records isn't reported as a false emptyWhen the ranked pool — rather than a lack of coverage — yields an empty or short result, the response discloses it with a recovery hint
libofcongress_browse_collections tool
Returns collection
slug— pass it tolibofcongress_searchascollection_slugto search inside that collectionSlugs come from the collection's loc.gov route, not its title — not guessable from the display name
Optional keyword filter by collection name/description; up to 100 collections per page
Item counts are approximate and omitted when the API doesn't provide them
libofcongress://item/{+item_id} resource
Returns the same full record as
libofcongress_get_item, asapplication/jsonitem_idcomes from alibofcongress_searchresult'sidfield, or fromlibofcongress_get_itemMulti-segment newspaper IDs keep their slashes intact (e.g.
libofcongress://item/sn95047246/1935-09-05/ed-1); percent-encoded slashes (%2F) also resolve
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Library of Congress-specific:
Module-level rate-limit enforcement: 20 req/min limit; 429 responses trigger a 1-hour block, and the error's recovery hint counts down the minutes left and names the time it lifts
Configurable pacing delay (default 3100ms, ~19 req/min) applied before every outbound LOC API request
HTML-response detection guards against silent rate-limit proxy pages that return 200 with HTML
Out-of-range page handling: a page past the end of the results (HTTP 404 on any page after the first) or past the retrieval ceiling (HTTP 400) returns an empty result with a notice, not an error — both point back to page 1 for the real page count, and the ceiling notice also explains how to partition a search that matches more than LOC will page through
Transient-fault resilience: network drops and timeouts retry with backoff behind a 30s per-request timeout ceiling; the 429 rate-limit path is never retried, since a retry would deepen LOC's 1-hour block
Agent-friendly output:
Empty results always include a
noticefield with recovery hints — echoes the applied filters and suggests how to broadenPagination status on every search response (
total,page,pages,has_next), capped at LOC's ~100,000-item retrieval ceiling, with a notice disclosing how to page past itocr_availableandis_itemdiscriminator fields let callers branch on data availability without parsing textRecovery hints on every typed error contract — actionable next steps for the agent on every failure mode
Getting started
Public Hosted Instance
A public instance is available at https://libofcongress.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"libofcongress-mcp-server": {
"type": "streamable-http",
"url": "https://libofcongress.caseyjhand.com/mcp"
}
}
}Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"libofcongress-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/libofcongress-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_SESSION_MODE": "stateless",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"libofcongress-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/libofcongress-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_SESSION_MODE": "stateless",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with Docker:
{
"mcpServers": {
"libofcongress-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"-e", "MCP_SESSION_MODE=stateless",
"ghcr.io/cyanheads/libofcongress-mcp-server:latest"
]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.4.0 or higher (or Node.js v24+).
No API key required — the LOC JSON API and LC Linked Data endpoints are open. LOC recommends a descriptive
LOC_USER_AGENTfor polite access.
Installation
Clone the repository:
git clone https://github.com/cyanheads/libofcongress-mcp-server.gitNavigate into the directory:
cd libofcongress-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# edit .env if you want to set LOC_USER_AGENT or LOC_REQUEST_DELAY_MSConfiguration
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
Variable | Description | Default |
| User-Agent header sent with LOC API requests. LOC recommends a descriptive value for polite access. |
|
| Delay in milliseconds between LOC API requests to stay under the 20 req/min rate limit. |
|
| Transport: |
|
| Port for HTTP server. |
|
| Auth mode: |
|
| HTTP session mode. This server is explicitly stateless. |
|
| Log level (RFC 5424). |
|
| Directory for log files (Node.js only). |
|
| Storage backend. |
|
| Enable OpenTelemetry instrumentation (spans, metrics, completion logs). |
|
See .env.example for the full list of optional overrides.
Running the server
Local development
Build and run:
# One-time build bun run rebuild # Run the built server bun run start:stdio # or bun run start:httpRun checks and tests:
bun run devcheck # Lint, format, typecheck, security bun run test # Vitest test suite bun run lint:mcp # Validate MCP definitions against spec
Docker
docker build -t libofcongress-mcp-server .
docker run --rm -p 3010:3010 libofcongress-mcp-serverThe Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/libofcongress-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Directory | Purpose |
|
|
| Server-specific environment variable parsing ( |
| Tool definitions ( |
| Resource definitions — |
|
|
|
|
| Unit and integration tests mirroring |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor request-scoped logging,ctx.statefor tenant-scoped storageRegister new tools and resources via the arrays in
src/index.tsWrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
Apache-2.0 — see LICENSE for details.
This server cannot be deployed
Maintenance
Related MCP Connectors
Chronicling America MCP — full-text search of ~150 years of digitized U.S.
Library of Congress (loc.gov) MCP — the world's largest library.
Internet Archive (archive.org) item search & metadata MCP.
Search 14.5M Smithsonian Open Access objects, get CC0 images, find cross-collection connections.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides access to the Library of Congress (loc.gov) data, enabling AI agents to search and retrieve information from the world's largest library.8 npmMIT
- FlicenseAqualityDmaintenanceEnables searching and retrieving historical records from the Library of Congress, including newspapers, photos, maps, manuscripts, audio, and film, via the Model Context Protocol.101-
- AlicenseAqualityBmaintenanceMCP server for searching and retrieving full-text pages from the Library of Congress, including newspapers, books, and manuscripts, via the loc.gov API.3Apache 2.0
- AlicenseNot gradedqualityAmaintenanceSearch and trace US federal rules across the Federal Register (proposed/final rules and notices), the eCFR (codified, point-in-time CFR full text, locally mirrored), and Regulations.gov (rulemaking dockets and public comments) via MCP.708 npm1Apache 2.0