Skip to main content
Glama

libofcongress-mcp-server

Search Historical Newspapers

libofcongress_search_newspapers
Read-only

Search historical newspaper pages in the Chronicling America corpus. Returns matching pages with OCR text excerpts (~500 characters), publication title, date, the states LOC indexes the title under, and the page URL needed for libofcongress_get_newspaper_page. Filters by keyword, date range, US state, and newspaper title. The OCR excerpts are sufficient for relevance assessment — call libofcongress_get_newspaper_page with the returned url field to read the full page text. OCR quality varies: 19th-century and degraded materials may contain fragmented or garbled text.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pageNo1-indexed page number for paginating results.
limitNoResults per page. Default 25, max 100.
queryYesKeyword search across OCR text and newspaper metadata.
stateNoFilter to newspapers published in this US state. Use the full state name, lowercase (e.g., "oklahoma", "new york").
date_endNoEnd year for date filter, inclusive (e.g., 1920). Omit for no upper bound.
date_startNoStart year for date filter, inclusive (e.g., 1900). Omit for no lower bound.
newspaper_titleNoFilter to a specific newspaper by title (partial match accepted).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
pageNoCurrent 1-indexed page number.
errorNoPresent when the call failed. Absent on success.
itemsNoNewspaper page results matching the search query and filters.
pagesNoTotal retrievable pages. For result sets larger than LOC will page through (~100,000 pages) this is capped, and a notice discloses how to reach the rest (partition by date/state).
totalNoTotal number of matching newspaper pages in the result set.
noticeNoRecovery hint when results are empty or a page is out of range — echoes applied filters and suggests how to broaden. Absent on successful result pages.
has_nextNoTrue when a retrievable next page follows this one. Never promises a page past LOC's ~100,000-item retrieval ceiling.
totalCountNoTotal matching newspaper pages — mirrors output.total for agent reasoning.
effectiveQueryNoThe keyword query as submitted to the Chronicling America API, after trimming.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `rate_limit_exceeded`: LOC API rate limit exceeded; requests are blocked for approximately 1 hour. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `invalid_date_range`: date_start is later than date_end. `rate_limit_exceeded`: LOC API rate limit exceeded; requests are blocked for approximately 1 hour. Other values are possible when a failure originates below the handler."
    • changedOutput schema / properties / error / properties / data / properties / reason / examples
      Previous value: -[
      -  "rate_limit_exceeded"
      -]New value: +[
      +  "invalid_date_range",
      +  "rate_limit_exceeded"
      +]
    • removedOutput schema / properties / items / items / properties / state
      Removed value: -{
      -  "description": "State where the newspaper was published.",
      -  "type": "string"
      -}
    • addedOutput schema / properties / items / items / properties / states
      Added value: +{
      +  "description": "Every US state LOC indexes this newspaper title under (e.g., [\"georgia\", \"south carolina\"]), in LOC order. A title indexed against its circulation area lists several, so no entry is necessarily the place of publication — libofcongress_get_newspaper_page returns place_of_publication. Absent when LOC lists no state.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
  2. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, lowering the burden on the description. The description adds useful behavioral nuance beyond annotations: it discloses the ~500-character OCR excerpt length, the specific returned fields, and the caveat that 19th-century/degraded materials may contain fragmented or garbled text. This is meaningful transparency without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: purpose, return contents, filter capabilities, downstream usage, and a quality caveat. Every sentence adds useful information, with no repetition of schema details and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations cover the read-only/open-world profile, the description is complete enough for correct invocation. It explains what results look like, how to use the returned url, and warns about OCR quality, leaving no essential gap for an agent deciding to call this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all seven parameters. The description adds a high-level summary of filter dimensions ('keyword, date range, US state, and newspaper title') but does not add meaning beyond what the parameter descriptions already provide. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search historical newspaper pages in the Chronicling America corpus.' It clearly distinguishes this tool from siblings by focusing on newspapers and by naming the companion tool libofcongress_get_newspaper_page, so an agent can tell it apart from the generic libofcongress_search or libofcongress_search_subjects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: for keyword, date, state, and title filtering over newspaper pages. It also explains the relationship with libofcongress_get_newspaper_page, saying the returned url field is needed to read full page text, and that OCR excerpts are sufficient for relevance assessment. It does not explicitly exclude sibling tools, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.