Skip to main content
Glama
john-walkoe

USPTO Patent Citation MCP Server

by john-walkoe

Citations_search_citations_minimal

Citations_search_citations_minimal
Read-only

Search USPTO patent citations using a minimal field set to surface high-volume patterns for early screening.

Instructions

Minimal citation search for discovery (90-95% context reduction).

Use for high-volume pattern discovery before detailed analysis. Essential 8 fields: application, publication, art unit, citation ID, category, tech center, date, examiner indicator.

Solr/Lucene Query Examples:

  • Field search: criteria='groupArtUnitNumber:2854'

  • Date range: criteria='officeActionDate:[2017-10-01 TO *]'

  • Boolean: criteria='citationCategoryCode:X AND techCenter:2100'

  • Wildcard: criteria='citedDocumentIdentifier:US*'

  • Combined: criteria='groupArtUnitNumber:2854 AND officeActionDate:[2023-01-01 TO 2023-12-31]'

Ultra-minimal mode: Pass custom fields list for 99% token reduction (2-3 fields only). Example: fields=['citedDocumentIdentifier', 'patentApplicationNumber'] for PFW integration.

Date handling: USPTO documents this API as office actions mailed 2017-10-01 to ~30 days ago. In practice ~44% of TC2100 records carry an earlier officeActionDate (verified against PFW document dates back to 2010-2012). Do NOT add a blanket officeActionDate:[2017-10-01 TO *] clause unless you specifically want the documented window — it discards records the index actually serves.

Lane routing — TRY BOTH: this is the ENRICHED lane (passage locations, claim mapping, quality scores, NPL flag, date filtering). For completeness-sensitive questions also run Citations_search_oa_citations_minimal (raw 892/1449 lists, statutory basis, broader applicant-IDS coverage) and union the results — neither lane is a superset of the other. See Citations_get_guidance(section='oa_citations').

CROSS-LANE JOIN KEY: every row carries referenceKey, the normalised reference identifier, and it is the ONLY correct key for unioning this lane with the OA lane. The two lanes write the same reference differently: on app 12849948 the OA parsedReferenceIdentifier reads '20060075466' while the enriched citedDocumentIdentifier reads 'US 2006/0075466 A1'. Joining those two raw fields finds zero overlap on every application; the true answer there is four references in both lanes. referenceKey is digits only (a leading US, spaces, slashes, hyphens and the kind code stripped, series markers such as RE kept), derived from publicationNumber first and citedDocumentIdentifier second, and carried on both lanes at every tier including a custom fields list.

ROWS WITH NO REFERENCE: referenceKey is null when the row carries no usable identifier, and the response envelope reports how many such rows the page holds as rows_without_reference_identifier (always present, 0 included). An absent citedDocumentIdentifier key, a null one and an empty string are ONE state, not three: a row can carry an empty publicationNumber with the citedDocumentIdentifier key missing from the JSON entirely. Measured: 2 of 5 on app 11752072, 4 of 8 on 12849948, 4 of 26 on 18407147. Those rows are real citations and must be reported as unresolved, never dropped.

IDENTIFIERS: patent_number takes either a GRANTED patent number (7-8 digits; commas, spaces and a US prefix are accepted) or an 11-digit pre-grant publication number. A granted patent number is crosswalked to its application serial with one USPTO ODP applications-search call and queried as patentApplicationNumber; an 11-digit value queries publicationNumber directly. The response reports which reading was used in patent_number_resolution {input, interpreted_as, resolved_application_number when crosswalked, source}. A number that resolves to no application is a 400 naming the accepted forms, not a zero-result. application_number remains the application serial; passing one that disagrees with the crosswalked patent number is also a 400.

Note: Returns citation metadata only. For the office action text itself, use the PFW MCP's PFW_get_oa_text / PFW_get_oa_rejections (direct, no document-bag + OCR round trip).

For complex workflows and cross-MCP integration, use Citations_get_guidance(section). Quick reference: 'fields' section for Solr syntax, 'workflows_pfw' for PFW integration.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
rowsNo
startNo
fieldsNo
art_unitNo
criteriaNo
date_endNo
date_startNo
tech_centerNo
patent_numberNo
applicant_nameNo
examiner_citedNo
application_numberNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral edge cases: documented date windows that don't match actual index contents, referenceKey normalization differences, null/missing identifier handling, 400 responses for unresolved patent numbers, and the fact that rows without references must be preserved. It also clarifies the response envelope includes rows_without_reference_identifier. These are exactly the non-obvious behaviors an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: query examples, date-window caveat, lane routing, cross-lane join key, null-row handling, identifier resolution, and cross-MCP integration. The most important purpose and usage guidance are front-loaded, and the internal headings make the density navigable. There is no filler or repeated schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 12 parameters, zero schema descriptions, and the presence of an output schema, this description is exceptionally complete: it covers query syntax, field selection, identifier resolution, error behavior, cross-tool routing, and data-quality pitfalls. The output schema already covers return shape, so the description need not restate it. An agent has enough context to call this tool correctly in a wide range of discovery workflows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the full burden, and it does add meaning for the hardest parameters: criteria (Solr/Lucene syntax with examples), fields (custom field list and token reduction), patent_number, and application_number (resolution, crosswalking, 400 behavior). It does not explicitly document rows, start, date_start, date_end, art_unit, tech_center, applicant_name, or examiner_cited, though several are inferable from examples or names. This is strong but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Minimal citation search for discovery' with an explicit 'Essential 8 fields' list. It also distinguishes itself from the OA lane by labeling this the 'ENRICHED lane' and naming the sibling tool to use for raw OA citations. The purpose is concrete and not a tautology of the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: 'Use for high-volume pattern discovery before detailed analysis.' It also provides strong when-not guidance, such as 'Do NOT add a blanket officeActionDate:[2017-10-01 TO *] clause' and 'For completeness-sensitive questions also run Citations_search_oa_citations_minimal ... and union the results.' It further routes office-action-text requests to PFW tools, making alternatives clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.