Skip to main content
Glama
john-walkoe

USPTO Patent Citation MCP Server

by john-walkoe

USPTO Patent Citation MCP Server (Enriched v3 + OA v2)

A high-performance Model Context Protocol (MCP) server providing access to two USPTO patent citation APIs β€” the Enriched Citations v3 (AI-extracted passage locations, claim mapping) and the Office Action Citations v2 (raw Form 892/1449 citation data) β€” with token-saving context reduction (90-95%), progressive disclosure workflows, interactive MCP Apps UI panels, and seamless cross-MCP integration for complete patent lifecycle analysis.

Platform Support Python FastMCP APIs MCP Apps License: MIT

Demo

A single prompt drives a full examiner citation analysis β€” Enriched Citations MCP and Patent File Wrapper MCP working together to profile examiner citation patterns across both the enriched and raw OA datasets, with ultra-minimal field selection to keep token usage lean. Ultra-minimal mode requests only the fields needed for each step, so in the video you'll see many fields showing as β€” (not requested). That's intentional, not missing data.

https://github.com/user-attachments/assets/668bd241-5d98-4932-90b3-129c28084bc6

Related MCP server: USPTO Patent MCP Server

πŸ“š Documentation

Document

Description

πŸ“₯ Installation Guide

Complete cross-platform setup with automated scripts

πŸ”‘ API Key Guide

Step-by-step guide with screenshots for obtaining USPTO and Mistral API keys

πŸ“– Usage Examples

Function examples, workflows, and integration patterns

🎯 Prompt Templates

Detailed guide to sophisticated prompt templates for citation analysis & research workflows

πŸ§ͺ Testing Guide

Test suite documentation and API key setup

πŸ”’ Security Guidelines

Comprehensive security best practices

πŸ›‘οΈ Security Scanning

Automated secret detection and prevention guide

Content Provenance

How retrieved citation text is labeled and scanned (never stripped) before reaching an AI model

βš–οΈ License

MIT License terms and conditions

⚑ Quick Start

Windows Install

Run PowerShell as Administrator, then:

# Navigate to your user profile
cd $env:USERPROFILE

# If git is installed:
git clone https://github.com/john-walkoe/uspto_enriched_citation_mcp.git
cd uspto_enriched_citation_mcp

# If git is NOT installed:
# Download and extract the repository to C:\Users\YOUR_USERNAME\uspto_enriched_citation_mcp
# Then navigate to the folder:
# cd C:\Users\YOUR_USERNAME\uspto_enriched_citation_mcp

# Run setup script (sets execution policy for this session only):
Set-ExecutionPolicy -ExecutionPolicy Unrestricted -Scope Process
.\deploy\windows_setup.ps1

# View INSTALL.md for sample script output.
# Close PowerShell Window.
# If you chose option to "configure Claude Desktop integration" then restart Claude Desktop

The PowerShell script will:

  • βœ… Check for and auto-install uv (via winget or PowerShell script)

  • βœ… Install dependencies and create executable

  • βœ… Prompt for USPTO API key (required) or Detect if you had installed the developer's other USPTO MCPs and ask if want to use existing keys from those installation.

  • πŸ”’ If entering in API keys, the script will automatically store API keys securely using Windows DPAPI encryption

  • βœ… Ask if you want Claude Desktop integration configured

  • πŸ”’ Offer secure configuration method (recommended) or traditional method (API keys in plain text in the MCP JSON file)

  • βœ… Backups and then automatically merge with existing Claude Desktop config (preserves other MCP servers)

  • βœ… Provide installation summary and next steps

Claude Desktop Configuration - Manual installs

{
  "mcpServers": {
    "uspto_enriched_citations": {
      "command": "uv",
      "args": ["--directory", "C:/Users/YOUR_USERNAME/uspto_enriched_citation_mcp",
               "run",
               "uspto_enriched_citation_mcp"
      ],
      "env": {
        "USPTO_API_KEY": "your_actual_api_key_here"
      }
    }
  }
}

For detailed installation, manual setup, and troubleshooting, see INSTALL.md

πŸ”‘ Key Benefits

  • πŸ—ΊοΈSmart Field Mapping & Selection - User-configurable field sets via YAML without code changes

  • πŸ“ŠProgressive Disclosure Workflow - Minimal discovery β†’ Balanced analysis β†’ Detailed citation examination

  • 🎯Token-Saving Context Reduction - 90-95% reduction to optimized search results

  • πŸ“‹Lucene Query Syntax Support - Full Apache Lucene Query Parser Syntax with validation

  • ✨Two Citation APIs - Enriched Citations v3 (AI-extracted) + Office Action Citations v2 (raw 892/1449)

  • πŸ–₯️MCP Apps UI - Interactive card-based panels for citation results, OA citations, and statistics

  • πŸ”—Cross-MCP Integration - Links to Patent File Wrapper, PTAB, and other USPTO MCPs

  • 🌐HTTP + stdio Dual Transport - stdio for Claude Desktop, HTTP mode for MCP Apps and reverse proxies

  • πŸ›‘οΈProduction-Ready Resilience - Structured logging, retry logic, circuit breaker, rate limiting

Workflow Design - All Performed by LLM with Minimal User Guidance

User Requests the following:

  • "Find all citations for application 18180061 and analyze the citation patterns"

  • "Show me citations in technology center 2100 for the applications this assignee owns" (applicant names are not searchable on either citation API; resolve them in the PFW MCP first, then query citations by application number)

  • "Research examiner citation patterns for art unit 2854"

  • "Analyze citation decision types for machine learning patents filed in 2023"

  • "Find citations that were CITED (not DISCARDED) in software classification 706"

LLM Performs these steps:

Step 1: Discovery minimal β†’ Step 2: Filter & Select β†’ Step 3: Analyze (optional) β†’ Step 4: Citation Details (optional) β†’ Step 5: Cross-MCP Integration (optional)

The field configuration supports an optimized research progression:

  1. Discovery minimal returns 50-100 citations efficiently with essential identifiers and decision info

  2. Filter & Select from results to identify most relevant citations for detailed analysis

  3. Analyze (optional) detailed research using balanced field set when comprehensive analysis needed

  4. Citation Details (optional) get complete citation record with optional citing context and passage analysis

  5. Cross-MCP Integration (optional) connect citations to prosecution history, PTAB challenges, or knowledge base research

🎯 Prompt Templates

The Citations MCP provides 5 guided workflow prompts accessible directly in Claude Desktop UI. These templates automate complex multi-step citation analysis workflows and eliminate the need to memorize tool syntax.

Prompts are opt-in server-side: set CITATIONS_ENABLE_PROMPTS=true to register them (default off β€” no prompts appear in the client until enabled).

For detailed prompt documentation, usage examples, and cross-MCP integration patterns, see PROMPTS.md.

Core Prompt Workflows

Prompt Name

Purpose

patent_citation_analysis

Complete citation analysis for specific patents with prosecution context

enhanced_examiner_behavior_intelligence_PFW_PTAB_FPD

Comprehensive examiner profiling with citations, petitions, PTAB correlation

litigation_citation_research_PFW_PTAB

Complete litigation research package with prosecution & PTAB analysis

technology_citation_landscape_PFW

Map prior art landscape for technology areas

art_unit_citation_assessment

Analyze art unit citation norms and examiner patterns

Key Features Across All Templates:

  • Enhanced Input Processing - Flexible identifier support (patent numbers, application numbers, title keywords)

  • Smart Validation - Automatic format detection and guidance

  • Cross-MCP Integration - Seamless workflows with PTAB, FPD, Citations, and Pinecone MCPs

πŸ“Š Available Functions (10 tools, plus 1 registration-gated admin tool)

Enriched Citations v3 Tools (AI-extracted passage locations, claim mapping)

Tool

Use Case

Requirements

Citations_search_citations_minimal

Ultra-fast citation discovery β€” 8 essential fields, 90-95% context reduction

USPTO_API_KEY

Citations_search_citations_balanced

Comprehensive citation analysis β€” 19 fields, 80-85% context reduction

USPTO_API_KEY

Citations_get_citation_details

Full single citation record by ID. ⚠️ Metadata only β€” use PFW for actual documents

USPTO_API_KEY

Citations_get_citation_statistics

Database statistics and aggregations for strategic planning

USPTO_API_KEY

Citations_get_available_fields

Discover Enriched Citations field names and query syntax (22 fields)

USPTO_API_KEY

Office Action Citations v2 Tools (raw Form 892/1449 data β€” broader coverage)

Tool

Use Case

Requirements

Citations_search_oa_citations_minimal

High-volume OA citation discovery β€” 7 key fields

USPTO_API_KEY

Citations_search_oa_citations_balanced

Detailed OA citation analysis β€” all 16 fields

USPTO_API_KEY

Citations_get_oa_citation_fields

Discover OA Citations field names and query syntax (16 fields)

USPTO_API_KEY

Utility Tools

Tool

Purpose

Requirements

Citations_validate_query

Validate Lucene syntax and get optimization suggestions

None

Citations_get_guidance

Context-efficient selective guidance sections

None

Admin Tool (OAuth deployments only)

Tool

Purpose

Requirements

citations_manage_users

Registered-user management with an MCP App panel

CITATIONS_ENABLE_USER_MANAGEMENT=true; visible only to identities with the citations:admin scope

When to use each API:

  • Enriched Citations v3: When you need passage locations, claim mapping, quality scores, nplIndicator, or officeActionDate filtering. Documented coverage starts 2017-10-01, but the index returns substantially older records in practice, so do not add a blanket date clause.

  • OA Citations v2: When you need the raw Form 892/1449 lists, legalSectionCode (102/103/112 statutory basis), actionTypeCategory, or the broader applicant-IDS inventory. It has no date field at all, and officeActionDate returns HTTP 400 here. Neither lane is a superset of the other, so run both for completeness-sensitive questions.

Detailed Citation Tier (Citations_get_citation_details): Single citation deep dive

  • Complete record: All available citation metadata with formatted presentation

  • Optional context: Include citing application details and passage-level analysis

  • Strategic insights: Decision reasoning, relevance scores, relationship mapping

LLM Guidance Function

Function (Display Name)

Purpose

Requirements

Citations_get_guidance

Context-efficient selective guidance sections

None

Context-Efficient Guidance System

Citations_get_guidance Tool - Solves MCP Resources visibility problem with selective guidance sections:

🎯 Quick Reference Chart - Know exactly which section to call:

  • πŸ” "Find citations by examiner/application/tech" β†’ fields

  • πŸ”€ "Which lane: OA citations or enriched citations?" β†’ oa_citations

  • πŸ“„ "Understand citation categories (X/Y/A + NPL via nplIndicator)" β†’ citation_codes

  • πŸ”– "Citation date coverage per lane" β†’ data_coverage

  • 🀝 "PFW workflow for office action documents" β†’ workflows_pfw

  • 🚩 "PTAB citation correlation" β†’ workflows_ptab

  • πŸ“Š "FPD petition citation patterns" β†’ workflows_fpd

  • 🏒 "Complete lifecycle analysis" β†’ workflows_complete

  • βš™οΈ "Tool guidance and parameters" β†’ tools

  • ❌ "Search errors or query issues" β†’ errors

  • πŸ’° "Reduce API costs and optimize" β†’ cost

Patent Attorney Workflows:

  • Due Diligence: Citation risk assessment and portfolio analysis

  • Prior Art Investigation: Patentability research and invalidity analysis

  • Prosecution Strategy: Examiner pattern analysis and response tactics

  • Competitive Intelligence: Technology landscape and market positioning

Cross-MCP Integration Workflows:

  • PFW + Citations: Connect citations to prosecution history for context

  • PTAB + Citations: Correlate citation outcomes with challenge success rates

  • FPD + Citations: Petition red flags and prosecution quality assessment

  • Complete Lifecycle Analysis: PFW β†’ Citations β†’ PTAB β†’ FPD integrated workflows

  • Knowledge Base Research: Integrate with Pinecone Assistant for MPEP guidance

Context Optimization Guidance:

  • Start with minimal discovery to identify key citations

  • Progress to balanced analysis only for strategically important citations

  • Use ultra-minimal mode with custom fields parameter for 99% token reduction

  • Validate complex queries before execution to avoid API waste

  • Leverage cross-MCP integration to prevent duplicate research

The tool provides specific workflows, field recommendations, API call optimization strategies, and cross-MCP integration patterns for maximum efficiency. See USAGE_EXAMPLES.md for detailed examples and integration workflows.

πŸ”§ Field Customization

User-Configurable Field Sets

The MCP server supports user-customizable field sets through YAML configuration at the project root. You can modify field sets that minimal and balanced searches bring back without changing any code!

Configuration file: field_configs.yaml (in project root)

Easy Customization Process

  1. Open field_configs.yaml in the project root directory

  2. Uncomment fields you want by removing the # symbol

  3. Save the file - changes take effect on next Claude Desktop restart

  4. Use the simplified tools with your custom field selections

Available Field Sets (Progressive Workflow)

  • citations_minimal - Ultra-minimal for citation discovery: 8 essential fields for high-volume discovery (50-100 results)

  • citations_balanced - Comprehensive citation analysis: 19 key fields for detailed citation analysis and portfolio research

Professional Field Categories Available

⚠️ API v3 Field Reality (22 total fields as of 2024-07-11)

  • Core Identifiers: citedDocumentIdentifier, patentApplicationNumber, publicationNumber, id

  • Citation Metadata: citationCategoryCode (X=anticipates/obviates alone, Y=obviates combined β€” only X and Y in dataset, per WIPO ST.14), examinerCitedReferenceIndicator, applicantCitedExaminerReferenceIndicator

  • Organizational: groupArtUnitNumber, techCenter, workGroupNumber

  • Temporal: officeActionDate, createDateTime

  • Content: passageLocationText, relatedClaimNumberText, qualitySummaryText, officeActionCategory

  • Reference Details: inventorNameText, kindCode, countryCode, nplIndicator (boolean β€” X/Y records can be NPL)

  • System: createUserIdentifier, obsoleteDocumentIdentifier

❌ Fields NOT Available (despite code references):

  • examinerNameText, firstApplicantName, decisionTypeCode, decisionTypeCodeDescriptionText

  • inventionTitle, uspcClassification, cpcClassificationBag, patentStatusCodeDescriptionText

  • For examiner analysis: Use PFW MCP β†’ get application numbers β†’ search citations

Example Customization

File: field_configs.yaml

predefined_sets:
  citations_minimal:
    description: "Essential fields for citation discovery (90-95% context reduction)"
    fields:
      # === CROSS-MCP INTEGRATION FIELDS ===
      - patentApplicationNumber               # β†’ Patent File Wrapper MCP
      - publicationNumber                     # β†’ PTAB MCP (if granted)
      - groupArtUnitNumber                   # β†’ All USPTO MCPs

      # === CITATION CORE FIELDS ===
      - citedDocumentIdentifier              # Citation reference
      - citationCategoryCode                 # X=anticipates or obviates alone, Y=obviates when combined (only X and Y appear in this dataset)
      - techCenter                          # Technology classification
      - officeActionDate                    # Temporal analysis
      - examinerCitedReferenceIndicator      # Examiner vs Applicant

πŸ”— Lucene Query Syntax Guide

Advanced Query Examples

Field-Specific Searches:

patentApplicationNumber:18180061                    # Exact application match
groupArtUnitNumber:1759                               # Art unit search
techCenter:2100                                       # Technology center match
inventorNameText:Smith*                               # Inventor name prefix wildcard

Boolean Logic:

techCenter:2100 AND groupArtUnitNumber:2854           # Boolean AND
citationCategoryCode:X OR citationCategoryCode:Y      # Only X and Y are populated in this dataset
techCenter:2100 NOT groupArtUnitNumber:1600          # Boolean NOT

Range and Wildcard:

officeActionDate:[2023-01-01 TO 2023-12-31]         # Date range
patentApplicationNumber:18*                           # Wildcard application search
citedDocumentIdentifier:US*                           # Cited document wildcard

Citation Indicators:

examinerCitedReferenceIndicator:true                  # Only examiner-cited references
nplIndicator:true                                     # Non-patent literature only (boolean field)
citationCategoryCode:X                                # X-category citations only

Complex Multi-Field:

(citationCategoryCode:X OR citationCategoryCode:Y) AND techCenter:2100
groupArtUnitNumber:[2000 TO 2999] AND examinerCitedReferenceIndicator:true
officeActionDate:[2023-01-01 TO 2024-12-31] AND nplIndicator:false

πŸ”— Cross-MCP Integration

This MCP is designed to work seamlessly with other USPTO MCPs for comprehensive patent lifecycle analysis:

MCP Server

Purpose

GitHub Repository

USPTO Patent File Wrapper (PFW)

Prosecution history & documents

uspto_pfw_mcp

USPTO Final Petition Decisions (FPD)

Petition decisions during prosecution

uspto_fpd_mcp

USPTO Patent Trial and Appeal Board (PTAB)

Post-grant challenges

uspto_ptab_mcp

Pinecone Assistant MCP

Patent law knowledge base (MPEP, examination guidance)

pinecone_assistant_mcp

Integration Overview

The USPTO Enriched Citation API v3 MCP provides AI-powered insight into patent evaluation process, revealing what prior art and references examiners consider when making decisions. When combined with other MCPs, it enables:

  • Citations β†’ PFW: Understand citation context by cross-referencing with prosecution history

  • Citations β†’ PTAB: Correlate citation patterns with post-grant challenge outcomes

  • PFW β†’ Citations: Identify citation red flags during prosecution history review

  • Citations β†’ RAG: Research MPEP guidance and examination standards before detailed analysis

  • Complete Lifecycle: PFW + Citations + PTAB + RAG for comprehensive analysis

⚠️ Document Retrieval: Citation Metadata vs. Actual Documents

CRITICAL: The Citation API returns METADATA only, NOT actual documents.

The Enriched Citation API provides AI-extracted citation data (who cited what, when, in which claims, passage locations) but does NOT provide:

  • ❌ Office action PDF documents

  • ❌ Cited patent full-text documents

  • ❌ Prosecution history documents

  • ❌ Any downloadable files

To get actual documents, use the 2-STEP PFW MCP workflow:

Step 1: Get Document List (Always Required)

# Use selective filtering to avoid context explosion
docs = PFW_get_application_documents(
    app_number='17896175',  # from citation['patentApplicationNumber']
    document_code='CTFR',   # See decoder below
    limit=20
)

Document Code Decoder (Citation-Related Documents):

  • CTNF: Non-Final Office Action (first rejection β€” where most citations appear)

  • CTFR: Final Office Action (final rejection)

  • NOA: Notice of Allowance (citation overcame or not used)

  • 892: Examiner's Search Strategy & Citations List

  • IDS: Applicant's Information Disclosure Statement

Step 2a: LLM Analysis (Extract Text for Questions)

# When user asks: "What did the examiner say about this citation?"
content = PFW_get_document_content_with_ocr(
    app_number='17896175',
    document_identifier=docs['documents'][0]['documentIdentifier']
)
# Analyze extracted text and answer user's question

Step 2b: User Download (Provide PDF Link)

# When user says: "Get me the office action" or "I want to review it"
download = PFW_get_document_download(
    app_number='17896175',
    document_identifier=docs['documents'][0]['documentIdentifier']
)
# Present as: **πŸ“ [Download Office Action]({download['proxy_download_url']})**

When to Use Each:

  • βœ… Use PFW_get_document_content_with_ocr when LLM needs to analyze content and answer questions

  • βœ… Use PFW_get_document_download when user explicitly requests document or needs proof

  • ❌ DON'T skip Step 1 - document_identifier is always required from PFW_get_application_documents

Key Integration Patterns

Cross-Referencing Fields:

  • patentApplicationNumber - Primary key linking citations to PFW prosecution history

  • publicationNumber - Patent numbers linking to PTAB post-grant challenges

  • groupArtUnitNumber - Art unit analysis across all MCPs

  • inventorNameText - Inventor matching across databases

⚠️ Examiner Analysis Requires Two-Step Workflow (Ultra-Minimal Mode):

Critical: Use wildcard-first strategy with custom fields for 99% token reduction!

  1. PFW MCP - Get application numbers with ultra-minimal fields:

    # βœ… CORRECT: Use _minimal tool with custom fields parameter
    pfw_apps = PFW_search_applications_minimal(
        query='examinerNameText:SMITH* AND filingDate:[2015-01-01 TO *]',
        fields=['applicationNumberText', 'applicationMetaData.examinerNameText'],
        limit=50
    )
    # Result: ~5KB for 50 apps (vs ~25KB preset minimal, ~500KB full data)
    
    # ❌ WRONG: Don't use convenience parameters (exact match often fails)
    # pfw_apps = PFW_search_applications_minimal(examiner_name='SMITH, JOHN')
    
    # ❌ WRONG: Don't use short field names (causes API errors)
    # fields=['applicationNumber', 'examinerName']  # Missing 'applicationMetaData.' prefix
  2. Citation MCP - Search citations by patentApplicationNumber:

    for app in pfw_apps[:20]:  # Limit to 20 to prevent token explosion
        citations = Citations_search_citations_minimal(
            criteria=f'patentApplicationNumber:{app.applicationNumberText}',
            rows=50
        )
  3. Fields like examinerNameText, firstApplicantName, decisionTypeCode are NOT in the Citation API

Progressive Workflow:

  1. Citation Discovery (Citations): Find relevant citation patterns with minimal search

  2. Prosecution Context (PFW): Cross-reference citing applications with prosecution history using ultra-minimal fields

  3. Challenge Assessment (PTAB): Check if cited patents faced post-grant challenges

  4. Knowledge Research (RAG): Research MPEP guidance and examination standards

  5. Strategic Analysis: Combine insights across multiple data sources for informed decision-making

Cost Optimization with Cross-MCP (Token Efficiency):

Integration Pattern

Old Approach

Ultra-Minimal Approach

Token Savings

Examiner Analysis

Preset minimal (15 fields)

Custom fields (2-3 fields)

80-87% reduction

50 PFW Apps

~25KB

~5KB

80% reduction

Art Unit Analysis

Exact match + 15 fields

Wildcard + 2 fields

87% + higher hit rate

Best Practices:

  • βœ… Use PFW_search_applications_minimal WITH custom fields parameter

  • βœ… Use wildcard-first strategy: examinerNameText:SMITH* (not exact match)

  • βœ… Use FULL field paths: applicationMetaData.examinerNameText (not short names)

  • βœ… Filter by date in query: filingDate:[2015-01-01 TO *] (citation-eligible apps only)

  • βœ… Limit cross-MCP analysis to top 20 results (prevents token explosion)

  • Use citations to identify relevant prosecution applications before PFW queries

  • Research PTAB outcomes for patents with specific citation patterns

  • Validate examination practices against MPEP guidance before expensive document extraction

For detailed integration workflows, cross-referencing examples, and complete use cases, see USAGE_EXAMPLES.md.

🌐 HTTP Mode & MCP Apps

This server supports dual transport: stdio (default, for Claude Desktop) and HTTP (for MCP Apps and browser-based clients).

Starting in HTTP Mode

FASTMCP_TRANSPORT=http uv run uspto-enriched-citation-mcp
# Server starts at http://localhost:8000/mcp

Env Var

Default

Description

FASTMCP_TRANSPORT

stdio

stdio for Claude Desktop, http for HTTP transport

FASTMCP_PORT

8000

HTTP port

FASTMCP_HOST

0.0.0.0

HTTP bind address

FASTMCP_STATELESS_HTTP

true

Stateless streamable HTTP (no server-side session table)

CORS_EXTRA_ORIGIN

(none)

Additional CORS origin for reverse proxy deployments

MCP_APP_EXTRA_DOMAINS

(none)

Comma-separated additional domains added to the MCP Apps Content-Security-Policy (e.g. https://your-proxy.example.com). Needed when the client loads the iframe through a reverse proxy or Docker host.

INTERNAL_AUTH_SECRET

(none)

Shared secret for endpoint authentication (x-api-key header). Required when FASTMCP_TRANSPORT=http and CITATIONS_AUTH_MODE is not oauth: the server logs an error and exits rather than serve an unauthenticated HTTP surface. Requests without the matching header are rejected with 401. Inject via reverse proxy so MCP clients do not need to configure it manually.

CITATIONS_AUTH_MODE

none

Set to oauth to enable OAuth 2.1 sign-in (HTTP mode only); default behavior unchanged. See docs/SSO_SETUP.md.

CITATIONS_ENABLE_USER_MANAGEMENT

false

Set to true to register the citations_manage_users admin tool (required for OAuth deployments).

LOG_LEVEL

INFO

Logging verbosity (applies to stdio and HTTP transports).

LOG_DIR

(auto)

Directory for rotating application/security log files. Default: /var/log/uspto_mcp if writable, else ~/.uspto_mcp/logs.

MCP Apps UI Panels

MCP Apps panels render in both stdio and HTTP mode. stdio is sufficient for Claude Desktop β€” HTTP mode is only needed for browser-based clients (e.g. basic-host test harness).

When run via an MCP Apps-capable client, three card-based UI panels render automatically:

Tool(s)

View

What You See

Citations_search_citations_minimal, Citations_search_citations_balanced

Citation Results

Color-coded citation cards (X=red, Y=orange, A=green), examiner/applicant badges, "Open in Patent Center" (citing application) and "View cited patent or application on Google Patents" links, pipe-separated passage locations

Citations_search_oa_citations_minimal, Citations_search_oa_citations_balanced

OA Citations

Office Action citation cards with 892/1449 source badges, legal section code badges, "Open in Patent Center" and "View cited patent or application on Google Patents" links

Citations_get_citation_statistics

Statistics

Summary stat cards and horizontal bar chart breakdowns

Testing MCP Apps (basic-host)

# Terminal 1 β€” start server in HTTP mode
FASTMCP_TRANSPORT=http uv run uspto-enriched-citation-mcp

# Terminal 2 β€” start basic-host
cd ~/ext-apps/examples/basic-host
SERVERS='["http://localhost:8000/mcp"]' npm start
# Open http://localhost:8080, run any search tool, card UI renders in the panel

πŸ› οΈ Installation & Setup

Prerequisites

  • Git installed or the source files

  • uv Package Manager - Handles Python installation automatically. If using Quick Start Windows PowerShell install, uv will be installed automatically if not present.

  • USPTO API Key (required, free from USPTO Open Data Portal) - See API Key Guide for step-by-step instructions with screenshots

  • Claude Desktop (for MCP integration)

Installation

See INSTALL.md for complete cross-platform installation guide.

Docker Deployment

A Dockerfile is included at the repo root for containerized deployments. For an all-in-one stack running all four USPTO MCPs (Citations on port 8000, PFW on 8001, PTAB on 8002, FPD on 8003) see the companion repo uspto_docker_mcp, which provides a single docker compose up entry point with shared volumes for logs and databases.

Claude Desktop Windows Configuration

For uv installations, use this config:

{
  "mcpServers": {
    "uspto_enriched_citation": {
      "command": "uv",
      "args": [
        "--directory",
        "C:/Users/YOUR_USERNAME/uspto_enriched_citation_mcp",
        "run",
        "uspto-enriched-citation-mcp"
      ],
      "env": {
        "USPTO_API_KEY": "your_actual_USPTO_api_key_here"
      }
    }
  }
}

Important Notes:

  • Replace YOUR_USERNAME with your actual username

  • Replace your_actual_USPTO_api_key_here with your real USPTO API key

  • No .env files needed - Configuration handled entirely through Claude Desktop environment variables

  • Follows the same pattern as other USPTO MCPs (uspto_fpd, uspto_pfw, uspto_ptab)

  • For testing scripts, you'll still need the environment variable set

  • See INSTALL.md for additional configuration options

πŸ“ˆ Performance Comparison

Method

Response Size

Context Usage

Features

Direct curl

~50KB+

High

Raw API access with full Lucene syntax support

MCP Balanced

~8KB

Medium

Key fields + classification + cross-reference capability

MCP Minimal

~2KB

Very Low

Essential identifiers + decision types only

πŸ§ͺ Testing

tests/TEST_SUITE.md contains 32 end-to-end tests with known-good inputs and verified expected outputs for every tool. Run these in Claude Desktop to confirm the MCP is working correctly against the live USPTO API.

This is especially useful because the USPTO enriched citations API has a data coverage cutoff and not all field values are populated in practice β€” hunting for data that fits can be frustrating. The test suite uses pre-verified records and queries so you get reliable results immediately.

To run: Open Claude Desktop, paste this prompt, then append the tests you want to run:

"Please perform these MCP tests in order. For each test, call the tool with the parameters shown and tell me whether the result matches the expected output. Report PASS, PARTIAL, or FAIL for each."

Last validated: 2026-07-09 (STDIO via Claude Code; MCP App iframe rendering requires Claude Desktop). Prior full run 2026-03-28, 29/29 PASS in both transports.

Automated Tests β€” pytest

With uv (Recommended):

# Test core functionality
uv run python tests/test_basic.py

# Test with pytest
uv run pytest

Notes:

  • uv run pytest is safe to run as-is: tests/test_unified_key_management.py is excluded via addopts in pyproject.toml because it overwrites the real DPAPI-stored API keys (opt back in with CITATIONS_RUN_KEY_TESTS=1 and an explicit path).

  • Integration tests (tests/test_integration.py) require a USPTO_API_KEY and skip without one.

With traditional Python:

python tests/test_basic.py
pytest

Expected Outputs

test_basic.py:

βœ… ALL TESTS PASSED - USPTO Enriched Citation MCP working correctly!

See tests/README.md for comprehensive testing guide.

πŸ“ Project Structure

uspto_enriched_citation_mcp/
β”œβ”€β”€ field_configs.yaml             # Root-level field customization (minimal + balanced sets)
β”œβ”€β”€ .gitignore                     # Git ignore patterns
β”œβ”€β”€ .pre-commit-config.yaml        # Pre-commit hooks configuration
β”œβ”€β”€ .secrets.baseline              # Secret scanning baseline
β”œβ”€β”€ .security/                     # Security scanning tools
β”‚   β”œβ”€β”€ check_prompt_injections.py # Prompt injection scanner with baseline system
β”‚   └── prompt_injection_detector.py # Detector patterns
β”œβ”€β”€ src/
β”‚   └── uspto_enriched_citation_mcp/
β”‚       β”œβ”€β”€ __init__.py            # Package initialization
β”‚       β”œβ”€β”€ __main__.py            # Entry point for -m execution
β”‚       β”œβ”€β”€ main.py                # Composition root: FastMCP 4 server, MCP Apps resources, tool registration
β”‚       β”œβ”€β”€ runtime.py             # Service singletons + initialize_services()
β”‚       β”œβ”€β”€ server_bootstrap.py    # Transport startup (stdio / HTTP)
β”‚       β”œβ”€β”€ middleware.py          # HTTP middleware (auth header, size limits, security headers)
β”‚       β”œβ”€β”€ app_uris.py            # MCP Apps resource URIs
β”‚       β”œβ”€β”€ fastmcp_compat.py      # FastMCP 4 / MCP SDK 2.x compat shim (keeps defer_loading on the wire)
β”‚       β”œβ”€β”€ shared_secure_storage.py # Cross-MCP API key storage (Windows DPAPI / Linux 600)
β”‚       β”œβ”€β”€ api/                   # API client modules
β”‚       β”‚   β”œβ”€β”€ base_citation_client.py # Shared transport, circuit breaker, retry, caching
β”‚       β”‚   β”œβ”€β”€ enriched_client.py # Enriched Citations v3 (api.uspto.gov, X-API-KEY auth)
β”‚       β”‚   β”œβ”€β”€ oa_citations_client.py # Office Action Citations v2 (api.uspto.gov)
β”‚       β”‚   β”œβ”€β”€ applications_client.py # ODP applications search (granted-patent crosswalk)
β”‚       β”‚   └── field_constants.py # Field name constants
β”‚       β”œβ”€β”€ auth/                  # Optional OAuth 2.1 sign-in (CITATIONS_AUTH_MODE=oauth)
β”‚       β”‚   β”œβ”€β”€ provider.py        # Dual-IdP (Google + Entra ID) authorization server
β”‚       β”‚   β”œβ”€β”€ store.py           # SQLite mcp_users store
β”‚       β”‚   └── pages.py           # Sign-in chooser pages
β”‚       β”œβ”€β”€ config/                # Configuration modules
β”‚       β”‚   β”œβ”€β”€ settings.py        # Environment configuration (incl. HTTP/CORS settings)
β”‚       β”‚   β”œβ”€β”€ constants.py       # API endpoint paths and defaults
β”‚       β”‚   β”œβ”€β”€ field_manager.py   # YAML-based field configuration
β”‚       β”‚   └── tool_reflections.py # LLM guidance (loads reference/tool_guidance.md)
β”‚       β”œβ”€β”€ services/              # Business logic layer
β”‚       β”‚   β”œβ”€β”€ citation_service.py # Enriched Citations operations
β”‚       β”‚   └── oa_citation_service.py # OA Citations operations
β”‚       β”œβ”€β”€ tools/                 # MCP tool implementations
β”‚       β”‚   β”œβ”€β”€ search.py          # Citations_search_citations_minimal / Citations_search_citations_balanced
β”‚       β”‚   β”œβ”€β”€ details.py         # Citations_get_citation_details
β”‚       β”‚   β”œβ”€β”€ oa.py              # search_oa_citations_* + Citations_get_oa_citation_fields
β”‚       β”‚   β”œβ”€β”€ statistics.py      # Citations_get_citation_statistics
β”‚       β”‚   β”œβ”€β”€ utility.py         # Citations_validate_query, Citations_get_available_fields, Citations_get_guidance
β”‚       β”‚   β”œβ”€β”€ admin.py           # citations_manage_users (registration-gated)
β”‚       β”‚   └── _shared.py         # query_info envelope helper
β”‚       β”œβ”€β”€ ui/                    # MCP Apps HTML views (stdio + HTTP)
β”‚       β”‚   └── views/
β”‚       β”‚       β”œβ”€β”€ citation_results_view.py  # Enriched citation card UI (filter pills)
β”‚       β”‚       β”œβ”€β”€ oa_citations_view.py      # OA citation card UI (filter pills)
β”‚       β”‚       β”œβ”€β”€ _common.py                # Shared view scaffolding
β”‚       β”‚       β”œβ”€β”€ statistics_view.py        # Statistics summary + bar chart
β”‚       β”‚       └── user_management_view.py   # Admin user-management panel
β”‚       β”œβ”€β”€ prompts/               # Multi-step analysis workflow templates (registered only when CITATIONS_ENABLE_PROMPTS=true)
β”‚       β”‚   β”œβ”€β”€ templates/         # Markdown bodies for the prompt templates
β”‚       β”‚   β”œβ”€β”€ patent_citation_analysis.py
β”‚       β”‚   β”œβ”€β”€ enhanced_examiner_behavior_intelligence_PFW_PTAB_FPD.py
β”‚       β”‚   β”œβ”€β”€ litigation_citation_research_PFW_PTAB.py
β”‚       β”‚   β”œβ”€β”€ technology_citation_landscape_PFW.py
β”‚       β”‚   └── art_unit_citation_assessment.py
β”‚       β”œβ”€β”€ shared/                # Shared utilities
β”‚       β”‚   β”œβ”€β”€ circuit_breaker.py # Circuit breaker pattern
β”‚       β”‚   β”œβ”€β”€ injection_scan.py  # Detection-only runtime injection scanner + provenance note
β”‚       β”‚   β”œβ”€β”€ uspto_shared_rate_limiter.py # Cross-process USPTO rate limiter (multi-MCP hosts)
β”‚       β”‚   β”œβ”€β”€ pfw_link.py        # PFW hand-off hint attached to OA responses
β”‚       β”‚   β”œβ”€β”€ error_utils.py     # Standardized error handling
β”‚       β”‚   β”œβ”€β”€ exceptions.py      # Custom exception classes
β”‚       β”‚   └── enums.py           # Enum definitions
β”‚       └── util/                  # Utility modules
β”‚           β”œβ”€β”€ query_builder.py   # QueryParameters, build_query, validate_string_param
β”‚           β”œβ”€β”€ patent_crosswalk.py # patent_number normalization + granted-patent resolution
β”‚           β”œβ”€β”€ query_validator.py # Lucene syntax validation
β”‚           β”œβ”€β”€ rate_limiter.py    # Token bucket rate limiting
β”‚           β”œβ”€β”€ retry.py           # Exponential backoff retry logic
β”‚           β”œβ”€β”€ cache.py           # LRU/TTL caching for API responses
β”‚           β”œβ”€β”€ logging.py         # Rotating file logger setup + SanitizingFilter
β”‚           β”œβ”€β”€ metrics.py         # Request metrics collection
β”‚           β”œβ”€β”€ security_logger.py # Security event logging (query fingerprints, never raw text)
β”‚           └── request_context.py # Request ID tracking
β”œβ”€β”€ reference/                     # API reference documents (gitignored user copies)
β”‚   └── tool_guidance.md           # Static LLM guidance content (loaded by tool_reflections.py)
β”œβ”€β”€ audits/                        # Code review audit reports
β”œβ”€β”€ deploy/                        # Deployment scripts
β”‚   β”œβ”€β”€ linux_setup.sh             # Linux setup (uv + shared key detection)
β”‚   β”œβ”€β”€ windows_setup.ps1          # Windows setup (DPAPI + Claude Desktop config)
β”‚   β”œβ”€β”€ manage_api_keys.ps1        # Windows key management utility
β”‚   β”œβ”€β”€ Validation-Helpers.psm1    # PowerShell key validators
β”‚   └── validation-helpers.sh      # Bash key validators + cross-MCP key detection
β”œβ”€β”€ docs/                          # Additional documentation
β”‚   β”œβ”€β”€ CONTENT_PROVENANCE.md      # Retrieved-text handling / injection-annotation posture
β”‚   β”œβ”€β”€ SSO_SETUP.md               # OAuth 2.1 sign-in setup (Google + Microsoft)
β”‚   └── graceful-degradation.md    # Circuit-breaker fallback + stale-cache design
β”œβ”€β”€ tests/                         # Test suite (uv run pytest; see Testing section)
β”‚   β”œβ”€β”€ README.md                  # Test suite documentation
β”‚   β”œβ”€β”€ TEST_SUITE.md              # Manual test cases (32 tests, STDIO mode)
β”‚   β”œβ”€β”€ conftest.py                # Shared mock_runtime fixture (mocked clients, real services)
β”‚   β”œβ”€β”€ test_basic.py              # Core functionality (no API key required)
β”‚   β”œβ”€β”€ test_auth_provider.py      # OAuth provider + SQLite store + admin gating
β”‚   β”œβ”€β”€ test_convenience_parameters.py # Convenience param β†’ query builder tests
β”‚   β”œβ”€β”€ test_field_configuration.py    # Field management tests
β”‚   β”œβ”€β”€ test_injection_scan.py         # Runtime injection scanner + envelope wiring tests
β”‚   β”œβ”€β”€ test_integration.py            # Integration tests (API key required)
β”‚   β”œβ”€β”€ test_oa_citations_client.py    # OA Citations client tests
β”‚   β”œβ”€β”€ test_oa_citation_service.py    # OA Citations service tests
β”‚   β”œβ”€β”€ test_resilience.py             # Circuit breaker and rate limiting tests
β”‚   β”œβ”€β”€ test_security.py               # Injection detection + input validation tests
β”‚   β”œβ”€β”€ test_shared_rate_limiter.py    # Cross-process shared rate limiter tests
β”‚   β”œβ”€β”€ test_statistics_tool.py        # Citation statistics tests
β”‚   β”œβ”€β”€ test_patent_crosswalk.py       # Granted-patent-number crosswalk (normalizer, client, tool wiring)
β”‚   β”œβ”€β”€ test_unified_key_management.py # API key storage tests (excluded by default, see below)
β”‚   └── ...                            # Additional tool/logging/query-validation tests
β”œβ”€β”€ scripts/                       # Operator utilities
β”‚   β”œβ”€β”€ manage_mcp_users.py        # Bootstrap and manage the OAuth mcp_users table
β”‚   └── rotate_internal_auth_secret.py # INTERNAL_AUTH_SECRET rotation with an overlap window
β”œβ”€β”€ Dockerfile                     # Container image (HTTP transport)
β”œβ”€β”€ docker-compose.yml             # Compose stack for the HTTP deployment
β”œβ”€β”€ .env.example                   # Environment template for the container
β”œβ”€β”€ pyproject.toml                 # Package configuration
β”œβ”€β”€ uv.lock                        # Dependency lockfile
β”œβ”€β”€ README.md                      # This file
β”œβ”€β”€ INSTALL.md                     # Installation guide
β”œβ”€β”€ USAGE_EXAMPLES.md              # Function examples and workflows
β”œβ”€β”€ PROMPTS.md                     # Prompt template documentation
β”œβ”€β”€ API_KEY_GUIDE.md               # API key setup guide with screenshots
β”œβ”€β”€ SECURITY_GUIDELINES.md         # Security best practices
└── SECURITY_SCANNING.md           # Secret detection and prevention guide

πŸ” Troubleshooting

Common Issues

API Key Issues

  • For Claude Desktop: API keys in config file are sufficient

  • For test scripts: Environment variables must be set

Setting USPTO API Key for Testing:

  • Windows Command Prompt: set USPTO_API_KEY=your_key

  • Windows PowerShell: $env:USPTO_API_KEY="your_key"

  • Linux/macOS: export USPTO_API_KEY=your_key

uv vs pip Issues

  • uv advantages: Better dependency resolution, faster installs

  • Mixed installation: Can use both uv sync and pip install -e .

  • Testing: Use uv run prefix for uv-managed projects

Lucene Query Issues

  • Cause: Invalid Lucene syntax or field names

  • Solution: Use Citations_validate_query tool to check syntax and get suggestions

  • Common: Missing quotes, unbalanced parentheses, wrong field names

Fields Not Returning Data

  • Cause: Field name not in API or configuration

  • Solution: Use Citations_get_available_fields to discover correct field names

Authentication Errors

  • Cause: Missing or invalid API key

  • Solution: Verify USPTO_API_KEY environment variable or Claude Desktop config

  • API Key Source: Get free API key from USPTO Open Data Portal

MCP Server Won't Start

  • Cause: Missing dependencies or incorrect paths

  • Solution: Re-run setup script, restart Claude Desktop and verify configuration

Getting Help

  1. Check the test scripts for working examples

  2. Review the field configuration in field_configs.yaml

  3. Verify your Claude Desktop configuration matches the provided templates

  4. Use Citations_get_guidance for workflow-specific guidance

πŸ›‘οΈ Security & Production Readiness

Enhanced Error Handling

  • Structured logging - Request ID tracking for better debugging and monitoring

  • CWE-532 sanitized logging - All loggers use get_logger() from util/logging.py, which applies a SanitizingFilter to redact API keys, tokens, and other sensitive values from log output before they reach any handler

  • Request timeout handling - Configurable timeouts for API reliability

  • Production-grade responses - Clean error messages without internal system details

  • Environment validation - API key format and presence checking

  • Input sanitization - Safe handling of user queries and parameters

Content Provenance & Injection Annotation

Free-text citation fields (passageLocationText, qualitySummaryText) are AI-extracted by the USPTO from office-action documents, which quote arbitrary applicant- and examiner-drafted text. The server treats that text as data, not instructions β€” and never strips or rewrites it (verbatim fidelity is the product):

  • Every text-bearing tool (Citations_search_citations_minimal/_balanced, Citations_get_citation_details, Citations_search_oa_citations_minimal/_balanced) attaches a provenance_note envelope field labeling retrieved text as quoted document content.

  • A detection-only runtime scanner (src/uspto_enriched_citation_mcp/shared/injection_scan.py) checks the free-text fields for instruction-override, prompt-extraction, and encoding-evasion language plus invisible-Unicode steganography. On a hit it attaches an injection_scan envelope key naming the flagged result with kind labels only β€” never the matched text; the key is absent entirely when results are clean.

  • The server instructions state the same posture so consuming models report instruction-like language found in retrieved text instead of acting on it.

Full write-up: docs/CONTENT_PROVENANCE.md. This runtime layer is complementary to the commit-time .security/ codebase scanner described below.

Security Features

  • πŸ”’ Secure API Key Storage (Windows DPAPI) - Encrypted API key storage using Windows Data Protection API

    • Per-user, per-machine encryption (only you on your machine can decrypt)

    • No plain-text API keys in configuration files

    • Automatic fallback to environment variables on non-Windows systems

    • Zero external dependencies (uses Python ctypes)

  • Prompt Injection Detection with Baseline System - Tracks known findings and only flags NEW patterns

    • Baseline file: .prompt_injections.baseline with SHA256 fingerprinting

    • Exit codes: 0 (no NEW findings), 1 (NEW findings detected), 2 (error)

  • Claude Desktop environment variables - No .env files, credentials passed securely through Claude Desktop

  • Follows USPTO MCP patterns - Consistent with uspto_fpd, uspto_pfw, and uspto_ptab MCPs

  • Comprehensive .gitignore - Prevents accidental credential commits

  • Security guidelines - Complete documentation for secure development practices

  • Structured error responses - No sensitive information leakage in error messages

  • API key validation - Format checking and presence validation

  • HTTP transport authentication - When running in HTTP mode (FASTMCP_TRANSPORT=http) without OAuth, the server validates the X-API-KEY header against INTERNAL_AUTH_SECRET. That secret is required in this mode: if it never resolves, the server logs an error and refuses to start rather than serve an open deployment. The /health endpoint is always unauthenticated and is exempt from the inbound rate limit.

  • Security headers - X-Content-Type-Options, X-Frame-Options, Strict-Transport-Security, and Content-Security-Policy headers applied to all HTTP responses

  • Rate limiting - Token-bucket rate limiter (100 req/min default); in multi-replica deployments, enforce at the reverse proxy layer

Request Tracking & Debugging

Every request is tagged with a UUID4 request ID, returned to the caller as request_id on the response envelope and carried on the server's log records for correlation. Log lines carry flow metadata only (tool, request id, status, counts): query text, request and response bodies, headers and URLs are never logged, and the SanitizingFilter on every handler enforces that.

Documentation

  • SECURITY_GUIDELINES.md - Comprehensive security best practices

  • tests/README.md - Complete testing guide with API key setup

  • Enhanced error messages with request IDs for better support

πŸ“ Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Add tests for new functionality

  4. Ensure all tests pass

  5. Submit a pull request

πŸ“„ License

MIT License

⚠️ Disclaimer

THIS SOFTWARE IS PROVIDED "AS IS" AND WITHOUT WARRANTY OF ANY KIND.

Independent Project Notice: This is an independent personal project and is not affiliated with, endorsed by, or sponsored by the United States Patent and Trademark Office (USPTO).

The author makes no representations or warranties, express or implied, including but not limited to:

  • Accuracy & AI-Generated Content: No guarantee of data accuracy, completeness, or fitness for any purpose. Users are specifically cautioned that outputs generated or assisted by Artificial Intelligence (AI) components, including but not limited to text, data, or analyses, may be inaccurate, incomplete, fictionalized, or represent "hallucinations" (confabulations) by the AI model.

  • Availability: USPTO API and third-party dependencies may cause service interruptions.

  • Legal Compliance: Users are solely responsible for ensuring their use of this software, and any submissions or actions taken based on its outputs, strictly comply with all applicable laws, regulations, and policies, including but not limited to:

  • Legal Advice: This tool provides data access and processing only, not legal counsel. All results must be independently verified, critically analyzed, and professionally judged by qualified legal professionals.

  • Commercial Use: Users must verify USPTO terms for commercial applications.

  • Confidentiality & Data Security: The author makes no representations regarding the confidentiality or security of any data, including client-sensitive or technical information, input by the user into the software's AI components or transmitted to third-party services. Users are responsible for understanding and accepting the privacy policies, data retention practices, and security measures of any integrated third-party services.

  • Foreign Filing Licenses & Export Controls: Users are solely responsible for ensuring that the input or processing of any data, particularly technical information, through this software's AI components does not violate U.S. foreign filing license requirements (e.g., 35 U.S.C. 184, 37 CFR Part 5) or export control regulations (e.g., EAR, ITAR). This includes awareness of potential "deemed exports" if foreign persons access such data or if AI servers are located outside the United States.

LIMITATION OF LIABILITY: Under no circumstances shall the author be liable for any direct, indirect, incidental, special, or consequential damages arising from use of this software, even if advised of the possibility of such damages.

USER RESPONSIBILITY: YOU ARE SOLELY RESPONSIBLE FOR THE INTEGRITY AND COMPLIANCE OF ALL FILINGS AND ACTIONS TAKEN BEFORE THE USPTO.

  • Independent Verification: All outputs, analyses, and content generated or assisted by AI within this software MUST be thoroughly reviewed, independently verified, and corrected by a human prior to any reliance, action, or submission to the USPTO or any other entity. This includes factual assertions, legal contentions, citations, evidentiary support, and technical disclosures.

  • Duty of Candor & Good Faith: You must adhere to your duty of candor and good faith with the USPTO, including the disclosure of any material information (e.g., regarding inventorship or errors) and promptly correcting any inaccuracies in the record.

  • Signature & Certification: You must personally sign or insert your signature on any correspondence submitted to the USPTO, certifying your personal review and reasonable inquiry into its contents, as required by 37 CFR 11.18(b). AI tools cannot sign documents, nor can they perform the required human inquiry.

  • Confidential Information: DO NOT input confidential, proprietary, or client-sensitive information into the AI components of this software without full client consent and a clear understanding of the data handling practices of the underlying AI providers.

  • Export Controls: Be aware of and comply with all foreign filing license and export control regulations when using this tool with sensitive technical data.

  • Service Compliance: Ensure compliance with all USPTO (e.g., Terms of Use for USPTO websites, USPTO.gov account policies, restrictions on automated data mining) terms of service.

  • Security: Maintain secure handling of API credentials and client information.

  • Testing: Test thoroughly before production use.

  • Professional Judgment: This tool is a supplement, not a substitute, for your own professional judgment and expertise.

By using this software, you acknowledge that you have read this disclaimer and agree to use the software at your own risk, accepting full responsibility for all outcomes and compliance with relevant legal and ethical obligations.

Note for Legal Professionals: While this tool provides access to patent research tools commonly used in legal practice, it is a data retrieval and AI-assisted processing system only. All results require independent verification, critical professional analysis, and cannot substitute for qualified legal counsel or the exercise of your personal professional judgment and duties outlined in the USPTO Guidance on AI Use.

πŸ’ Support This Project

If you find this USPTO Enriched Citation MCP Server useful, please consider supporting the development! This project was developed during my personal time over many hours to provide a comprehensive, production-ready tool for the patent community.

Donate with PayPal

Your support helps maintain and improve this open-source tool for everyone in the patent community. Thank you!

Acknowledgments

  • USPTO for providing the Enriched Citation API v3 with AI-powered data extraction capabilities

  • Model Context Protocol for the MCP specification

  • Claude Code for exceptional development assistance, architectural guidance, documentation creation, PowerShell automation, test organization, and comprehensive code development throughout this project

  • Claude Desktop for additional development support and testing assistance


Questions? See INSTALL.md for complete installation guide or review the test scripts for working examples.

Shared USPTO rate limiting (multi-MCP deployments)

If you run all 4 USPTO MCPs (Citations, PFW, PTAB, FPD) as HTTP containers on the same box, serving multiple users, under one USPTO API key, each server's own in-process limiter can't see what the other 3 processes are doing β€” and USPTO's documented limits are per-key (burst=1, 4-15 req/sec depending on call type, plus weekly quotas), not per-process. Point all 4 containers at one bind-mounted directory and they share a single cross-process token bucket + a bounded pool of in-flight-request slots, arbitrated via POSIX file locks (crash-safe β€” a dead process's lock is released by the kernel). Single-MCP or STDIO deployments need nothing; the limiter is off unless the directory variable is set.

# docker-compose.yml (excerpt, all 4 USPTO MCP services)
volumes:
  uspto-rate-limit: {}
services:
  citations-mcp:
    volumes:
      - uspto-rate-limit:/var/run/uspto-shared-rate-limit
    environment:
      USPTO_SHARED_RATE_LIMIT_DIR: /var/run/uspto-shared-rate-limit
      USPTO_SHARED_RATE_LIMIT_RPS: "4"       # default; total across ALL 4 MCPs
      USPTO_SHARED_MAX_CONCURRENT: "2"       # default; shared in-flight slots

One token bucket and 2 concurrency slots are shared across every process mounting the directory β€” a heavier MCP naturally draws more of the budget under load, and a long PDF download occupies a slot for its full duration (not just connection setup), per USPTO's burst=1 guidance.

OAuth sign-in (optional)

Set CITATIONS_AUTH_MODE=oauth to protect the HTTP endpoint with Google + Microsoft sign-in (OAuth 2.1 with dynamic client registration β€” works as a Claude.ai / Claude Desktop custom connector). Access is controlled by a local SQLite user list; role admin unlocks the citations_manage_users user-management tool and its MCP App panel. The default (none) and STDIO are unchanged. Full walkthrough: docs/SSO_SETUP.md. Important: Set CITATIONS_ENABLE_USER_MANAGEMENT=true or the admin tool will not appear in tools/list.

Available Tools

10 tools
Citations_get_available_fieldsCitations_get_available_fieldsA
Read-only

Get all searchable fields from USPTO Enriched Citation API. Fields, available fields, columns, schema, what can I query, field names, query syntax for the enriched citations lane.

Use for: Field discovery, query syntax validation, understanding data structure. Returns: Complete field list with descriptions and types.

For field selection strategies and Solr/Lucene syntax examples, use Citations_get_guidance(section='fields').

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true, and the description reinforces that it only returns field information, which aligns. It does not describe the output format or any potential limitations (e.g., whether the field list is exhaustive or versioned), but for a zero-parameter read-only tool, the description is adequate. No contradiction found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized: a clear primary statement, a block of alternative search terms, a 'Use for' section, and a 'Returns' statement, all in a few sentences. It ends with a pointer to a sibling tool, making every sentence earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool, the description fully covers what the agent needs: purpose, usage scenarios, output, and differentiation from alternatives. The output schema is present, so no need to describe the return structure in detail. The guidance to another tool for advanced syntax is a bonus.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameters, and the description clarifies what the tool returns ('Complete field list with descriptions and types'). Since there are no parameters to document, the description provides sufficient context about the output, which adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves searchable fields from the USPTO Enriched Citation API, with a specific verb 'Get' and resource. It includes synonyms and use cases, making it distinct from siblings like Citations_get_oa_citation_fields and Citations_get_guidance, though it doesn't explicitly contrast with these.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly lists when to use: 'Use for: Field discovery, query syntax validation, understanding data structure.' It also gives an alternative for 'field selection strategies and Solr/Lucene syntax examples' by directing to Citations_get_guidance(section='fields'), making the distinction clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_get_citation_detailsCitations_get_citation_detailsA
Read-only

Get complete details for specific citation by ID. Full record for one citation, all fields, cited passage, column and line locator, figure, mapped claim, quality summary.

Use for deep analysis of strategically important citations. Full record with all fields and complete citing context.

⚠️ IMPORTANT: Returns citation METADATA only, NOT actual documents.

2-STEP PFW MCP WORKFLOW: Step 1: PFW_get_application_documents(app_number='{app_number}', document_code='CTNF', limit=20)

Document Code Decoder:

  • CTNF: Non-Final Office Action (where most citations appear β€” start here)

  • CTFR: Final Office Action Rejection

  • NOA: Notice of Allowance

  • 892: Examiner's Search Strategy & Citations List

  • IDS: Applicant's Information Disclosure Statement

Step 2a (LLM analysis): PFW_get_document_content_with_ocr(app_number, document_identifier) β†’ Extract text for analysis Step 2b (User download): PFW_get_document_download(app_number, document_identifier) β†’ PDF download link

For complete cross-MCP workflows, use Citations_get_guidance(section='workflows_pfw') for detailed integration patterns.

ParametersJSON Schema
NameRequiredDescriptionDefault
citation_idYes
include_contextNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds the important behavioral limitation that it returns citation metadata only and NOT actual documents. It also provides a concrete 2-step workflow for retrieving document content and downloads through related tools, which meaningfully discloses what this tool does not do and how to supplement it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is clear and front-loaded, but the description repeats the 'full record' idea and includes a long PFW workflow and document code decoder that are tangential to this tool's immediate invocation. The content is structured, but much of it could be moved to guidance content rather than the tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's primary caveat and cross-tool workflow, and an output schema exists, so return structure does not need to be explained. However, it omits explicit parameter semantics and does not clearly guide the agent from a search result to this tool's citation_id input, leaving some context incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is zero schema description coverage, and the description does not explain citation_id or include_context. The parameter names are somewhat self-explanatory, but the description fails to compensate for the missing schema details, especially what include_context toggles or how the ID should be supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves complete details for a specific citation by ID, enumerates the content returned (all fields, cited passage, locators, figure, mapped claim, quality summary), and distinguishes it from the sibling search tools by focusing on ID-based retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use it for deep analysis of strategically important citations and makes clear it is for looking up an existing citation by ID. It does not explicitly contrast with sibling search tools or state when not to use it, but the intended use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_get_citation_statisticsCitations_get_citation_statisticsA
Read-only

Get database statistics and aggregations for strategic planning. Counts, totals, aggregate, how many, distribution, breakdown by art unit or tech center, trends over time, citation volume.

⚠️ ENRICHED LANE ONLY. This tool aggregates the Enriched Citations (v3) index and nothing else. criteria is validated against the enriched field whitelist, so an OA-only clause (legalSectionCode, actionTypeCategory, paragraphNumber, referenceIdentifier, parsedReferenceIdentifier, workGroup) is a 400 here rather than a wrong answer, and there is no lane parameter: the OA Citations (v2) index has no statistics path on this server. That is a documented limit, not a bug to work around by rephrasing the clause.

To aggregate the OA lane, count it yourself with the OA search tools and read response.numFound, which is the whole-result total and not the page size: Citations_search_oa_citations_minimal(criteria='techCenter:2100 AND legalSectionCode:103', rows=1) One call per bucket gives the same breakdown shape this tool returns for the enriched lane. Any cross-lane comparison must state which lane each number came from; the two indexes are independent and neither is a superset of the other.

Returns for the enriched lane: total_citations, examiner_cited_count, applicant_cited_count, and breakdowns by citation category (X/Y/A) and by who cited.

ParametersJSON Schema
NameRequiredDescriptionDefault
criteriaNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses important behavioral traits: criteria is validated against an enriched field whitelist, OA-only clauses cause a 400, there is no lane parameter, and the OA index has no statistics path. It also states the returned fields for the enriched lane. This goes well beyond the annotation and gives the agent an accurate model of the tool's behavior and constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but every section earns its place: purpose, lane limitation, workaround, and return summary. The bold ENRICHED LANE ONLY warning is front-loaded, and the example is compact and directly useful. Given the complexity of the lane distinction, the length is appropriate and well structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one optional parameter, an output schema, and readOnlyHint already present, the description covers the essential contextual gaps: what the tool aggregates, what lane it targets, what happens on invalid criteria, how to handle the OA lane, and what the return shape is. The agent has enough information to select and invoke this tool correctly in almost any scenario.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only a 'criteria' string with a default of '' and zero description coverage, so the description carries the full burden. It does explain that criteria is validated against the enriched field whitelist and gives an example query ('techCenter:2100 AND legalSectionCode:103'), but it does not fully define the general query syntax or what an empty default means. Still, it adds substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource: 'Get database statistics and aggregations for strategic planning.' It then enumerates what counts as statistics (counts, totals, aggregate, distribution, breakdown by art unit/tech center, trends over time, citation volume), which sharply distinguishes it from sibling search tools. This makes it immediately clear this is an aggregation tool, not a search or details tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'ENRICHED LANE ONLY' and explains when to use an alternative: 'To aggregate the OA lane, count it yourself with the OA search tools.' It even provides a concrete example call and explains the limitation is 'a documented limit, not a bug to work around by rephrasing the clause.' This is model usage guidance with both when-to-use and when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_get_guidanceCitations_get_guidanceA
Read-only

Get selective USPTO Citation guidance sections for context-efficient workflows

🎯 QUICK REFERENCE - What section for your question?

πŸ” "Find citations by examiner/application/tech" β†’ fields πŸ”€ "Which lane: OA citations or enriched citations?" β†’ oa_citations πŸ“„ "Understand citation categories (X/Y/A + NPL via nplIndicator)" β†’ citation_codes πŸ”– "Citation date coverage per lane" β†’ data_coverage 🀝 "PFW workflow for office action documents" β†’ workflows_pfw 🚩 "PTAB citation correlation" β†’ workflows_ptab (updated for 2026 PTAB API) πŸ“Š "FPD petition citation patterns" β†’ workflows_fpd 🏒 "Complete lifecycle analysis" β†’ workflows_complete βš™οΈ "Tool guidance and parameters" β†’ tools ❌ "Search errors or query issues" β†’ errors πŸ’° "Reduce API costs and optimize" β†’ cost

Available sections:

  • overview: Available sections and tool summary

  • workflows_pfw: Citation + PFW integration workflows

  • workflows_ptab: Citation + PTAB integration workflows (updated 2026-01-17)

  • workflows_fpd: Citation + FPD integration workflows

  • workflows_complete: Four-MCP complete lifecycle analysis

PTAB Integration (updated 2026-01-17):

  • Trials: PTAB_search_trials_minimal/balanced/complete

  • Documents: PTAB_get_documents, PTAB_get_document_download, PTAB_get_document_content

  • See: Citations_get_guidance(section='workflows_ptab') for integration patterns

  • citation_codes: X/Y/A category decoder; NPL identified by nplIndicator:true field

  • oa_citations: OA (v2) vs enriched (v3) routing rule, measured coverage, field matrix

  • data_coverage: per-lane date coverage and date handling

  • fields: Field selection strategies and Solr/Lucene syntax

  • tools: Tool-specific guidance and parameters

  • errors: Common error patterns and troubleshooting

  • cost: Cost optimization strategies

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionNoWhich guidance section to retrieve (default: overview)overview

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations declare readOnlyHint=true, and the description consistently reflects a read-only documentation tool with no destructive behavior. It adds useful behavioral context by listing available sections, indicating PTAB content was updated on 2026-01-17, and noting that sections are retrieved selectively for context-efficient workflows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized and front-loaded with a clear purpose, but it is longer than necessary and contains redundancy: the quick-reference section, the 'Available sections' list, and the later detailed bullet list overlap significantly. The emoji formatting helps scannability but does not fully justify the repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with one optional parameter, a readOnlyHint annotation, a fully documented schema, and an output schema, the description is complete. It lists every available guidance section, explains what each covers, and gives enough context for an agent to select the right section and invoke the tool successfully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single 'section' parameter at 100% coverage, but the description adds substantial meaning by enumerating all valid section values, their purposes, and the default 'overview'. Since the schema has no enum, this enumeration is valuable beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence uses a specific verb ('Get') and resource ('selective USPTO Citation guidance sections'), and the quick-reference table clarifies it is a documentation/guidance tool rather than a data-retrieval tool. However, it does not explicitly contrast itself with the sibling search/data tools, so an agent might initially wonder whether 'guidance sections' is a search operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The quick-reference list maps user intents to specific sections (e.g., 'Which lane: OA citations or enriched citations?' β†’ oa_citations), which gives strong in-tool usage guidance. It lacks explicit exclusions or instructions about when to use a sibling search tool instead, but the guidance-vs-search distinction is largely implied by the content.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_get_oa_citation_fieldsCitations_get_oa_citation_fieldsA
Read-only

Get all searchable fields from the USPTO Office Action Citations API v2. Fields, available fields, columns, schema, what can I query, field names, query syntax for the OA citations lane.

Returns the complete field list for building Lucene queries against the OA Citations dataset. OA Citations v2 is the simpler counterpart to the AI-enriched citations β€” it provides raw citation data from Form 892 and Form 1449 office actions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint: true, which the description aligns with by stating it 'Returns the complete field list' – no mutation. The description adds context about the data source (raw citation data from Form 892 and Form 1449) and the lane distinction, which is useful beyond the annotation. However, it doesn't disclose any specifics about output format, pagination, or error behavior, which is acceptable for a simple metadata retrieval but not rich in additional behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is redundant in places: 'Fields, available fields, columns, schema, what can I query, field names, query syntax for the OA citations lane' is a list of synonyms that could be condensed. The core information is in the first sentence and the following two sentences, but the initial repetition adds noise. It's not extremely lengthy, but it could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters and an output schema (known to exist, though not provided), the description covers the essential context: it states the resource (OA Citations API v2), the purpose (return field list for Lucene queries), the lane distinction (vs AI-enriched), and the data source (Form 892/1449). This is sufficient for an agent to understand when and how to use it, leaving little missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description doesn't need to explain parameters because there are none, and the schema coverage is 100% (vacuous). The description does explain that it returns a field list, which is the only meaningful semantic context, but it doesn't need to compensate for any parameter documentation gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Get all searchable fields from the USPTO Office Action Citations API v2' and specifies its purpose: returning the complete field list for building Lucene queries. It distinguishes itself from the generic sibling 'Citations_get_available_fields' by noting that it targets the OA Citations lane and contrasts with the AI-enriched citations. This gives a specific verb, resource, and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: for retrieving field information specifically for the OA Citations v2 dataset, as opposed to the AI-enriched citations. It mentions 'simpler counterpart' and 'provides raw citation data from Form 892 and Form 1449,' which helps an agent choose this tool over the generic 'Citations_get_available_fields.' However, it doesn't explicitly state when not to use it or name direct alternatives, so it's clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_search_citations_balancedCitations_search_citations_balancedA
Read-only

Balanced citation search for analysis (80-85% context reduction). Prior art references cited by an examiner, cited passage, column and line locator, figure, mapped claim, art unit, tech center, examiner vs applicant citation.

Use after minimal search for detailed study of selected citations (10-20 results). 19 fields including passages, claims, office action category.

Solr/Lucene Query Examples:

  • Field search: criteria='groupArtUnitNumber:2854'

  • Date range: criteria='officeActionDate:[2023-01-01 TO 2023-12-31]'

  • Boolean: criteria='(citationCategoryCode:X OR citationCategoryCode:Y) AND techCenter:2100'

  • NPL only: criteria='nplIndicator:true AND techCenter:2100'

  • Complex: criteria='groupArtUnitNumber:2854 AND citationCategoryCode:X AND officeActionDate:[2020-01-01 TO *]'

NOT searchable: examinerNameText and firstApplicantNameText do NOT exist on this API. Examiner queries 400; applicant queries silently return 0. Resolve examiners and applicants through the PFW MCP, then query citations by application number.

Ultra-minimal mode: Pass custom fields list for 99% token reduction (2-3 fields only). Example: fields=['citedDocumentIdentifier', 'citationCategoryCode', 'passageLocationText']

Date handling: documented window is office actions mailed 2017-10-01 to ~30 days ago, but ~44% of TC2100 records carry an earlier officeActionDate in practice. Add an officeActionDate:[2017-10-01 TO *] clause only when you want the documented window specifically. For completeness, also query the OA lane and union.

Convenience parameters (balanced mode only):

  • decision_type: Office action type β€” use "CTNF" (non-final rejection) or "CTFR" (final rejection)

  • category_code: Citation relevance code β€” X (anticipatory Β§102/103), Y (combined Β§103), A (background)

  • examiner_cited: Boolean filter for examiner-cited references (true/false)

  • art_unit: Group art unit number (e.g., '2128', '3600')

CROSS-LANE JOIN KEY: every row carries referenceKey, the normalised reference identifier, and it is the ONLY correct key for unioning this lane with the OA lane. The two lanes write the same reference differently: on app 12849948 the OA parsedReferenceIdentifier reads '20060075466' while the enriched citedDocumentIdentifier reads 'US 2006/0075466 A1'. Joining those two raw fields finds zero overlap on every application; the true answer there is four references in both lanes. referenceKey is digits only (a leading US, spaces, slashes, hyphens and the kind code stripped, series markers such as RE kept), derived from publicationNumber first and citedDocumentIdentifier second, and carried on both lanes at every tier including a custom fields list.

ROWS WITH NO REFERENCE: referenceKey is null when the row carries no usable identifier, and the response envelope reports how many such rows the page holds as rows_without_reference_identifier (always present, 0 included). An absent citedDocumentIdentifier key, a null one and an empty string are ONE state, not three: a row can carry an empty publicationNumber with the citedDocumentIdentifier key missing from the JSON entirely. Measured: 2 of 5 on app 11752072, 4 of 8 on 12849948, 4 of 26 on 18407147. Those rows are real citations and must be reported as unresolved, never dropped.

IDENTIFIERS: patent_number takes either a GRANTED patent number (7-8 digits; commas, spaces and a US prefix are accepted) or an 11-digit pre-grant publication number. A granted patent number is crosswalked to its application serial with one USPTO ODP applications-search call and queried as patentApplicationNumber; an 11-digit value queries publicationNumber directly. The response reports which reading was used in patent_number_resolution {input, interpreted_as, resolved_application_number when crosswalked, source}. A number that resolves to no application is a 400 naming the accepted forms, not a zero-result. application_number remains the application serial; passing one that disagrees with the crosswalked patent number is also a 400.

Note: Returns citation metadata only. For the office action text itself, use the PFW MCP's PFW_get_oa_text / PFW_get_oa_rejections (direct, no document-bag + OCR round trip).

For complex workflows and cross-MCP integration, use Citations_get_guidance(section). Quick reference: 'oa_citations' for OA-vs-enriched routing, 'fields' for Solr syntax, 'workflows_pfw'/'workflows_ptab'/'workflows_fpd' for integration patterns.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsNo
startNo
fieldsNo
art_unitNo
criteriaNo
date_endNo
date_startNo
tech_centerNo
category_codeNo
decision_typeNo
patent_numberNo
applicant_nameNo
examiner_citedNo
application_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only readOnlyHint=true as an annotation, the description carries the full burden of behavioral disclosure, and it does so extensively. It discloses date handling quirks (documented window vs actual practice), the behavior for rows without a reference identifier (must be reported as unresolved, never dropped), the cross-lane join key semantics, error responses for invalid patent numbers, and the ultra-minimal mode. There is no contradiction with the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, but it is well-structured with clear sections (Solr examples, convenience parameters, join key, identifiers, limitations) and front-loads the core purpose and usage. Every sentence contributes essential information for a complex 14-parameter tool, though some redundancy exists in the referenceKey and rows-without-reference sections. It is not concise in absolute terms but is appropriately sized for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (14 params, output schema present, many edge cases), the description is remarkably complete. It covers the purpose, usage, error handling, data quirks, integration with sibling tools and PFW MCP, and even explains the output envelope (rows_without_reference_identifier). Nothing an agent needs to invoke it correctly or interpret results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all parameters. It explicitly explains convenience parameters (decision_type, category_code, examiner_cited, art_unit), patent_number vs application_number resolution, and provides Solr query examples for criteria and fields. It also notes that applicant_name queries silently return 0 and examinerNameText is not searchable. This adds substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: a balanced citation search for analysis with 80-85% context reduction, and enumerates the types of data it retrieves (prior art references, passages, locators, claims, etc.). It distinguishes itself from sibling tools like the minimal search and the OA lane, explicitly positioning itself as the follow-up to minimal search for detailed study of 10-20 results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Use after minimal search for detailed study of selected citations (10-20 results)'. It also names alternatives: for office action text use PFW MCP's get_oa_text/get_oa_rejections, for complex workflows use Citations_get_guidance(section), and for examiner/applicant resolution use the PFW MCP. It clearly states what is NOT searchable and the consequences (400 vs silent 0), giving unambiguous routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_search_citations_minimalCitations_search_citations_minimalA
Read-only

Minimal citation search for discovery (90-95% context reduction).

Use for high-volume pattern discovery before detailed analysis. Essential 8 fields: application, publication, art unit, citation ID, category, tech center, date, examiner indicator.

Solr/Lucene Query Examples:

  • Field search: criteria='groupArtUnitNumber:2854'

  • Date range: criteria='officeActionDate:[2017-10-01 TO *]'

  • Boolean: criteria='citationCategoryCode:X AND techCenter:2100'

  • Wildcard: criteria='citedDocumentIdentifier:US*'

  • Combined: criteria='groupArtUnitNumber:2854 AND officeActionDate:[2023-01-01 TO 2023-12-31]'

Ultra-minimal mode: Pass custom fields list for 99% token reduction (2-3 fields only). Example: fields=['citedDocumentIdentifier', 'patentApplicationNumber'] for PFW integration.

Date handling: USPTO documents this API as office actions mailed 2017-10-01 to ~30 days ago. In practice ~44% of TC2100 records carry an earlier officeActionDate (verified against PFW document dates back to 2010-2012). Do NOT add a blanket officeActionDate:[2017-10-01 TO *] clause unless you specifically want the documented window β€” it discards records the index actually serves.

Lane routing β€” TRY BOTH: this is the ENRICHED lane (passage locations, claim mapping, quality scores, NPL flag, date filtering). For completeness-sensitive questions also run Citations_search_oa_citations_minimal (raw 892/1449 lists, statutory basis, broader applicant-IDS coverage) and union the results β€” neither lane is a superset of the other. See Citations_get_guidance(section='oa_citations').

CROSS-LANE JOIN KEY: every row carries referenceKey, the normalised reference identifier, and it is the ONLY correct key for unioning this lane with the OA lane. The two lanes write the same reference differently: on app 12849948 the OA parsedReferenceIdentifier reads '20060075466' while the enriched citedDocumentIdentifier reads 'US 2006/0075466 A1'. Joining those two raw fields finds zero overlap on every application; the true answer there is four references in both lanes. referenceKey is digits only (a leading US, spaces, slashes, hyphens and the kind code stripped, series markers such as RE kept), derived from publicationNumber first and citedDocumentIdentifier second, and carried on both lanes at every tier including a custom fields list.

ROWS WITH NO REFERENCE: referenceKey is null when the row carries no usable identifier, and the response envelope reports how many such rows the page holds as rows_without_reference_identifier (always present, 0 included). An absent citedDocumentIdentifier key, a null one and an empty string are ONE state, not three: a row can carry an empty publicationNumber with the citedDocumentIdentifier key missing from the JSON entirely. Measured: 2 of 5 on app 11752072, 4 of 8 on 12849948, 4 of 26 on 18407147. Those rows are real citations and must be reported as unresolved, never dropped.

IDENTIFIERS: patent_number takes either a GRANTED patent number (7-8 digits; commas, spaces and a US prefix are accepted) or an 11-digit pre-grant publication number. A granted patent number is crosswalked to its application serial with one USPTO ODP applications-search call and queried as patentApplicationNumber; an 11-digit value queries publicationNumber directly. The response reports which reading was used in patent_number_resolution {input, interpreted_as, resolved_application_number when crosswalked, source}. A number that resolves to no application is a 400 naming the accepted forms, not a zero-result. application_number remains the application serial; passing one that disagrees with the crosswalked patent number is also a 400.

Note: Returns citation metadata only. For the office action text itself, use the PFW MCP's PFW_get_oa_text / PFW_get_oa_rejections (direct, no document-bag + OCR round trip).

For complex workflows and cross-MCP integration, use Citations_get_guidance(section). Quick reference: 'fields' section for Solr syntax, 'workflows_pfw' for PFW integration.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsNo
startNo
fieldsNo
art_unitNo
criteriaNo
date_endNo
date_startNo
tech_centerNo
patent_numberNo
applicant_nameNo
examiner_citedNo
application_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral edge cases: documented date windows that don't match actual index contents, referenceKey normalization differences, null/missing identifier handling, 400 responses for unresolved patent numbers, and the fact that rows without references must be preserved. It also clarifies the response envelope includes rows_without_reference_identifier. These are exactly the non-obvious behaviors an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: query examples, date-window caveat, lane routing, cross-lane join key, null-row handling, identifier resolution, and cross-MCP integration. The most important purpose and usage guidance are front-loaded, and the internal headings make the density navigable. There is no filler or repeated schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 12 parameters, zero schema descriptions, and the presence of an output schema, this description is exceptionally complete: it covers query syntax, field selection, identifier resolution, error behavior, cross-tool routing, and data-quality pitfalls. The output schema already covers return shape, so the description need not restate it. An agent has enough context to call this tool correctly in a wide range of discovery workflows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the full burden, and it does add meaning for the hardest parameters: criteria (Solr/Lucene syntax with examples), fields (custom field list and token reduction), patent_number, and application_number (resolution, crosswalking, 400 behavior). It does not explicitly document rows, start, date_start, date_end, art_unit, tech_center, applicant_name, or examiner_cited, though several are inferable from examples or names. This is strong but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Minimal citation search for discovery' with an explicit 'Essential 8 fields' list. It also distinguishes itself from the OA lane by labeling this the 'ENRICHED lane' and naming the sibling tool to use for raw OA citations. The purpose is concrete and not a tautology of the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: 'Use for high-volume pattern discovery before detailed analysis.' It also provides strong when-not guidance, such as 'Do NOT add a blanket officeActionDate:[2017-10-01 TO *] clause' and 'For completeness-sensitive questions also run Citations_search_oa_citations_minimal ... and union the results.' It further routes office-action-text requests to PFW tools, making alternatives clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_search_oa_citations_balancedCitations_search_oa_citations_balancedA
Read-only

Search Office Action Citations (v2) with all 16 available fields. Prior art cited against an application, Form 892, Form 1449, IDS references, 102 103 112 statutory basis, action type, paragraph number.

Use after Citations_search_oa_citations_minimal for detailed analysis of selected applications. All fields: patentApplicationNumber, groupArtUnitNumber, techCenter, referenceIdentifier, parsedReferenceIdentifier, actionTypeCategory, legalSectionCode, examinerCitedReferenceIndicator, applicantCitedExaminerReferenceIndicator, officeActionCitationReferenceIndicator, workGroup, paragraphNumber, createDateTime, createUserIdentifier, obsoleteDocumentIdentifier, id.

This tier is where the OA-only analytical fields live: legalSectionCode (102/103/112 statutory basis) and actionTypeCategory ('rejected') have no equivalent in the enriched lane, and paragraphNumber locates the citation within the office action.

⚠️ APPLICANT-CITED (1449/IDS) COVERAGE IS PARTIAL. This lane is documented upstream as transcribing Form 892 AND Form 1449, but on IDS-heavy files it returns close to what the examiner applied and little else, in every era. Measured against the patents' own References Cited pages (union of BOTH lanes): US 7,971,071 -> 5 of 91 US 9,496,922 -> 1 of 251 US 9,135,462 -> 0 of about 620 US 11,656,067 -> 3 of 15, prosecuted 2021-2023 INSIDE the documented window A reference's absence here is NO evidence that the applicant did not disclose it, and a count from this lane is not the applicant's full IDS. For a complete 1449 record, read the IDS documents through the PFW MCP.

CROSS-LANE JOIN KEY: every row carries referenceKey, the normalised reference identifier, and it is the ONLY correct key for unioning this lane with the enriched lane. On app 12849948 this lane's parsedReferenceIdentifier reads '20060075466' while the enriched citedDocumentIdentifier reads 'US 2006/0075466 A1'; joining those two raw fields finds zero overlap on every application, when the true answer there is four references in both lanes. referenceKey is digits only (a leading US, spaces, slashes, hyphens and the kind code stripped, series markers such as RE kept) and is carried on both lanes at every tier. It is null on a row whose identifier does not reduce to a document number.

OA Citations v2 documented window: office actions mailed 2017-10-01 to ~30 days prior to today (older records have been observed in practice). No office-action date field exists β€” do not add an officeActionDate clause (HTTP 400).

⚠️ publicationNumber IN criteria IS A DELIBERATE 400 HERE, AND THAT IS A FEATURE. The raw upstream API answers that field with HTTP 200 and numFound 0, which reads as "this patent was never cited" and is silently wrong. This server refuses the clause so the mistake is visible. Use the patent_number parameter for the subject patent, or parsedReferenceIdentifier to find where a patent was CITED.

IDENTIFIERS: application_number is the APPLICATION serial. This index has no patent-number field (publicationNumber returns HTTP 400), so patent_number is crosswalked here: pass a GRANTED patent number (7-8 digits; commas, spaces and a US prefix accepted) and it is resolved to its application serial with one USPTO ODP applications-search call, then queried as patentApplicationNumber. The response reports the mapping in patent_number_resolution {input, interpreted_as, resolved_application_number, source}. An 11-digit pre-grant publication number is refused here (use Citations_search_citations_balanced for those), an unresolvable number is a 400 naming the accepted forms, and a patent_number that disagrees with a supplied application_number is a 400 rather than a query that can only return zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsNo
startNo
fieldsNo
art_unitNo
criteriaNo
tech_centerNo
patent_numberNo
examiner_citedNo
application_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true. The description adds extensive behavioral context: partial applicant-cited coverage with concrete patent metrics, the cross-lane join key (referenceKey), the documented time window, the deliberate HTTP 400 for publicationNumber, and the patent_number crosswalk behavior. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and placement, then organizes critical caveats into distinct sections. It is verbose, but each section addresses a unique concern (coverage warning, join key, window, 400 behavior, identifier rules) that materially affects correct use. The length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all essential operational details: the 16 fields, the unique analytical fields, the partial coverage caveat with evidence, the join key, the time window, the deliberate 400 behavior, and the patent_number resolution. With an output schema present, an agent has everything needed to call it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description compensates well for the ambiguous parameters: patent_number (crosswalk to application serial), application_number (application serial), parsedReferenceIdentifier (join key usage), and criteria (publicationNumber refusal). It does not explicitly explain rows, start, fields, art_unit, tech_center, or examiner_cited, but these are self-evident from names and defaults. Overall strong compensation for the high-stakes parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it searches Office Action Citations (v2) with all 16 fields, and positions itself as the detailed tier after the minimal search. It distinguishes from siblings by naming the unique analytical fields (legalSectionCode, actionTypeCategory, paragraphNumber) and by referencing the enriched lane.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use after Citations_search_oa_citations_minimal for detailed analysis, and gives alternatives: use Citations_search_citations_balanced for pre-grant publications, and use PFW MCP for complete 1449 records. Also warns about partial applicant-cited coverage, guiding when not to rely on this lane.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_search_oa_citations_minimalCitations_search_oa_citations_minimalA
Read-only

Search Office Action Citations (v2) for high-volume discovery (8 key fields).

OA Citations v2 is the raw citation list transcribed from Form PTO-892 (examiner) and Form PTO-1449 (applicant IDS). Usually broader than the enriched lane in bulk (measured TC2100: 4.87M vs 4.32M records), with most of the surplus being applicant IDS references β€” but NOT a superset: on a given application the enriched lane can return more (measured: app 12849948 returns 4 here vs 8 enriched). For any completeness-sensitive question, run BOTH lanes and union the results.

⚠️ APPLICANT-CITED (1449/IDS) COVERAGE IS PARTIAL. This lane is documented upstream as transcribing Form 892 AND Form 1449, but on IDS-heavy files it returns close to what the examiner applied and little else, in every era. Measured against the patents' own References Cited pages (union of BOTH lanes): US 7,971,071 -> 5 of 91 US 9,496,922 -> 1 of 251 US 9,135,462 -> 0 of about 620 (both lanes return zero) US 11,656,067 -> 3 of 15, prosecuted 2021-2023 INSIDE the documented window, and all three are the examiner's own double-patenting family citations, none of them the twelve references a later IPR petition relied on. Treat a reference's absence here as NO evidence that the applicant did not disclose it, and never present a count from this lane as the applicant's full IDS. For a complete 1449 record, read the IDS documents themselves through the PFW MCP.

Key fields returned: patentApplicationNumber, groupArtUnitNumber, techCenter, referenceIdentifier, parsedReferenceIdentifier, actionTypeCategory, examinerCitedReferenceIndicator, createDateTime. The OA API ignores fl, so this set is enforced client-side β€” the tier really does return only these eight. The PFW hand-off is stated once on the response envelope as pfw_link, not repeated on every row. legalSectionCode and paragraphNumber are NOT here; use the balanced tier or pass an explicit fields list for them.

CROSS-LANE JOIN KEY: every row carries referenceKey, the normalised reference identifier, and it is the ONLY correct key for unioning this lane with the enriched lane. The two lanes write the same reference differently: on app 12849948 this lane's parsedReferenceIdentifier reads '20060075466' while the enriched citedDocumentIdentifier reads 'US 2006/0075466 A1'. Joining those two raw fields finds zero overlap on every application; the true answer there is four references in both lanes. referenceKey is digits only (a leading US, spaces, slashes, hyphens and the kind code stripped, series markers such as RE kept), derived from parsedReferenceIdentifier first and the raw referenceIdentifier second, and carried on both lanes at every tier including a custom fields list. It is null on a row whose identifier does not reduce to a document number, which is an unjoinable row rather than a missing one.

Solr/Lucene Query Examples:

  • By application: criteria='patentApplicationNumber:18180061'

  • By tech center: criteria='techCenter:2100'

  • By art unit: criteria='groupArtUnitNumber:2854'

  • Examiner-cited only: criteria='examinerCitedReferenceIndicator:true'

  • Statutory basis (OA-ONLY capability): criteria='techCenter:2100 AND legalSectionCode:103'

  • Where a patent was cited: criteria='parsedReferenceIdentifier:9280610'

⚠️ NO DATE FIELD. officeActionDate does not exist here and returns HTTP 400 β€” the index already IS the 2017-10-01+ window, so omit any date clause. createDateTime is an ETL load stamp, NOT the office action date β€” never present it as prosecution chronology.

⚠️ publicationNumber IN criteria IS A DELIBERATE 400 HERE, AND THAT IS A FEATURE. The raw upstream API does not reject that field: it answers HTTP 200 with numFound 0, which reads exactly like "this patent was never cited" and is silently wrong. This server refuses the clause instead so the mistake is visible. Use the patent_number parameter, which crosswalks a granted patent number to the application serial this index does hold, or query parsedReferenceIdentifier to find where a patent was CITED.

⚠️ Use parsedReferenceIdentifier (normalized) rather than referenceIdentifier for reference lookups β€” the raw string format varies for the same patent.

IDENTIFIERS: application_number is the APPLICATION serial. This index has no patent-number field (publicationNumber returns HTTP 400), so patent_number is crosswalked here: pass a GRANTED patent number (7-8 digits; commas, spaces and a US prefix accepted) and it is resolved to its application serial with one USPTO ODP applications-search call, then queried as patentApplicationNumber. The response reports the mapping in patent_number_resolution {input, interpreted_as, resolved_application_number, source}. An 11-digit pre-grant publication number is refused here (use Citations_search_citations_minimal for those), an unresolvable number is a 400 naming the accepted forms, and a patent_number that disagrees with a supplied application_number is a 400 rather than a query that can only return zero.

Use Citations_search_oa_citations_balanced for full 16-field detail (adds legalSectionCode, paragraphNumber, parsedReferenceIdentifier). For passage locations, claim mapping, NPL flags, or date filtering, use Citations_search_citations_minimal β€” and run it alongside this tool by default. Coverage: USPTO documents both APIs as office actions mailed 2017-10-01 to ~30 days ago; in practice both have been observed serving older records, so do not treat an older application as out of scope without querying. Routing detail: Citations_get_guidance(section='oa_citations').

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsNo
startNo
fieldsNo
art_unitNo
criteriaNo
tech_centerNo
patent_numberNo
examiner_citedNo
application_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only readOnlyHint=true in annotations, the description carries the full burden of behavioral disclosure, and it excels. It reveals significant caveats: partial applicant-cited coverage (with specific measured examples), the absence of a date field (officeActionDate returns 400), the deliberate rejection of publicationNumber in criteria, the referenceKey join mechanism, and the meaning of createDateTime. No contradiction with annotations; the read-only nature is consistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but it earns its length given the tool's complexity (9 parameters, many pitfalls). It uses clear section headers (warning symbols, query examples, identifiers) and bullet-like structure, front-loads the purpose, and avoids redundancy. It is not perfectly conciseβ€”some caveats could be tightenedβ€”but it is well-organized and every paragraph adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 9-parameter schema with zero descriptions, the presence of an output schema, and the intricate domain (patent citations), this description is exceptionally complete. It covers coverage limitations, cross-lane join keys, absence of date fields, publicationNumber handling, identifier semantics, and sibling routingβ€”so nothing an agent needs to invoke the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no parameter descriptions), so the description must compensateβ€”and it does comprehensively. It explains `criteria` with five concrete Solr query examples, defines `patent_number` as a crosswalked granted patent number (with accepted formats and error behavior), clarifies `application_number` as the application serial, and notes that `fields` is ignored by the OA API and the 8-field set is enforced client-side. Every parameter is given practical meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb ('Search'), a precise resource ('Office Action Citations (v2)'), and the use case ('high-volume discovery (8 key fields)'). It further distinguishes from siblings by naming the balanced and minimal tiers, and explicitly notes the 8-key-field limitation. The purpose is unmistakable even without reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to use this tool and when to use alternatives. It advises running BOTH this lane and the enriched lane for completeness-sensitive questions, states 'Use Citations_search_oa_citations_balanced for full 16-field detail' and 'Use Citations_search_citations_minimal for passage locations, claim mapping, NPL flags, or date filtering', and even suggests defaulting to running minimal alongside. This is exemplary routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Citations_validate_queryCitations_validate_queryA
Read-only

Validate Lucene query syntax and provide optimization suggestions. Check my query, syntax error, is this query valid, Lucene, Solr, escaping, dry run before searching, why did my search fail.

Solr/Lucene Syntax Examples:

  • Field search: 'groupArtUnitNumber:2854'

  • Date range: 'officeActionDate:[2023-01-01 TO 2023-12-31]'

  • Boolean operators: 'citationCategoryCode:X AND techCenter:2100'

  • OR logic: '(citationCategoryCode:X OR citationCategoryCode:Y)'

  • NOT operator: 'techCenter:2100 NOT groupArtUnitNumber:1600'

  • Wildcard: 'citedDocumentIdentifier:US*'

  • NPL only: 'nplIndicator:true'

  • Open-ended range: 'officeActionDate:[2017-10-01 TO *]'

For comprehensive query syntax guide, use Citations_get_guidance(section='fields').

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
field_setNocitations_minimal

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only. The description adds useful behavioral context: it performs a dry-run validation rather than executing a search, and it returns optimization suggestions. This aligns with readOnlyHint and clarifies the side-effect-free nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then uses examples to make Lucene syntax concrete. It is slightly long, but the example list earns its place for a validation tool where syntax is the main concern.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the read-only annotation, output schema, and sibling context, most of what an agent needs to invoke the tool is present. The main gap is the undocumented 'field_set' parameter, which the description neither defines nor illustrates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must explain the parameters. It thoroughly explains the 'query' parameter with Solr/Lucene syntax examples, but it never mentions the 'field_set' parameter or its allowed values/meaning, leaving a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Validate Lucene query syntax and provide optimization suggestions.' This clearly distinguishes it from sibling search/guidance tools and includes concrete examples of what the tool accepts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly lists likely user intents ('Check my query, syntax error, is this query valid, ... dry run before searching, why did my search fail') and tells the agent to use Citations_get_guidance for a comprehensive syntax guide. It does not explicitly say when not to use it in favor of a search tool, but the validation context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv1.0.0
    • First observedCitations_get_available_fields
    • First observedCitations_get_citation_details
    • First observedCitations_get_citation_statistics
    • First observedCitations_get_guidance
    • First observedCitations_get_oa_citation_fields
    • First observedCitations_search_citations_balanced
    • First observedCitations_search_citations_minimal
    • First observedCitations_search_oa_citations_balanced
    • First observedCitations_search_oa_citations_minimal
    • First observedCitations_validate_query

TDQS

A4.4/5.0

Scored across 10 tools

Disambiguation5/5

Each tool targets a clearly distinct function: field discovery, search per lane/tier, detail lookup, query validation, statistics, and guidance. The enriched vs OA lane split and minimal vs balanced tiers are explicitly and repeatedly disambiguated in the descriptions.

Naming Consistency5/5

All tools share the Citations_ prefix and follow a consistent snake_case verb_noun pattern, with modifiers like minimal/balanced and lane qualifiers like oa applied uniformly. The naming makes the tool relationships and purposes predictable.

Tool Count5/5

Ten tools is well-scoped for a patent citation search domain: two search tiers across two data lanes, field discovery, detail retrieval, validation, statistics, and guidance. Each tool serves a distinct part of the workflow without redundancy.

Completeness5/5

The surface covers the full read-only citation lifecycle: field discovery, query validation, minimal and balanced search on both citation lanes, detailed record retrieval, and aggregate statistics. Cross-MCP handoffs to PFW and PTAB are also documented, so agents are not left at dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers