PageIndex MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
| resources | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| process_documentB | Upload and process PDF documents from URLs or local files. Supports OCR processing, hierarchical content extraction, and intelligent document analysis. Returns a unique doc_id for subsequent operations. Processing typically takes 0-3 minutes depending on document size (estimate: 2 seconds per page). Supports files up to 100MB. |
| browse_documentsA | Primary document retrieval tool. After orienting with get_folder_structure() (when available), use this for all document-related questions. The bare call returns root-level sub-folders and documents; pass folder_id to drill into a sub-folder level by level. Use sort="relevance" + query for semantic ranking. Do NOT jump to search_documents() first — it is an escalation path, only after browse_documents(sort="relevance") has failed. |
| search_documentsA | ESCALATION tool — never the first step. Use only after |
| get_folder_structureA | Orientation step: show the folder hierarchy as a tree (like |
| get_documentA | Check a document's processing status and metadata. |
| get_document_structureA | Extract a document's hierarchical outline (headers, sections, page references). REQUIRED for documents over 20 pages — call this first to locate relevant sections, then pass their page numbers to |
| get_page_contentA | Extract page content from a processed document. Use tight, targeted page ranges — never the whole document at once. For documents over 20 pages, call |
| get_document_imageA | Retrieve an image from a document — pass an |
| remove_documentA | Permanently delete documents and all associated data. Only invoke when the user explicitly names the documents AND confirms deletion. Returns |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| getting-started | How to connect to the PageIndex MCP server: endpoints, OAuth and API-key authentication, and self-serve signup. |
| about-pageindex | What PageIndex is: vectorless, reasoning-based document retrieval with exact page references. |
TDQS
Scored across 9 tools
Each tool has a distinct role in the document lifecycle, and overlapping discovery tools (browse vs. search) are explicitly differentiated. However, get_folder_structure vs. get_document_structure could momentarily confuse an agent despite their different targets.
All tools follow a consistent snake_case verb_noun pattern, with get_* used uniformly for retrieval-oriented operations and clear action verbs like process, remove, browse, and search. There are no mixed conventions or vague generic names.
Nine tools is well within the ideal range and each tool earns its place, covering ingestion, deletion, navigation, search, status checking, structure extraction, page content, and image retrieval without redundancy.
The tool surface covers the full document workflow: add/remove documents, browse/search the corpus, check processing status, navigate structure, extract page content, and retrieve images. There are no obvious dead ends or missing operations for the stated domain.