Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{}
resources
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
process_documentB

Upload and process PDF documents from URLs or local files. Supports OCR processing, hierarchical content extraction, and intelligent document analysis. Returns a unique doc_id for subsequent operations. Processing typically takes 0-3 minutes depending on document size (estimate: 2 seconds per page). Supports files up to 100MB.

browse_documentsA

Primary document retrieval tool. After orienting with get_folder_structure() (when available), use this for all document-related questions. The bare call returns root-level sub-folders and documents; pass folder_id to drill into a sub-folder level by level. Use sort="relevance" + query for semantic ranking. Do NOT jump to search_documents() first — it is an escalation path, only after browse_documents(sort="relevance") has failed.

search_documentsA

ESCALATION tool — never the first step. Use only after browse_documents(sort="relevance", query=...) missed a document you strongly believe exists by precise term, acronym, or file-name fragment. Query must be keywords only — see the query schema describe. Each result has a score (6-10, higher is more relevant). On result: single best match → read it; several equally good → ask user; none → fall back to browse_documents(sort="relevance").

get_folder_structureA

Orientation step: show the folder hierarchy as a tree (like tree -d). Call this before browse_documents() to plan targeted retrieval — folder names reveal content domains (e.g. "Research", "Finance") and return folder IDs you pass to browse_documents(folder_id=…). Wait for this result before issuing browse_documents() or search_documents(). For folder_id="root" (default), returns the entire folder tree; for a specific folder_id, returns that subtree only.

get_documentA

Check a document's processing status and metadata. status is one of "pending", "queued", "processing", "completed", or "failed" — call this before get_document_structure() or get_page_content() to confirm the document is ready.

get_document_structureA

Extract a document's hierarchical outline (headers, sections, page references). REQUIRED for documents over 20 pages — call this first to locate relevant sections, then pass their page numbers to get_page_content(). Use the part parameter to iterate large outlines until pagination.has_more is false.

get_page_contentA

Extract page content from a processed document. Use tight, targeted page ranges — never the whole document at once. For documents over 20 pages, call get_document_structure() first to pick relevant sections. Embedded image paths in the response feed into get_document_image().

get_document_imageA

Retrieve an image from a document — pass an image_path from a get_page_content() response: either an embedded image (from a content block) or a full rendered page (from the page_images array, useful when a page is scanned or its meaning depends on visual layout). Returns the image as an MCP image content block (base64 data + MIME type). Do NOT repeat the raw base64 value in your text response.

remove_documentA

Permanently delete documents and all associated data. Only invoke when the user explicitly names the documents AND confirms deletion. Returns results — one entry per requested document: { doc_name, status: "deleted" | "not_found" | "failed", error? }. Inspect each entry for per-document failures. This action is irreversible.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
getting-startedHow to connect to the PageIndex MCP server: endpoints, OAuth and API-key authentication, and self-serve signup.
about-pageindexWhat PageIndex is: vectorless, reasoning-based document retrieval with exact page references.

TDQS

A4.2/5.0

Scored across 9 tools

Disambiguation4/5

Each tool has a distinct role in the document lifecycle, and overlapping discovery tools (browse vs. search) are explicitly differentiated. However, get_folder_structure vs. get_document_structure could momentarily confuse an agent despite their different targets.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern, with get_* used uniformly for retrieval-oriented operations and clear action verbs like process, remove, browse, and search. There are no mixed conventions or vague generic names.

Tool Count5/5

Nine tools is well within the ideal range and each tool earns its place, covering ingestion, deletion, navigation, search, status checking, structure extraction, page content, and image retrieval without redundancy.

Completeness5/5

The tool surface covers the full document workflow: add/remove documents, browse/search the corpus, check processing status, navigate structure, extract page content, and retrieve images. There are no obvious dead ends or missing operations for the stated domain.

Maintenance

ActivityStale
ResponsivenessWithin a week