Skip to main content
Glama

Citra

Give your AI agent eyes for PDFs — with proof.

Local-first PDF evidence for agents. Structured text, tables, OCR, visual crops, and page-level citations your agent can defend — not invent.

Canonical package @sylphx/citra · bin citra · MCP io.github.SylphxAI/citra · live 5.0.0

npm version License: MIT stars

Zero-config in one line

npx -y @sylphx/citra

No Docker. No API key. No global install. Spawns a stdio MCP server agents can use immediately.

Client

Setup

Any agent / CLI

npx -y @sylphx/citra

Claude Code

claude mcp add citra -- npx -y @sylphx/citra

Claude Desktop / Cursor / VS Code / Codex

"command": "npx", "args": ["-y", "@sylphx/citra"]

Global CLI

npm i -g @sylphx/citracitra

Related MCP server: MCP PDF Reader

Why Citra feels unfairly good

Plain-text PDF tools make agents guess. Citra returns an Agent Document Twin they can cite.

Pain today

With Citra

Page numbers invented or missing

Page + geometry + provenance

Tables flattened into soup

Rows · columns · cells · bounding boxes

Scanned PDFs become noise

OCR path linked to evidence

Install / config / “hope it works”

npx -y — done

Silent engine fallbacks

Fail closed if the native binary is missing

Five reasons teams pick Citra

  1. Zero-config — real npx MCP, not a 20-step bootstrap.

  2. Evidence, not vibes — citations agents can show a human.

  3. Local-first — PDFs stay on the machine; no required cloud vision API.

  4. Brand-sole — one package, one bin, one story (@sylphx/citra / citra).

  5. Instrument family — compose with Iris (image), Cue (video), Spine, Lookout, Locus.

See the difference

Plain text vs evidence

Without evidence

With Citra

“Revenue was about $12M”

“Page 14, Table 3, cell (row 4, col 2) = $12.4M

Lost table structure

Rows, columns, cells, bounding boxes

Scanned PDF = garbage text

OCR with page-linked evidence

Hidden / adversarial text ignored

Trust signals when requested

What you get

Three tools. One product surface.

Tool

What agents use it for

read_pdf

Smart default: markdown, tables, structure, OCR, citations

search_pdf

Find page + snippet matches before deep reading

pdf_evidence

Crops, renders, inspect, focused evidence ops

Minimal call:

{
  "sources": [{ "path": "/absolute/path/to/report.pdf" }]
}

Flagship use cases

  1. Financial reports — extract table cells agents can cite by page and geometry

  2. Research papers — headings, reading order, page-level quotes

  3. Scanned documents — OCR path with evidence, not a text soup

Platforms

One optional native package is selected for your host only:

Platform

Native package

macOS arm64

@sylphx/citra-darwin-arm64

macOS x64

@sylphx/citra-darwin-x64

Linux x64

@sylphx/citra-linux-x64-gnu

Linux arm64

@sylphx/citra-linux-arm64-gnu

Windows x64

@sylphx/citra-win32-x64-msvc

Missing native → fail closed (no silent TypeScript PDF engine).

Product docs

Doc

Purpose

docs/POSITIONING.md

Strategic positioning

docs/COMPETITIVE.md

Peer anchors and wedge

docs/EVIDENCE_CONTRACT.md

Evidence = result contract

docs/TOOL_SURFACE.md

Few clear tools policy

docs/PRODUCT_INDEPENDENCE.md

This repo is SSOT

docs/IPPB.md

Independent public product bar

docs/PUBLISH.md

npm / git publish status

docs/guide/installation.md

Install & host config

skills/citra/SKILL.md

Agent skill surface

Surfaces (MCP · CLI · SDK)

MCP (default agent path)

npx -y @sylphx/citra

Claude Desktop / Cursor / VS Code / Codex

{
  "mcpServers": {
    "citra": {
      "command": "npx",
      "args": ["-y", "@sylphx/citra"]
    }
  }
}

Dual-era hosts that send server/discover before initialize (e.g. Gemini Antigravity CLI) are supported on stdio.

CLI

npx -y @sylphx/citra --help

SDK

  • @sylphx/citra/sdkCitra (read / search / evidence)

  • @sylphx/citra/pure-rust → low-level client helpers

  • Same tools as MCP: read_pdf · search_pdf · pdf_evidence

  • Requires the platform optional native package (same as MCP)

Install footprint (honest)

Compare full clean installs, not “JS wrapper tarball vs native executable”:

Metric (measured clean install, linux-x64)

Historical TS 3.0.14

Sole-Rust 4.1.0 lineage

Main package on disk

~403 KB

~77 KB

Full node_modules

~82.3 MiB

~24.4 MiB (~3.4× smaller)

Installed files

4,101

20 (~205× fewer)

Production npm deps

PDF.js + MCP TS SDK + more

{} + one platform native

The native binary is multi-megabyte because it is the PDF engine. That is expected — and still a cleaner install than shipping PDF.js + a large JS tree.

Details: installed footprint comparison

Performance (method-bounded)

Controlled same-host linux-x64 dual-mode A/B vs historical @sylphx/pdf-reader-mcp@3.0.14, using registry-installed sole-Rust natives (measured on the 4.1.x lineage; method applies to current sole-Rust packages):

Mode

What it measures

Result

persistent_warm

long-lived server, repeated identical local read_pdf after warm-up

≥ ~10× median latency improvement on all 8 required fixture classes

startup_inclusive

spawn + initialize + one task

large advantage on the same fixtures

persistent_warm includes a process-local cache for identical local path+options. First request in a process still pays full parse cost.

Not a multi-host guarantee. Details: 4.1.0 report · claims policy

Engine note

Current production is a native Rust engine on supported platforms via a thin Node launcher.

Local-first. Five platform packages. One clean install. Fail closed without the matching native.

Unusually formed or broken ToUnicode CMaps are handled without crashing; the release binary is panic-unwind so a worker-thread panic fails the request instead of aborting the process (#608).

Engineering history and recovery pins: docs/migration.md — not the product pitch.


Stop PDF hallucinations. Give agents proof.

npx -y @sylphx/citra

Available Tools

1 tool
read_pdfB

Reads content/metadata from one or more PDFs (local/URL). Each source can specify pages to extract.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_full_textNoInclude the full text content of each PDF (only if 'pages' is not specified for that source).
include_metadataNoInclude metadata and info objects for each PDF.
include_page_countNoInclude the total number of pages for each PDF.
sourcesYesAn array of PDF sources to process, each can optionally specify pages.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool reads content/metadata and allows page specification, but lacks details on permissions, rate limits, error handling, or output format. For a tool with no annotations, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and front-loaded: two sentences that efficiently convey core functionality without waste. Every sentence earns its place by stating the main purpose and a key feature.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 4 parameters with full schema coverage, the description is minimally adequate. It covers the basic action and a feature, but lacks context on behavioral traits, output, or error handling, making it incomplete for a tool with no structured support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all parameters. The description adds minimal value beyond the schema by mentioning 'pages to extract,' which aligns with the 'pages' parameter but doesn't provide additional semantics. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Reads content/metadata from one or more PDFs (local/URL).' It specifies the action (reads), resource (PDFs), and scope (content/metadata, multiple sources). However, it doesn't differentiate from siblings since none exist, so it can't achieve a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides minimal usage guidance. It mentions 'Each source can specify pages to extract,' which hints at when to use page specification, but offers no explicit when/when-not scenarios, prerequisites, or alternatives. With no sibling tools, this is less critical, but the guidance remains basic.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.2/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool 'read_pdf' has a clear and distinct purpose focused on extracting content and metadata from PDFs.

Naming Consistency5/5

Since there is only one tool, naming consistency is inherently perfect. The tool name 'read_pdf' follows a clear verb_noun pattern, which would be consistent if more tools were added.

Tool Count2/5

A single tool is too few for a server named 'PDF Reader MCP Server', as this suggests a broader domain that might include operations like search, annotate, convert, or edit PDFs. The scope feels incomplete with just reading functionality.

Completeness2/5

The tool set is severely incomplete for a PDF reader domain. While 'read_pdf' covers extraction, there are obvious gaps such as searching within PDFs, manipulating pages, adding annotations, converting formats, or handling PDF metadata updates, which are common in such applications.

Maintenance

ActivityActive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/SylphxAI/pdf-reader-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server