okular-mcp
Integrates with KDE's Okular PDF viewer, reading the current document, page, mouse selection, and saved annotations, and writing highlights, notes, and file attachments back into the PDF.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@okular-mcpsummarize the highlights in this PDF"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
okular-mcp
An MCP server that lets an LLM session read and annotate a PDF alongside you in Okular: it sees which document and page you are on, the text under your mouse, and the highlights and notes you have made, and it can highlight passages back and drive the viewer to a page.
The PDF file is the shared state. Okular saves its annotations as standard PDF annots; this server reads and writes those same annots with PyMuPDF, so nothing lives in a private database and the notes travel with the file.
Tools
Tool | What it does |
| Every open Okular window: document path, current page (1-based), page count |
| Text currently selected with the mouse ( |
| Numbered text lines of a page range, defaulting to the page shown in Okular; |
| The paragraph around the current mouse selection (or |
| Links on a page: anchor text and line, then the target page and line for internal links (citations, sections, figures) or the URL, so the model can follow a citation with |
| The bookmark tree (sections with page numbers), a map of the paper |
| All saved annotations: page, type, author, highlighted text, note |
| Mark every occurrence of |
| Comment icon in the margin: level with |
| Embed a text file ( |
| Contents of a file-attachment annotation, decoded as text |
| Delete an annotation by the |
| Jump Okular to a page, opening the document first if needed. Meant for "show me where that is defined", not for the model to navigate on its own |
Marks written by the server carry a distinct author (llm, override with
OKULAR_MCP_AUTHOR) and a light-blue colour (blue for underlines and notes, red for strikeout), so they
are told apart from yours. Every writing tool takes color: a palette name (blue,
yellow, green, orange, pink, purple, red, cyan, grey), #rrggbb, or
r,g,b in 0–1, so a session can keep colours consistent by usage (one for
definitions, another for open questions).
Related MCP server: PDF Agent MCP
Requirements
Linux with KDE's Okular on the session D-Bus, plus:
qdbus(Qt tools; installed with Okular on most distributions)wl-clipboardon Wayland, orxclip/xselon X11, for the selection toolsPython 3.10+;
mcpandpymupdfare pulled in as dependencies
Install
With uv:
uv tool install git+https://github.com/david-wang-0/okular-mcpThen register it with your MCP client. Claude Code, user scope:
claude mcp add okular -s user -- okular-mcpCodex CLI:
codex mcp add okular -- okular-mcpAny other client: run okular-mcp as a stdio server.
Workflow
Read in Okular. Highlight with the highlighter tool, add notes to highlights.
Save (Ctrl+S). Okular keeps annotations in memory until then, and the server only sees what is in the file.
Ask the LLM to collect the highlights (
annotations), discuss the page you are on (viewer_state,page_text), or explain the phrase under your mouse in its paragraph (get_selection,context). It can follow citations and section references itself (links, thenpage_textat the target line) and get a map of the paper (outline).Let it mark back (
mark_text;style="underline"withnote_in_marginis the least intrusive), answer a note of yours as a threaded reply (add_notewithreply_to), attach a Markdown or Mermaid file next to a passage (attach_file, read back withread_attachment), or clean up its own marks (remove_annotation). The server saves incrementally and asks Okular to reload, so the change appears in the viewer.
Gotchas
Save before any writing tool (
mark_text,add_note,attach_file,remove_annotation). If Okular has unsaved annotations when the file changes on disk, it offers a reload and the unsaved marks are lost if you accept. The tool descriptions tell the model to ask you to save first.Tabs. With several documents in tabs of one window, Okular's D-Bus interface reports the first-opened tab, not the active one (KDE bug 502482). Use separate windows if you rely on
viewer_state.Selection text is Okular's text layer, returned verbatim: hyphenation and column order can be messy.
Minimal client environments. Some MCP clients (Codex, the Python SDK's client) start servers with only
HOME,PATH,SHELLandUSER. The server then recoversDBUS_SESSION_BUS_ADDRESSandWAYLAND_DISPLAYfromXDG_RUNTIME_DIR(or/run/user/<uid>), so no per-clientenvblock is needed on a systemd desktop.Lines are reconstructed. A PDF has no lines;
page_textandcontextuse PyMuPDF's block and line grouping, so a two-column page reads column by column and equations or footnotes can land in odd places.Highlighted text is recovered from the annotation's quads, so it can pick up a neighbouring word on tight line spacing.
Development
uv venv && uv pip install -e '.[test]'
.venv/bin/pytestokular_mcp/viewer.py is the only Okular-specific module (windows,
current, goto, reload); selection.py and pdf.py are viewer-agnostic,
so another viewer backend can be added without changing the tool surface.
License
MIT
Available Tools
13 toolsadd_notePDF: margin note or replyA
Put a comment icon in the margin, or reply to an annotation.
Comment icon in the margin: level with near_text, or as a threaded reply to the annotation reply_to (an xref from annotations, e.g. to answer the user's own note), else at the top of the page. color as in mark_text. Same save-Okular-first caveat as mark_text.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | ||
| page | No | ||
| path | No | ||
| color | No | ||
| reply_to | No | ||
| near_text | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the write/safety profile is covered. The description adds one genuine behavioral fact — the 'save-Okular-first' caveat — but only by deferring to mark_text, so the agent must already know that sibling to act on it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the two modes and back-loading the anchor elaboration. Dense and waste-free, though the backtick-heavy cross-references cost some readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation, and the annotation set covers safety. But for a non-idempotent mutation with 6 undocumented parameters, the description omits whether near_text and reply_to are mutually exclusive and what happens to page/path defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the schema gives no semantics. The description compensates for three of them (near_text, reply_to, color) with clear roles, but leaves note, page, and path entirely unexplained, including which anchor fields are mutually exclusive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (place a comment icon / reply to an annotation) on a specific resource, and explicitly splits the tool into two modes: a margin note anchored via near_text, or a threaded reply via reply_to. An agent can distinguish it from mark_text and annotations by the named anchors alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives selection logic for the anchor: use near_text for a margin-level icon, reply_to (an xref from annotations) to answer an existing note, otherwise it falls back to the top of the page. It names sibling tools but never states the when-not condition versus mark_text, which is the closest alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
annotationsPDF: highlights and notesBRead-only
All highlights, notes and attachments saved in the PDF.
All annotations in the PDF (default: the open document): page, type, author, highlighted text, note. Only annotations saved to the file are visible.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds genuinely useful behavior beyond that: only annotations saved to the file are visible, and the default scope is the currently open document. It doesn't cover pagination or result ordering, but the key visibility caveat is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the first sentence ('All highlights, notes and attachments saved in the PDF') is largely restated by the second sentence ('All annotations in the PDF'), so one of the two sentences is partly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value formatting needn't be spelled out, and the description still names the fields. Safety is covered by annotations. The remaining gap is the undocumented 'page'/'path' parameters, which is minor for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither 'page' nor 'path' is documented in the schema. The description's word 'page' appears in a list of return fields, not as a filter explanation, and 'path' is only obliquely hinted at via 'default: the open document.' The parameters are largely left for the agent to infer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource and operation: retrieving all highlights, notes and attachments saved in a PDF, including the fields returned (page, type, author, text, note). It is distinct enough from write siblings like add_note/mark_text/remove_annotation, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is a scoping hint that the default target is "the open document," but no explicit when-to-use guidance and no routing to alternatives such as get_selection or page_text. An agent gets no criteria for choosing this tool over its read siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_filePDF: attach a text fileB
Embed a text file (Markdown, Mermaid) as a paperclip annotation.
Embed a text file (e.g. notes.md, diagram.mmd) as a paperclip annotation in the margin, level with near_text or at the top of the page; or replace the file of an existing attachment replace_xref. Okular saves it via the icon's context menu. Same save-Okular-first caveat as mark_text.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| note | No | ||
| page | No | ||
| path | No | ||
| color | No | ||
| content | Yes | ||
| near_text | No | ||
| replace_xref | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, but the description adds genuinely useful operational context: Okular persists the attachment via the icon's context menu, and there is a 'save-Okular-first' caveat shared with mark_text. The replace_xref mode (overwriting an existing attachment's file) is disclosed, which is meaningful given destructiveHint=false. It stops short of explaining auth/permission requirements or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key constraint is front-loaded, but the second sentence restates 'Embed a text file ... as a paperclip annotation' almost verbatim from the opening line, wasting space. The remaining clauses earn their place by conveying placement and save behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the mutation/save caveat is covered. But with 8 parameters at 0% schema coverage and no explanation of the required ones (name, content) or the placement interplay between page/near_text/path, the description is only minimally sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 8 parameters, so the description carries the full burden — yet it only adds meaning for `near_text` (placement, level with text) and `replace_xref` (target existing attachment). The six remaining parameters (name, note, page, path, color, content) are left completely undetermined by both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Embed') plus resource ('a text file ... as a paperclip annotation') and gives concrete file-type examples (Markdown, Mermaid). An agent can distinguish this from add_note and mark_text, though the description never explicitly contrasts with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when this tool applies by describing placement modes ('level with `near_text` or at the top of the page') and a replace flow ('replace the file of an existing attachment `replace_xref`'), and points to `mark_text` for a shared caveat. However it never states when to choose attach_file over add_note or read_attachment, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contextPDF: paragraph around the selectionARead-only
The paragraph around the user's mouse selection.
The paragraph around the user's current mouse selection (or text) on the page shown, with neighbours paragraphs before and after and the line numbers, so you can see what they are pointing at without reading the whole page.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| path | No | ||
| text | No | ||
| neighbours | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes a safe read, so the description's job is to add beyond that — it does by disclosing what comes back (the paragraph, `neighbours` paragraphs before/after, and line numbers) and that `text` can substitute for the live mouse selection. It does not cover fallback behavior when no selection exists, but the added context is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and short, but the opening sentence 'The paragraph around the user's mouse selection' is a near-verbatim restatement of the title and is then repeated at the start of the longer second sentence. One sentence of duplication is wasted space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no prose explanation, and the safety profile is covered by annotations. Still, with four parameters at 0% schema description coverage and no clarification of what `page` or `path` do (especially with zero required parameters), an agent cannot fully infer how to target the correct document.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the schema documents nothing and the description must compensate. It explains `neighbours` (paragraphs before and after) and `text` (an alternative to the live selection), covering two of four parameters, but leaves `page` and `path` completely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a precise resource — the paragraph containing the user's mouse selection, plus neighbouring paragraphs and line numbers. It implicitly contrasts with page-level retrieval via 'without reading the whole page', but never names the distinction against siblings like get_selection or page_text explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a rationale for use ('so you can see what they are pointing at without reading the whole page'), which implies 'use when you need around-the-selection context cheaply'. However there is no explicit guidance on when to prefer get_selection or page_text instead, and no stated preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_selectionOkular: mouse selectionARead-only
Text the user has selected with the mouse, with page and lines.
Text currently selected with the mouse ('primary') or on the clipboard ('clipboard'), with the document path, page and line numbers when it is found on the page Okular shows (line numbers match page_text and context).
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | primary |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=true annotation already establishes this as a safe read. The description adds real context beyond that: selection may come from the mouse or clipboard, and line numbers are only populated when the text is found on the current page. It still omits what happens when nothing is selected or how the two sources behave differently.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The two sentences overlap heavily: the first ('Text the user has selected with the mouse, with page and lines') largely restates the second. It is short but not maximally economical; the lead could be tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the safety profile is carried by annotations. Given the one-parameter, read-only scope, the description covers what an agent needs. Minor gaps remain around the default source behavior and empty-selection handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'source' parameter has only a default and no documentation. The description compensates by naming the accepted conceptual values ('primary' for mouse selection, 'clipboard' for the clipboard), which is the key semantic an agent needs to set that parameter correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: retrieving the text the user has selected with the mouse, plus page and line numbers. It implicitly distinguishes itself from siblings like page_text and context by noting the line numbers match those tools' output. It does not name an alternative it is not, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clarifies the two retrieval modes ('primary' mouse selection vs 'clipboard') and that line numbers appear only when the text is found on the currently shown page, which implies usage. However, there is no explicit statement of when to reach for this instead of page_text, context, or viewer_state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gotoOkular: show a page to the userA
Jump Okular to a page for the user.
Jump Okular to page of path (default: current document; opens it if needed).
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the bar is lower. The description adds real behavioral context by noting the target defaults to the current document and that the document will be opened if needed — a genuine side effect not captured by annotations — but says nothing about page-number validity, range errors, or what happens if the jump fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the precise semantics of the parameters. The opening sentence slightly restates the second, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a two-parameter navigation tool whose annotations cover the safety profile, the description supplies the essentials (target page, default document, open-if-needed); only page-numbering and error behavior are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden and largely does: it explains that `page` is the destination and that `path` defaults to the current document (opening it if necessary). It does not specify page numbering (1-based?) or out-of-range behavior, leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (jump) and resource (Okular page/document), so an agent can tell it apart from siblings such as page_text, viewer_state, or outline. It is clear and concrete, though it never explicitly contrasts itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: 'for the user' indicates this is the navigation/display action, and the parenthetical clarifies the current-document default. There is no explicit when-to-use guidance or reference to alternatives like viewer_state or page_text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linksPDF: links on a pageARead-only
Links on a page with their targets, to follow citations yourself.
Links on a page (default: the page shown): anchor text and line, then the target page and line for internal links (citations, sections, figures) or the URL. Follow one with page_text(page=..., lines="-...").
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already declares the safe read profile, and an output schema exists to describe return values. The description does add the meaningful default that omitting page means 'the page shown' and clarifies the internal-vs-external target distinction, but there is little else (no pagination, ordering, or volume limits).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly two sentences, front-loaded with the core purpose, and the follow-up call is a compact inline example. Minor redundancy from restating 'Links on a page' twice, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with an output schema, the description covers purpose, return structure, the page default, and the workflow continuation. The only material gap is the unexplained 'path' parameter, which no other structured field covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry both parameters. It compensates for 'page' by explaining the null default ('default: the page shown') and shows it used in the page_text example, but 'path' is never explained, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Links on a page with their targets') and states exactly what is returned: anchor text and line, plus target page/line for internal links or the URL. It implicitly separates itself from siblings like page_text by casting itself as the citation-following step rather than the text-fetch step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context ('to follow citations yourself') and an explicit follow-up recipe: reach the target with page_text(page=..., lines="<target_line>-..."). It stops short of stating when NOT to use it or naming an alternative tool by name, so it is clear but not fully routable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_textPDF: highlight or underline textA
Highlight, underline, squiggle or strike out a passage, with a note.
Mark every occurrence of text on a page (default: the one shown) as highlight, underline, squiggly or strikeout, saved into the PDF under a distinct author; quote enough words to be unique. A note sits on the mark, or with note_in_margin on a comment icon in the margin replying to it (least intrusive). color: blue, yellow, green, orange, pink, purple, red, cyan, grey, #rrggbb or r,g,b in 0-1 (default light blue; red for strikeout). The user must save Okular first or their unsaved marks are lost.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| page | No | ||
| path | No | ||
| text | Yes | ||
| color | No | ||
| style | No | highlight | |
| note_in_margin | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-readonly, non-idempotent, non-destructive mutation, but the description adds substantial context beyond them: marks are saved under a distinct author, the note placement model (on the mark vs a margin icon replying to it), and critically 'The user must save Okular first or their unsaved marks are lost.' That save-state warning is exactly the kind of behavioral disclosure annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and no filler sentences, but the remaining content is packed into one dense paragraph, mixing parameter details and a warning. Slightly denser than ideal, though every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the mutation behavior, note model, and save requirement are well covered. The only gap is that `path` semantics are never explained, leaving the agent to infer the target file.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it documents text, note, page (default the shown page), color (full enum list plus default and strikeout override), style, and note_in_margin. Only `path` is left unexplained, so the coverage is strong but not total.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (highlight, underline, squiggle, strike out) on a clear resource (a text passage in the PDF), and the title 'PDF: highlight or underline text' reinforces it. An agent can distinguish this from add_note, which adds a standalone note, without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable guidance: mark every occurrence of `text`, default to the shown page, and 'quote enough words to be unique.' It also frames note_in_margin as the least intrusive option. It stops short of explicitly routing between siblings like add_note or remove_annotation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outlinePDF: table of contentsCRead-only
The document's table of contents.
The document's bookmark tree (sections with page numbers), a map of the paper.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds almost nothing behavioral beyond that – it doesn't say what happens when a document has no bookmarks, whether output nests, or pagination/limits. The mention of 'sections with page numbers' is content, not behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the second sentence largely repeats the first ('table of contents' vs 'bookmark tree... a map of the paper'), so it is somewhat redundant rather than tightly informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is a simple read. Still, for a tool with an undocumented parameter and no edge-case behavior, the description leaves gaps that an agent would need to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single `path` parameter has 0% schema description coverage and is never mentioned in the description. The phrase 'The document's' weakly implies a current/default document, but it does not explain what path selects or that omitting it uses the active document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('table of contents' / 'bookmark tree (sections with page numbers)') that an agent can distinguish from read-only siblings like page_text, links, or annotations. It lacks a clear verb and slightly restates the title, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. Nothing routes the agent between outline, page_text, links, or context, so the agent must infer usage from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_textPDF: page text, numbered linesBRead-only
Numbered text lines of a page range.
Numbered text lines of a page range (default: the page shown in Okular); lines such as "12-30" narrows to a window. Line numbers match context.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| page | No | ||
| path | No | ||
| lines | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered by structured data. The description adds genuinely useful behavior beyond that: the default page source (Okular's current page) and the `lines` window syntax ("12-30"). It stops short of describing output shape, but that is largely covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The remainder is tight and front-loaded, but the first sentence is repeated word-for-word, so a full sentence does not earn its place. The default-page fact is also spliced awkwardly into a parenthetical mid-sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with an output schema and no required parameters, the definition covers the essentials (default page, window filter, line-number alignment with `context`). It is still incomplete on `path` vs `page` resolution and on what the numbered output looks like when no page range is given.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and four parameters exist, so the description carries the full burden. It partially compensates by explaining `lines` ("12-30") and implying page-range pairing via `page`/`to` defaults, but `path` and the exact meaning of `to` are left entirely undocumented, leaving half the parameters opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb+resource ('Numbered text lines of a page range'), which is distinct from siblings like get_selection or annotations, and it explicitly ties line numbers to the `context` tool. However, the opening sentence is duplicated verbatim, which dilutes the otherwise clear statement of purpose and adds no differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions the default behavior (the page shown in Okular) and that line numbers align with `context`, which is a mild routing hint, but there is no explicit when-to-use, when-not-to-use, or comparison against siblings such as get_selection or context. An agent must infer when page_text is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_attachmentPDF: read an attached fileBRead-only
Read an attached file back as text.
Contents of a file-attachment annotation (an xref from annotations with a file field), decoded as text.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| xref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description confirms the read-only, decoded-to-text nature of the operation. It adds the useful detail that the content is decoded as text and that only file-attachment xrefs qualify, but says nothing about failure behavior for invalid/non-file xrefs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the precise semantic constraint. No filler, though the second sentence is somewhat dense/jargon-laden.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary. For a two-parameter read tool the description covers the key linkage to `annotations`, but the undocumented `path` parameter leaves a gap in what an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains the required `xref` well (an xref from `annotations` carrying a `file` field), but the optional `path` parameter is completely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read an attached file back as text') and pins the scope to file-attachment annotations, which separates it from the generic `annotations` sibling. It does not explicitly contrast with the reverse operation `attach_file`, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it tells the agent the xref must come from `annotations` with a `file` field, which is a de facto precondition. There is no explicit when-not guidance or named alternative for adding attachments.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_annotationPDF: delete an annotationBDestructive
Delete an annotation by xref.
Delete an annotation by the xref reported by annotations (default: the open document) and save. Same save-in-Okular-first caveat as mark_text.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| xref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation semantics are covered. The description adds that it saves and defaults to the open document, plus a 'same save-in-Okular-first caveat as mark_text', which adds context but only by cross-referencing a sibling tool rather than spelling out the caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical routing information (xref source, save behavior) is presented early, but the first two sentences duplicate the same 'Delete an annotation by xref' information, so not every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a destructive tool, though, deferring the crucial Okular-save caveat to 'same ... as mark_text' leaves an agent without the actual precondition in-context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It does address both parameters—xref's origin (from `annotations`) and path's default (the open document)—but gives no format detail for `path`, leaving half the surface only partially compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Delete an annotation') and the addressing key ('by xref'), which cleanly distinguishes it from write-oriented siblings like mark_text and add_note. It falls short of a 5 only because the opening sentence is a near-verbatim restatement of the title before the substantive second sentence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an implied workflow by telling the agent to use the xref 'reported by `annotations`', which is useful routing. However, it never states when to use this versus alternatives (e.g., modifying or hiding an annotation) nor any preconditions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
viewer_stateOkular: open documents and pagesARead-only
Which PDF and page Okular is showing.
Open Okular windows: document path, current page (1-based) and page count.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes a safe read operation, and the description is consistent with that. It adds a short summary of what is returned (document path, 1-based current page, page count) and notes open Okular windows, but does not disclose broader behavioral traits such as permissions or limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, front-loaded sentences with no filler. It states the core purpose first and then summarizes the returned fields, which is appropriately sized for a simple read-only state query.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema and readOnlyHint annotation, the description is nearly complete for correct invocation: it identifies the resource and summarizes the payload. It lacks sibling routing or usage guidance, but that is not required for a zero-parameter read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero input parameters, so there are no parameter semantics for the description to add. The empty schema is self-explanatory, and the description appropriately focuses on the state being queried.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource—Okular's current PDF/page state—and enumerates the returned fields (document path, current page, page count). It is clear enough to distinguish this from generic state getters, though it does not explicitly contrast with siblings like page_text or goto.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance or mention of alternatives. The description implies querying the current viewer state, but does not say when to choose this over goto, page_text, or other sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
add_note - First observed
annotations - First observed
attach_file - First observed
context - First observed
get_selection - First observed
goto - First observed
links - First observed
mark_text - First observed
outline - First observed
page_text - First observed
read_attachment - First observed
remove_annotation - First observed
viewer_state
TDQS
Scored across 13 tools
Each tool targets a distinct resource or action: viewer state, selection, page text, context around selection, links, outline, annotations, annotation creation/reading/removal, and navigation. Slight conceptual overlap exists between context/page_text and mark_text/add_note, but the descriptions clearly delimit them.
All names use snake_case, but the pattern mixes noun phrases (viewer_state, annotations, outline) with verb_noun forms (get_selection, mark_text, remove_annotation, goto). It remains readable, but there is no single predictable convention across the set.
13 tools is well within the ideal range and matches a PDF reading/annotation server's scope. Each tool represents a distinct operation, and none feels redundant or excessively granular.
The surface covers reading, navigation, annotation creation/reading/deletion, and attachments. Minor gaps remain: no explicit save or full-text search operation, and annotations cannot be modified in place, but the core lifecycle is covered.
Maintenance
Related MCP Connectors
Parse, extract, split, and ask over digital PDFs (text layer, no OCR) from Cursor and Claude.
Generate PDFs from templates via AI chat. Works with Claude, ChatGPT, Cursor, and any MCP client.
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Merge, split, extract, rotate, reorder and stamp PDF pages from your AI chat, all offline.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to search and query PDF documents through a local RAG system with vector embeddings. Provides semantic document search capabilities while keeping all data stored locally without external dependencies.-
- FlicenseAqualityCmaintenanceEnables AI agents to efficiently process large local and online PDFs through selective extraction of text, images, and metadata. It provides tools for content search and document outline navigation to optimize context window usage.715-
- AlicenseBqualityDmaintenanceEnables PDF manipulation including text, images, annotations, form fields, page operations, and metadata through natural language.16MIT
- AlicenseAqualityBmaintenanceEnables AI agents to edit LaTeX documents with live PDF preview, commenting, and visual editing via the Model Context Protocol.781 npmAGPL 3.0