Skip to main content
Glama

omniseek_read

Read text from any URL or document file, auto-routing to the right extractor. Returns structured content with outline and media info.

Instructions

Read text from any URL OR document FILE — OmniSeek's single "read this deep" verb. AUTO-ROUTES.

Fully-qualified MCP name: mcp__omniseek__omniseek_read (server name is omniseek; there is no omniseek-eye server).

ROUTING: if target is a local filesystem path OR ends with a document extension (.pdf / .pptx / .docx / .xlsx / .txt / .md / .csv, case-insensitive, a ?query is tolerated) it routes to the DOCUMENT reader (below); otherwise it routes to the URL reader. A LOCAL path ending in .html / .htm also routes to the document reader, which returns the page's extracted text; an http(s) .html URL stays on the URL branch, where the adapters live. start_char / max_chars window the body on BOTH branches (see below); export_media / ocr apply only to the document branch (a URL read has no image-extraction path) and are IGNORED on the URL branch.

URL BRANCH: fetch + normalize ONE URL. Tries each registered adapter until one claims it — a specific article link (a Reddit post, an arXiv paper, a Bluesky post) as a normalized document. arXiv is two-tier by design: an /abs/<id> URL returns abstract-level metadata (title / authors / abstract, a fast lookup), while an /pdf/<id> URL routes to the PDF extractor and returns the WHOLE body (e.g. 2203.02155v1 → 68 pages of full text). Pass the URL whose depth you want. vs the open web: reads ONE specific URL you already have; to FIND open-web pages use WebSearch first, then omniseek_read to normalize the page (a common pairing). The normalized body is WINDOWED by start_char / max_chars (default 24000), exactly like the document branch: a big page (a SEC 10-K/20-F is ~2 MB → ~200k chars, a long article) would otherwise return one blob that overflows the tool channel and is unreadable. When truncated is true, re-call with start_char bumped by returned_chars to page through the rest. A small page (< max_chars) returns whole, truncated=false — unchanged from before. URL branch returns: {"url", "matched": bool, "document": Document as dict | None, "total_chars", "returned_chars", "start_char", "truncated"} (the last four only when matched). On matched:false a reason is added: walled (anti-bot challenge -> retry the source via CDP, e.g. omniseek_search(sources=[...], raw=True, full=True)) vs empty vs blocked, so you can tell "gated, drill it another way" from "genuinely nothing there".

DOCUMENT BRANCH (pptx / docx / xlsx / pdf / txt / md / csv): read the FILE into readable, structured text — the document counterpart of omniseek_transcribe (speech). Free, keyless, cached. WHERE THE FILE LIVES:

  • the operator's machine: scp it to OmniSeek host inbox first — scp "" :omniseek-inbox/ then call with "omniseek-inbox/".

  • Anywhere on the web: just pass the URL (conference slide decks, a shared docx, a PDF). WHAT COMES BACK: outline = per slide/sheet/page {label, chars, media} — the MAP of the whole document, always complete and tiny; text = the readable content ("## Slide 3" / "## Sheet: budget" / "## Page 5" headers), windowed by start_char/max_chars for big docs (truncated=true + total_chars tell you to re-call with start_char to continue); media/media_total = the image inventory per section. THE IMAGE HALF (be honest about it): a figure deck or scanned doc carries its meaning in IMAGES — text extraction alone is NOT the document. Two ways to read it: omniseek_view delivers the figures to your OWN vision in-band (judging the figure is yours); ocr=True here runs OCR over every embedded image and folds the recognized text-in-pixels (scanned page body, chart labels, palette HEX/RGB codes) into the body under a '图中文字 (OCR)' section — mechanical text transcription, NOT figure interpretation, and labeled as possibly imperfect. Use ocr for text-bearing images (scans, labels); use omniseek_view to SEE the figure. Document branch returns: {source, format, title, outline, text, total_chars, returned_chars, start_char, truncated, media_total, media, media_dir, ocr_images?, cached} — or {source, error, inbox_files?}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
ocrNo
targetYes
max_charsNo
start_charNo
export_mediaNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure—and it delivers. It reveals the auto-routing logic, the arXiv two-tier behavior, the windowing/truncation mechanism, the error categories (walled vs empty vs blocked), and the image/OCR handling. It even discloses that OCR is mechanical and possibly imperfect. No contradictions with annotations (none present).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every sentence earns its place given the tool's complexity. It is front-loaded with the core purpose and routing, then organized with clear headers (URL BRANCH, DOCUMENT BRANCH, IMAGE HALF) and bullet-style breakdowns. No redundancy or filler; the density is justified by the need to convey nuanced behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with no output schema, so the description must fully specify return values—and it does. For the URL branch it lists the exact fields returned, including the added 'reason' on matched:false. For the document branch it details outline, text, media, and caching. It also covers the file-transfer mechanism (scp) and the two ways to handle images. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description is the sole source of parameter meaning. It explains target (URL or path), start_char and max_chars (windowing and paging), export_media (document branch only), and ocr (image text extraction). It also clarifies their scoped behavior on each branch. This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific verb+resource: 'Read text from any URL OR document FILE' and immediately distinguishes itself from siblings (omniseek_search, omniseek_transcribe, omniseek_view). The auto-routing between URL and document branches is explicitly defined, so an agent knows exactly what this tool does and how it differs from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit routing rules (file extension vs. URL), tells when to use WebSearch first, when to use omniseek_view for figures, and how to page through truncated results. It also explains when parameters like export_media and ocr are ignored. This is a textbook example of usage guidance with clear alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.