Skip to main content
Glama

zim_search

Read-only

Search offline ZIM archives via full-text, title, or prefix autocomplete modes. Pick the right mode for keyword snippets, exact titles, or typeahead suggestions.

Instructions

Search a ZIM archive — three modes, one tool.

EXTRACT the search intent before calling. Pick the mode that matches what the user actually needs; the wrong mode silently returns the wrong shape of results.

MODES (pass one as mode):

  • "fulltext" (default) — Xapian BM25 search with optional namespace / content_type filters. Use for queries with multiple keywords ("history of Rome", "Tesla electricity") or when the caller wants snippets, not just titles. Cross-archive via cross_file=True.

  • "title" — Exact / typo-tolerant title lookup. Returns titles matching the query (case ladder + suggestion expansion + Levenshtein-1). Use when the caller knows the article name and wants to confirm it exists or find near matches ("find article titled Detroit"). Single-archive applies Z3/Z4/OPP-1 promotion; cross-archive (cross_file=True) returns raw matches (promotion is per-archive).

  • "suggest" — Prefix autocomplete via libzim SuggestionSearcher. Returns title candidates only — no snippets, no body. Use for typeahead-style completion ("prefix Det"). Does NOT support cross_file=True (per-archive).

ALIASES: callers may say "search", "find", "lookup", or "autocomplete". All route through THIS tool — pick the matching mode.

PARAMETERS: query REQUIRED. Plain terms, all AND-ed (drop one to widen); AND/OR/NOT, quotes and wildcards are not parsed (matched as literal words). mode One of {"fulltext", "title", "suggest"}. Default "fulltext". zim_file_path Optional. Omit to auto-select the single loaded archive, or use cross_file=True to fan out. cross_file Default False. Set True to fan out across every loaded archive (modes "fulltext" and "title" only; "suggest" rejects this with invalid_combination). namespace Only valid in mode="fulltext". Restricts search to one ZIM namespace letter (e.g. "C" for content). Rejected in title/suggest modes. content_type Only valid in mode="fulltext". Restricts search to one MIME bucket (e.g. "text/html"). limit Max results; cap 50 title/suggest/cross_file, 100 filtered, else 1000. offset Pagination offset (default 0); single-archive fulltext only. Next page: offset + page_info.source_consumed (else returned_count). cursor Unsupported; page fulltext via offset.

RESPONSE: fulltext rows carry path, title, snippet; title rows carry path, title, score — pass path as entry_path to zim_get. cross_file=True nests hits under results[].result.results. suggest items are {text, path, type}. Title-mode _meta.promotion_applied is True when a candidate was hoisted; False with a hint when cross-archive blocked it.

ERRORS: Returns a ToolErrorPayload on: - mode="suggest" with cross_file=True (invalid_combination) - limit out of range, negative offset (invalid_limit, invalid_offset) - missing archive when zim_file_path is required but cross_file is False and auto-selection fails

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNofulltext
limitNo
queryYes
cursorNo
offsetNo
namespaceNo
cross_fileNo
content_typeNo
zim_file_pathNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv3.2.5

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint annotation by explaining mode-specific response shapes, cross-file nesting, error payloads, limit caps, pagination semantics, unsupported cursor, and title-mode promotion behavior. It gives a clear model of what the tool will do and what can go wrong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: mode selection, parameter details, response shapes, and error cases are all needed for a tool of this complexity. It is well-structured with clear headers, front-loads the most important guidance, and avoids repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With nine parameters, three modes, no output schema, and no parameter descriptions in the schema, the description must be highly self-sufficient. It covers result shapes, pagination, cross-file nesting, limit caps, and error conditions, leaving no critical gap for an agent deciding how to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries full responsibility for parameter semantics. It explains every parameter: query term handling, mode enum, namespace/content_type constraints, cross_file behavior, zim_file_path auto-selection, limit/offset rules, and cursor unsupported. This is comprehensive and highly actionable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool searches a ZIM archive and immediately distinguishes three search modes (fulltext, title, suggest). The verb-resource pairing is explicit, and the mode breakdown makes it easy to understand what the tool does versus the related sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs the agent to extract search intent before calling and provides detailed when-to-use guidance for each mode, including examples and exclusions (e.g., suggest does not support cross_file=True). Aliases are mapped, and the wrong-mode consequence is disclosed, which is strong usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.