Skip to main content
Glama
cluefinch

Cluefinch MCP

Official

research_collect

Read-only

Gather bounded evidence from selected URLs or search queries, extract literal passages, and report gaps to support multi-source research.

Instructions

Collect bounded evidence from several selected or searched sources.

USE THIS for multi-source evidence gathering when you have explicit URLs,
several search formulations, or both. It searches/fetches candidate documents,
extracts literal passages and reports acquisition gaps. It does NOT crawl links
found inside those documents. If a strong source must first be navigated to a
chapter, appendix or related page, use web_links and then pass the selected
URLs here or to web_fetch.

Supply queries=["variant one", "variant two"] and/or urls=["https://..."].
topic only ranks excerpts; topic alone does not search. Explicit URLs consume
the source budget first. If they already fill max_sources, no search is run.
Remaining candidate slots are distributed round-robin across query variants.
For domain/exclude_domains filtering, use web_search first and pass selected
result URLs here.

Returns raw material, not synthesis: query_variants, sources and gaps. Each
source has stable source_id, final URL, source_type, content_hash,
text_truncated and literal excerpts with Unicode offsets. BM25 selects passages;
source_type and ranking are not credibility judgments.

For more context around an excerpt, copy excerpt.expand.arguments into the
indicated tool; URL, offset, budget and expected_content_hash are already
supplied. content_changed returns no stale-coordinate slice.

Always inspect gaps before judging coverage. They can report failed/empty
searches, unresponsive engines, blocked/unreadable pages, insufficient text,
redirect duplicates, omitted candidates and truncation. Empty gaps do not prove
topic completeness. Expected request failures return {error, hint}; individual
source failures can coexist with successful sources.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsNoExact source URLs to fetch first; useful after web_search with domain/exclude_domains. Supply urls and/or queries. At most 100 input URLs; max_sources still controls the fetch budget.
topicNoOptional topic used only to rank excerpts alongside queries. Does not initiate a search without queries.
queriesNoYour search variants, normally 2-3; at most 10. The server does not generate queries. May be combined with explicit URLs.
languageNoLanguage/locale forwarded to search; unused in URL-only mode.
time_rangeNoSearch recency: day, week, month or year. Does not assert a source publication date.
max_sourcesNoCandidate fetch budget: default 5, configured cap 10. Failed/duplicate candidates can leave fewer successful sources. Explicit URLs take priority.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.4

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover only readOnlyHint and openWorldHint; the description carries the rest, disclosing budget semantics (explicit URLs consume the source budget first, round-robin distribution over remaining slots, no search run if URLs fill max_sources), failure semantics (empty gaps don't prove completeness; {error, hint} for expected failures; source failures coexisting with successes), and non-judgment caveats (BM25 ranking and source_type are not credibility judgments). That is substantive behavior beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the primary use rule before mechanics; almost every sentence carries a constraint, routing rule, or failure caveat. It is nonetheless long and dense for a collection tool, and a few clauses (excerpt expansion, content_changed) sit below the critical usage block.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description usefully summarizes the return shape (query_variants, sources, gaps, stable source_id, offsets, content_hash) and the excerpt.expand handoff, plus the content_changed signal. Combined with the usage and failure guidance, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already documented, but the description adds cross-parameter semantics the schema cannot express: URLs are consumed before queries, remaining candidate slots are distributed round-robin across query variants, and topic alone ranks excerpts without initiating a search. Some points (max_sources cap, topic-does-not-search) restate the schema, so it is not a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (collect bounded evidence from selected/searched sources) and immediately bounds scope: it searches/fetches candidates, extracts literal passages, reports gaps, and explicitly does NOT crawl links found inside documents. It is clearly distinguishable from web_search, web_fetch and web_links, all three of which are named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions (multi-source gathering with explicit URLs, several query formulations, or both), a when-not rule (no crawling of in-document links), and explicit routing to siblings: use web_links first if a source must be navigated to a chapter/appendix, and use web_search first for domain/exclude_domains filtering. Alternatives and the conditions selecting them are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools