research_collect
Gather bounded evidence from selected URLs or search queries, extract literal passages, and report gaps to support multi-source research.
Instructions
Collect bounded evidence from several selected or searched sources.
USE THIS for multi-source evidence gathering when you have explicit URLs,
several search formulations, or both. It searches/fetches candidate documents,
extracts literal passages and reports acquisition gaps. It does NOT crawl links
found inside those documents. If a strong source must first be navigated to a
chapter, appendix or related page, use web_links and then pass the selected
URLs here or to web_fetch.
Supply queries=["variant one", "variant two"] and/or urls=["https://..."].
topic only ranks excerpts; topic alone does not search. Explicit URLs consume
the source budget first. If they already fill max_sources, no search is run.
Remaining candidate slots are distributed round-robin across query variants.
For domain/exclude_domains filtering, use web_search first and pass selected
result URLs here.
Returns raw material, not synthesis: query_variants, sources and gaps. Each
source has stable source_id, final URL, source_type, content_hash,
text_truncated and literal excerpts with Unicode offsets. BM25 selects passages;
source_type and ranking are not credibility judgments.
For more context around an excerpt, copy excerpt.expand.arguments into the
indicated tool; URL, offset, budget and expected_content_hash are already
supplied. content_changed returns no stale-coordinate slice.
Always inspect gaps before judging coverage. They can report failed/empty
searches, unresponsive engines, blocked/unreadable pages, insufficient text,
redirect duplicates, omitted candidates and truncation. Empty gaps do not prove
topic completeness. Expected request failures return {error, hint}; individual
source failures can coexist with successful sources.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | Exact source URLs to fetch first; useful after web_search with domain/exclude_domains. Supply urls and/or queries. At most 100 input URLs; max_sources still controls the fetch budget. | |
| topic | No | Optional topic used only to rank excerpts alongside queries. Does not initiate a search without queries. | |
| queries | No | Your search variants, normally 2-3; at most 10. The server does not generate queries. May be combined with explicit URLs. | |
| language | No | Language/locale forwarded to search; unused in URL-only mode. | |
| time_range | No | Search recency: day, week, month or year. Does not assert a source publication date. | |
| max_sources | No | Candidate fetch budget: default 5, configured cap 10. Failed/duplicate candidates can leave fewer successful sources. Explicit URLs take priority. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||