Infobroker
Infobroker is a multi-provider research MCP server that searches, extracts, verifies, cites, and caches information through one unified tool surface.
Web & multi-source search (
search_web) — query one or many engines at once, with automatic provider selection, fallback chains, batched queries, autocomplete, query expansion, deep passage ranking, and a research compile mode.Content extraction (
fetch_page) — pull clean Markdown from any URL via Jina Reader or native/ Wikipedia/ Archive/ arXiv/ Stack Exchange renderers, ask a page a question, crawl same-origin pages, extract structured metadata, and detect last-updated dates.Claim verification (
verify_claims) — run a multi-pass truth-finding loop across independent sources and get confidence-scored confirmed, contested, and unverified findings with provenance.Academic citations (
get_citations) — get BibTeX references with title, authors, year, venue, and URL for scholarly writing.Local knowledge base (
manage_kb) — cache searches and fetches, ingest text or URLs, list/get/delete entries, view stats, archive reports, and manage optional at-rest encryption.Provider intelligence (
inspect_providers) — list provider state, run live health checks, and view quota, latency, uptime, and build/spec identity.Configuration control (
reload_config) — hot-reload provider, rate-limit, and knowledge-base settings without restarting, with optional schema migration.Built-in reliability — 26 providers (21 zero-config), automatic fallback and hedging, per-provider rate limits, content policy enforcement, and persistent quotas that survive restarts.
Enables web search via the Brave Search API, supporting higher throughput and specialized queries when an API key is configured.
Provides web search results from DuckDuckGo as a zero-config provider, used in unified search across multiple backends.
Offers access to archived web content through a dedicated renderer, enabling source-specific extraction from the Internet Archive.
Enables web search via a self-hosted SearXNG instance, providing advanced search capabilities when an API key is configured.
Provides search across Wikipedia articles and dedicated content extraction from Wikipedia pages via its specific renderer.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Infobrokersearch web for recent breakthroughs in fusion energy and verify with multiple sources"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Infobroker
One server. Every source. Research that delivers.
Infobroker is a multi-provider MCP server that unifies web search, structured knowledge, academic, archive, and content-extraction APIs behind a single tool surface. Twenty-one zero-config providers ship in the box — search the web, look up facts, fetch articles — with nothing to configure. Five more providers unlock with API keys or self-hosting. A built-in corroboration engine cross-references independent sources to separate established facts from contested claims. Bundled client skills transform raw research into polished writing. Free first. Privacy always.
North Star
Infobroker is the Bothan Spynet as a tool — a decentralized intelligence network that queries independent sources and routes results through a single, impartial interface. In intelligence-cycle terms, you supply the direction and get the dissemination; the server handles the collection and processing.
Related MCP server: MCP Info Gatherer
Quick Start
cd Infobroker && npm install && npm run startAdd this to your OpenCode config (~/.config/opencode/opencode.json):
{
"instructions": [
"<path-to-Infobroker>/instructions/search-preferences.md"
],
"skills": {
"paths": [
"<path-to-Infobroker>/skills",
"<path-to-opencode-config>/skills"
]
},
"mcp": {
"infobroker": {
"type": "local",
"command": ["node_modules/.bin/tsx", "src/index.ts"],
"cwd": "<path-to-Infobroker>",
"environment": {
"INFOBROKER_CONFIG": "<path-to-Infobroker>/config.json"
}
}
}
}The mcp block starts the server; the instructions and skills blocks
are what activate the bundled client skills. Without them the skills ship
in the repository but stay inert.
Free providers work immediately. API-keyed providers — Brave, Exa, Tavily, Yep — unlock higher throughput and specialized search; self-hosted SearXNG gives full query privacy:
export INFOBROKER_BRAVE_API_KEY="your-key"
export INFOBROKER_EXA_API_KEY="your-key"Requirements: Node.js 20+.
MCP Server
Your research backend. Seven tools, twenty-six providers, one corroboration engine. The complete feature inventory is documented in the feature taxonomy in the spec.
Unified Search
"Search for the location of the second Death Star." "Find scholarly papers on hyperspace travel theories." "Search the latest astromech specs and show me the passages that answer: does the R2 unit pre-date the Clone Wars?"
search_web sends one query to every provider that can answer it. Search
across DuckDuckGo, Wikipedia, academic databases, news, code repositories —
or describe your task and the server picks the best source. Pass an array of
queries to batch several searches in one call. Ask for a deep read and it
fetches the top results and ranks each page's passages against your query,
so you get the specific text that answers the question instead of links.
Failed providers fall back silently through a configurable chain so you get
results, not error messages. Other search tools lock you to one engine;
Infobroker routes every query to the right provider and keeps going when one
fails.
Content Extraction
"Fetch the article on the Battle of Yavin and summarize it." "Get the text of that page about the Death Star plans." "Where, in that report, does it mention the reactor core?"
fetch_page hands any URL to Jina Reader, which renders it as clean
Markdown optimized for LLM consumption. Falls back to native HTTP when
Jina is throttled. Wikipedia and Internet Archive have dedicated
renderers for source-specific extraction. Ask a page a question — pass
question to fetch_page and it returns the passages that answer it, each
scored and ranked, instead of the whole document. Built-in web fetchers
return raw HTML; Infobroker gives you clean, readable content from any
source — ready for summarization or analysis. Fetch also reports the page's
last-updated date when it can determine one, so you know how current your
source is.
Citations
"Give me BibTeX references for papers on hyperdrive field dynamics." "Cite the paper that first described the hyperdrive field equations."
get_citations searches scholarly sources and returns each reference as a formatted
BibTeX entry with its fields — title, authors, year, venue, and URL — ready
to paste into a reference list.
Provider Intelligence
"Which source should I use to research the Death Star's weakness?" "Show me all available sources and their quota status."
The server knows its own capabilities. search_web auto-selects the
best backend for your task, weighing capability, quota, and latency —
or routes by your intent when you ask for privacy, speed, or free-only
sources. inspect_providers surfaces every configured source and drills into a
single provider's uptime and error history. No other search MCP server
gives you operational visibility into every backend.
Multi-Source Verification
"Verify whether the Empire really destroyed Alderaan." "Find the consensus on who fired first — Han or Greedo."
verify_claims runs a multi-pass truth-finding loop: broad
search across your highest-authority providers — search engines,
encyclopedias, and scholarly indexes — then claim extraction,
cross-source reconciliation, and targeted follow-up for gaps, dispatched
in parallel and stopping early once the truth is pinned down. Claims
corroborated across independent sources score high confidence, weighted
by each source's authority; every source is bound to the claim it
supports. Contradictions are surfaced with all perspectives. Gaps
trigger refined queries that broaden to the rest of your providers. It
also remembers: prior findings in your knowledge base participate as
corroborating sources before it queries the network. You get a
structured report — confirmed, contested, and unverified findings — with
source provenance, per-source claims, and confidence scores. Every other
search tool returns a list of links; Infobroker finds the truth and
tells you how sure it is.
Knowledge Base
"Search what you already found about the Rebel Alliance fleet." "Ingest this article so it's cached for next time."
Every search, fetch, and corroboration run is cached in a local knowledge
base. manage_kb checks the cache before hitting external providers — only
falling back to the network when the cached results aren't fresh enough
or relevant enough. Its actions ingest new text or a URL by hand, report
what's cached, and remove content. Content is age-scored, expired on a
freshness schedule, and deduplicated by source. Retrieval runs on your
machine with a configurable in-process embedding model — your content is
never sent to a third party to be embedded. Beyond
the cache, manage_kb archives the reports you generate: ingest with
source_type: "report"
(and default to the knowledge base) and revisit them with manage_kb list and
manage_kb get, or write them to a local directory instead. Each archived report
records its source's last-updated date, so you can compare it against the
live source and refresh only what has actually changed. Other search MCP
servers re-fetch the same facts every session; Infobroker remembers and
reuses what it already found.
Research Pipeline
"Research the construction of the Death Star, then draft a summary." "Fact-check these claims about Darth Vader's origin."
Infobroker doesn't stop at search results. Bundled client skills chain its tools into writing pipelines, routing every request through a solved workflow shape and the writing sub-skills until a finished document comes out the other end. Everything lives in the repository — no external paths or separate install. The full pipeline — the six skills, the workflow shapes, and the escalation path — is detailed in the Skills section. Other search MCP servers produce search results; Infobroker produces finished work.
Operational Visibility
"Show server health." "Hot-reload my config without restarting."
Quota counters persist to disk and survive restarts. Rate limits are
enforced per-provider, not globally. Configuration is hot-reloadable
via reload_config — change providers, adjust chains, or tweak
thresholds without dropping connections, or pass a patch to merge a
change into your user layer with a backup. search_web doubles as
DuckDuckGo query autocomplete. inspect_providers reports the server's build
health and request stats. You always know what your search server is
doing and how much capacity remains.
Skills
The MCP server is one half of the product. The bundled skills are the other. Six client skills ship in the repository — no external dependency, no separate install — and they turn raw research into finished work.
The orchestrator skill (infobroker) opens with a classify gate that
maps your request to a workflow shape: research-and-write, fact-check,
deep-dive, competitive evaluation, literature review, monitoring,
red-team, vetting, or gated analysis. Each shape composes the same
primitives — recall from the knowledge base, search, extract, verify,
write, and cite — into its own sequence and ends with a grep-able
completion token so you can confirm the outcome. Four writing sub-skills
execute the writing phases: summarization condenses findings before
writing, technical-writing drafts reports and docs, proofreading
polishes language, and translation produces multilingual output.
Gated analysis is the escalation shape. When a question is high-stakes or
decision-driving, the classify gate routes to the analysis-loop skill —
a disciplined path with confidence-scored findings, source-reliability
grading, and structured analytic techniques chosen by fit and named with a
rationale — rather than the lighter research-and-write route. It shares the
same primitives and Infobroker tools but runs its own gated workflow, so you
get the rigor without leaving the pipeline.
A single instruction file, search-preferences.md, routes your client
toward these tools: the knowledge base first, external providers only
when the cache falls short. Wire it and the skills directory into your
OpenCode config once — the Quick Start above shows the exact snippet —
and every research request follows the pipeline automatically.
Write your own skill into skills/ to add a workflow shape of your own.
The pipeline diagram lives in references/pipeline-map.md and the
workflow-shape definitions in references/workflows.md. Other search MCP
servers return links; Infobroker ships the writers that turn them into
documented answers.
Providers
Twenty-six providers. Twenty-one work with zero configuration.
Provider | Tier | Type | Key Required |
DuckDuckGo | Built-in | Web search | No |
Jina Reader | Free HTTP | Content extraction | No |
Wikipedia | Free HTTP | Encyclopedia | No |
Wiktionary | Free HTTP | Dictionary | No |
Wikidata | Free HTTP | Structured facts | No |
OpenStreetMap | Free HTTP | Geocoding | No |
Internet Archive | Free HTTP | Historical | No |
arXiv | Free HTTP | Academic | No |
Semantic Scholar | Free HTTP | Academic | Optional |
Stack Exchange | Free HTTP | Code Q&A | Optional |
GitHub | Free HTTP | Code search | Optional |
CORE | Free HTTP | Open access | Optional |
OpenAlex | Free HTTP | Academic | No |
Europe PMC | Free HTTP | Academic | No |
Hacker News | Free HTTP | News | No |
GDELT | Free HTTP | News | No |
SEC EDGAR | Free HTTP | Financial filings | No |
World Bank | Free HTTP | Economic data | No |
Marginalia | Built-in | Small web | No |
Mojeek | Built-in | Independent index | No |
Wiby | Built-in | Small web | No |
Brave Search | Keyed HTTP | Web, News | Yes |
Exa | Keyed HTTP | Semantic | Yes |
Tavily | Keyed HTTP | Synthesis | Yes |
Yep | Keyed HTTP | Web, Semantic | Yes |
SearXNG | Self-hosted | Full privacy | Yes (self) |
Built-in and free-HTTP providers are active out of the box. Keyed providers enable with an API key. Self-hosted providers point at a server you run yourself:
export INFOBROKER_BRAVE_API_KEY="BSA-..."
export INFOBROKER_SEARXNG_URL="http://localhost:8080"Then set "enabled": true in config.json for the provider.
Keyed providers also accept an ordered credential pool via
INFOBROKER_<NAME>_API_KEYS (comma-separated). Infobroker rotates to the
next key when one is rejected or rate-limited, and reports per-key
availability through inspect_providers without ever surfacing key
material.
SearXNG is the only shipped self-hosted provider, and it is optional through and through. Nothing in the server requires it, and nothing is bundled or installed on its behalf — SearXNG runs as a container you operate, and Infobroker queries its JSON endpoint like any other backend. Leave it disabled (the default) and you lose nothing: the privacy-critical chain still serves via DuckDuckGo and Mojeek. Enable it only when you want full query privacy, in which case only your own SearXNG instance sees your queries.
Configuration
Four environment variables tune a deployment. INFOBROKER_CONFIG points
at a different config file (default ./config.json),
INFOBROKER_CONFIG_LOCAL at a user config layer (default
config.local.json), INFOBROKER_<NAME>_API_KEY supplies a keyed
provider's credential, and INFOBROKER_<NAME>_URL points at a self-hosted
provider.
config.json ships with the repository and holds the defaults: which
providers are enabled, their priority in fallback chains, rate limits,
corroboration parameters, and the task-to-provider dispatch table.
Hot-reloadable via reload_config — edit the file and call the tool, or
pass a patch to merge configuration programmatically; changes take
effect without a restart.
Your own overrides live in a separate user layer — config.local.json
in the project directory (or a path you set via INFOBROKER_CONFIG_LOCAL).
This file is git-ignored, so pulling updates from the repository never
overwrites your settings. Values in the user layer take precedence over
the shipped defaults; anything left out falls back to config.json.
config.json carries a schema stamp (config_version). If your user layer
was written for an older schema, or holds a key the current schema no longer
recognizes, the server reports the drift at startup and in every
reload_config response without changing your file. Apply the registered
migrations on demand by calling reload_config with migrate true: the
server first copies your layer to a timestamped *.bak-* file, then updates
it atomically, leaving anything it does not recognize untouched.
The knowledge base ships empty. Retrieval runs on a configurable in-process
embedding model selected by kb.embedding_model. By default the store
writes to a user-scoped path (~/.local/share/infobroker/knowledge-base)
outside the repository, so the content you research and cache stays on your
machine and is never committed. Each deployed instance accumulates its own
store.
Knowledge base encryption
Research reports and cached pages can be sensitive, and the knowledge
base stores them in a single file in your home directory. Enable optional
at-rest encryption by adding a kb.encryption block and supplying a key:
{
"kb": {
"encryption": { "enabled": true, "key_file": "~/.config/infobroker/kb.key" }
}
}The key file (plain, 0600) is the most reliable source across MCP clients
and operating systems; INFOBROKER_KB_KEY (a 32-byte key) or
INFOBROKER_KB_PASSPHRASE (a passphrase) also work. Generate a key with
openssl rand -base64 32. Encryption protects the store and disk-saved
reports from anyone who obtains the files without the key — device theft,
backup or cloud-sync leaks, other local accounts. It does not protect
against a malicious MCP client on the same machine, or malware, which
full-disk encryption covers.
Two rules keep this safe. First, encryption is your opt-in: if the key is missing or wrong, the knowledge base locks and reports an error rather than touching your data — so back up the key (a forgotten key or passphrase means the store is unrecoverable by design). Second, the server never writes a partial file: every save is atomic, and an unrecognized or newer store format is never overwritten.
The manage_kb tool's encryption action is the day-to-day surface for this
journey, and it never echoes secret material — generate_key and backup
return file paths, and rekey reads a key file rather than a raw key.
Enable by generating a key, backing it up, adding the kb.encryption
block, and reloading; the store is encrypted in place immediately.
Disable by removing the block and reloading; the store is decrypted to
plaintext immediately (keep the key available during the transition so the
server can read the store to decrypt it). Recover a locked store with
status to see the state, verify to confirm a candidate key before
committing it, backup to restore a copy of your key file, and rekey to
move to a new key without losing content. After re-keying, point
kb.encryption.key_file at the new key, reload, then run verify again to
confirm the new key opens the store.
infobroker_manage_kb action=encryption operation=generate_key key_file=~/.local/share/infobroker/keys/kb.key
infobroker_manage_kb action=encryption operation=backup key_file=~/.local/share/infobroker/keys/kb.key.bak
infobroker_reload_configTool-surface key operations (generate_key, backup, rekey's target)
are confined to the keys directory: kb.keys_dir when configured, else the
keys sibling of the knowledge base storage path
(~/.local/share/infobroker/keys by default). A path outside that
directory is refused. The kb.encryption.key_file configuration value is
operator-owned and not subject to the confinement.
Add the kb.encryption block to config.local.json before reloading to
enable, or remove it before reloading to disable. When the store is locked,
status, verify, and rekey remain reachable so you can recover without
first unlocking.
Content policy
Infobroker reads the open web and caches what it retrieves, so retrieved
content is assessed against a configurable policy before it is stored (and,
in the strictest mode, before it is returned). The policy flags content
that matches heuristic categories — prompt-injection instructions,
credential phishing, malware/exploit material, and adult content — and can
consult an external assessment service when one is configured. It is on by
default in flag mode: flagged content is still returned to you for
legitimate research, but it is never written to the knowledge base, and
every flag is recorded in the audit trail.
{
"content_policy": {
"mode": "flag",
"threshold": 0.2,
"patterns": { "prompt_injection": ["ignore previous instructions"] }
}
}Modes: off disables assessment; flag (default) returns but never stores
flagged content; block refuses flagged content to the caller. threshold
tunes sensitivity (0–1, default 0.2 — a single match flags). patterns
extend or override the built-in categories per category name. external_url_env
names an environment variable holding the URL of an external assessor, and
external_api_key_env optionally names one holding its bearer key; when the
external service is unreachable the built-in assessment applies.
Security-relevant events — refused network targets, policy flags, config
reloads, encryption transitions, key operations, and quota exhaustion — are
appended to an owner-only audit log at output.audit_log_path (default
~/.local/share/infobroker/audit.log).
Bring your own endpoint
Any HTTP search endpoint can become an Infobroker provider without
touching the source tree. Declare it in config.local.json as a
generic_http provider, then reference it from a dispatch chain:
{
"providers": {
"my_search": {
"tier": "generic_http",
"capabilities": ["web_search"],
"enabled": true,
"priority": 20,
"endpoint": "https://api.example.com/search",
"query_param": "q",
"results_path": "data.items",
"field_map": { "title": "name", "url": "link", "snippet": "summary" }
}
},
"dispatch": { "general_web": ["my_search", "duckduckgo"] }
}The server GETs endpoint?query_param=<query>, walks results_path
(dot-separated into the response JSON), and maps each result to the
common shape using field_map. Add the slug to your config.local.json
override and call reload_config to use it immediately.
Per-provider status and provenance
Two optional keys tune per-provider behavior in config.json:
degraded_latency_ms— a provider whose recent average latency exceeds this many milliseconds is reporteddegradedby theinspect_providershealth action, even while reachable. A globaloutput.degraded_latency_msacts as the fallback when a provider omits its own.resells— settrueon aggregator/reseller backends (search engines that surface other publishers' pages, like DuckDuckGo, Brave, or SearXNG). The server reports each result'soriginal_sourcewhere the backing API exposes one (e.g. Brave'sprofilename); first-party sources (Wikipedia, arXiv) leave it empty because the page is the origin.
Hedged fallback
search_web and fetch_page fall back with a hedge instead of waiting
out a slow provider's full timeout: the primary (first-choice) provider
runs alone for a latency-derived window, then the remaining providers
race and the first result wins. The common path uses one provider call;
the hedge fires only when the primary is slow or failing. fetch_page
additionally prefers the primary renderer in a short grace window so a
marginally slow jina is not displaced by a lower-quality native_fetch.
A renderer whose content is an anti-bot challenge page (a CAPTCHA or
verification interstitial rather than the target page) is treated as a
failed render, so fetch_page falls through to the next renderer instead
of serving the challenge as content.
Tune the window with output.hedge_enabled, hedge_min_delay_ms,
hedge_max_delay_ms, and hedge_grace_ms; set hedge_enabled to
false for the sequential chain. A provider that returns a rate-limit
or anti-bot response is held in a per-provider cooldown (output.rate_limit_cooldown_ms)
so a burst of requests stops re-hammering it, and when a non-general_web
chain exhausts, the server retries the general_web chain before failing.
How It Compares
Tool name | What you're used to | How Infobroker differs |
Built-in | One search engine, one fetch mode, no configuration, no visibility into what backend is used | Twenty-one zero-config providers with a unified tool surface. Choose the right source for each task. Fall back automatically on failure. See every provider's status and quota. |
Raw API calls | Manual HTTP requests, per-provider auth, per-provider response parsing, no fallback, no quota tracking | One interface for every provider. API keys configured once. Results normalized to a common shape. Rate limits and quota tracked automatically. |
Dedicated search APIs | Pay-per-query, vendor lock-in, opaque routing | Free-first design. DuckDuckGo, Wikipedia, and nineteen other providers work with zero configuration. Upgrade paths for Brave, Exa, Tavily, and Yep. Self-hosted SearXNG for full privacy. |
Other search MCP servers | Single-provider focus, no fallback, no corroboration, no writing pipeline | Multi-provider with automatic fallback. Corroboration engine cross-references independent sources. Bundled writing skills transform research into finished documents. |
AI with built-in search | The model picks the search engine, serves stale cache, no reproducibility | You control the provider chain. Queries are reproducible. Fallback behavior is visible. The corroboration engine verifies facts across independent sources. |
Every other search MCP server asks you to pick a provider and trust it. Infobroker gives you a fleet — and picks the right one for each task. When a provider fails, the next one takes over without you noticing. When a claim matters, the corroboration engine finds agreement, contradiction, and gaps. The bundled skills close the loop from raw research to finished writing. One server. Every source. Research that delivers.
Last updated: 2026-10-02.
Contribute
Node.js 20+.
node --version. Get it at nodejs.org.npm install && npm run typecheckBundle your own skill in
skills/to extend the research pipeline.Validate README structure:
npm run validate-readmeVersioning: CalVer (
YYYY.MM.DD).npm run version-bumpstamps today's date into all version references. Pre-commit hooks verify consistency.npm run pushchecks, tags, and pushes.MCP protocol: modelcontextprotocol.io
Providers: DuckDuckGo · Jina Reader · Wikipedia API
Canonical origin: git.gay/flukeatzerocool/Infobroker. This GitHub repository is a read-only mirror.
License
MIT. Free to use, modify, and redistribute. The bundled client skills and instruction files ship under the same license, so the full research pipeline — server, skills, and documentation — is freely reusable in commercial and open-source work alike. Third-party providers remain subject to their own terms and API keys.
Spec
The server is built from a single source specification, infobroker.md
(v2026.10.02), which defines every requirement and the gates that verify it.
Each requirement traces to an implementation file, and npm run check
reconciles the code, the spec, and this README so what is documented is what
the server actually delivers.
Available Tools
7 toolsinfobroker_fetch_pageFetch Page ContentAIdempotent
Fetch a URL and extract clean content via a renderer (Jina Reader by default; native-HTTP, Wikipedia, Internet Archive, arXiv, and Stack Exchange alternatives). Use when you have a URL and need readable text, passages ranked against a question, or the page's last-updated date. Do NOT use for topic search (use infobroker_search_web) or cross-source claim verification (use infobroker_verify_claims). question returns passages ranked against it, sized by passage_size and capped by max_passages; crawl bounds the same-origin crawl to config caps; extract adds JSON-LD, OpenGraph, and microdata; renderer selects the backend, and native_fetch is the keyless fallback when Jina is throttled or anti-bot challenged. max_length truncates the returned text only; a truncated page is written in full to a temp file. Makes external HTTP calls and needs no API key. Fetched pages auto-index into the knowledge base unless the content policy flags them (see infobroker_manage_kb): flag mode returns without storage, block mode refuses. Unreachable URLs return an [ERROR] envelope with remediation; success returns an [OK] envelope with status, provider, results, and meta.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to fetch: a single URL, or up to five URLs fetched in parallel | |
| crawl | No | Bounded same-origin crawl: recursively fetch same-origin pages up to config caps (default off) | |
| extract | No | Return structured metadata (JSON-LD, OpenGraph, microdata) alongside the content (default off) | |
| question | No | Question to extract ranked passages for, instead of returning the whole page | |
| renderer | No | Renderer: jina (default), native_fetch, wikipedia, internet_archive, arxiv, or stack_exchange | |
| max_length | No | Maximum characters to return (default 50000) | |
| detect_date | No | Detect and report the page's last-updated date (default from config) | |
| max_passages | No | Number of passages to return (default from config) | |
| passage_size | No | Target words per passage (default from config) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false) leave room, and the description fills it substantially: external HTTP calls with no API key, native_fetch as the keyless fallback when Jina is throttled or anti-bot challenged, auto-indexing into the knowledge base with content-policy flag/block semantics, temp-file write on truncation, and the [OK]/[ERROR] envelope shapes. This is well beyond what annotations convey and does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then usage/exclusions, then parameter behavior, then envelope/indexing behavior — a sensible ordering with no filler sentences. It is dense and somewhat overstuffed (many semicolon-chained clauses), but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing both the [OK] and [ERROR] envelopes and their fields. Combined with the mutation/indexing side effects and keyless-fallback behavior, an agent has everything needed to invoke this 9-parameter tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics: the question/passage_size/max_passages interaction, max_length truncating only returned text while the full page is written to a temp file, and renderer acting as backend selection with native_fetch as fallback. Some clauses (crawl, extract) largely restate the schema descriptions, keeping it short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Fetch a URL and extract clean content via a renderer') and immediately characterizes the renderer backends. It explicitly names the two sibling tools it is not (infobroker_search_web for topic search, infobroker_verify_claims for claim verification), so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the triggering condition ('Use when you have a URL and need readable text, passages ranked against a question, or the page's last-updated date') and gives explicit exclusions with named alternatives ('Do NOT use for topic search... or cross-source claim verification'). Both the when and the when-not are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_get_citationsGet CitationsARead-onlyIdempotent
Return academic references for a query as BibTeX citations with title, authors, year, venue, and URL. Use when scholarly writing needs a reference list. Do NOT use for general web search (use infobroker_search_web) or contested-claim verification (use infobroker_verify_claims). query is a natural-language topic; max_results sets the reference count, and larger values take longer and span more sources. Needs no API key when at least one scholarly source is reachable; if every source fails it returns an [ERROR] envelope with remediation, and queries respect per-provider rate limits. Returns an [OK] or [ERROR] JSON envelope.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query | |
| max_results | No | Maximum references to return (1-30, default 8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, and non-destructive, but the description adds behavior they cannot express: no API key required when a scholarly source is reachable, per-provider rate limits, and an [ERROR] envelope with remediation when every source fails. This is meaningful operational context for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage rules, then parameter and failure behavior — a sensible ordering. It is dense (four sentences covering a lot), but every clause carries distinct information rather than restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description still specifies the return format (BibTeX fields) and the envelope contract ([OK]/[ERROR] with remediation). For a two-parameter, read-only tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description earns extra credit by clarifying that `query` is a natural-language topic and that larger `max_results` values 'take longer and span more sources' — a latency/coverage tradeoff the schema does not state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return academic references for a query') plus the exact output shape (BibTeX with title, authors, year, venue, URL). An agent can distinguish this from search_web and verify_claims without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger ('when scholarly writing needs a reference list') and two exclusions with the sibling to use instead (infobroker_search_web for general web search, infobroker_verify_claims for contested-claim verification). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_inspect_providersInspect ProvidersARead-onlyIdempotent
Inspect configured search providers: list their state, run a live health check, or report build and spec identity. Use when searches return empty or slow results and you want provider status, quota, or latency, or when choosing which backend to trust. Do NOT use to search (use infobroker_search_web) or to read a page (use infobroker_fetch_page). Read-only: it never modifies configuration, providers, or stored data. list snapshots provider state locally and status filters it; health runs a live outbound probe against provider (required for the health action; subject to that provider's rate limits, so it can be slow); spec reports build and spec identity locally. Returns an [OK] or [ERROR] JSON envelope with status, provider, and results.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Operation to perform | |
| status | No | Filter for list action | |
| provider | No | Provider slug (required for health) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint, so the safety profile is covered; the description goes further by disclosing that `health` performs a live outbound probe subject to provider rate limits and may be slow, that `list`/`spec` are local snapshots, and that the tool never modifies configuration. That is exactly the operational context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, the negative guidance follows, and the per-action breakdown is dense but each clause carries distinct information. The middle sentence is long and semi-colon chained, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description tells the agent what comes back ('an [OK] or [ERROR] JSON envelope with status, provider, and results'), covers every action, the required parameter, and the read-only guarantee. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage the baseline is 3, but the description adds real meaning: `provider` is required specifically for the `health` action and is a slug, and `status` acts as a filter on the list snapshot rather than an action. It only stops short of documenting the concrete `active`/`all` enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb ('Inspect') and resource ('configured search providers') and enumerates the three sub-operations (list state, live health check, build/spec identity). It explicitly distinguishes itself from infobroker_search_web and infobroker_fetch_page, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a positive trigger ('when searches return empty or slow results', 'when choosing which backend to trust') and explicit exclusions with named alternatives ('Do NOT use to search (use infobroker_search_web) or to read a page (use infobroker_fetch_page)'). This is the full when/when-not/alternative pattern.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_manage_kbKnowledge BaseADestructive
Manage the local knowledge base: search cached content, ingest text or URLs, list/get/delete entries, view stats, and manage at-rest encryption. Use when you need to archive a generated report (ingest with source_type 'report' and save_to 'kb'), revisit stored content, or manage encryption keys. Do NOT use for fresh external search (use infobroker_search_web) or to fetch a new page (use infobroker_fetch_page). The delete action is destructive and cannot be undone; a lost encryption key makes the store unrecoverable by design. Parameter semantics: action selects the operation; query for search, text/url plus title and collection for ingest, source_url for get/delete, and operation with key_file for the encryption sub-actions; source_type and freshness_tier tag an ingest and filter search/list; save_to defaults to 'kb', format to 'markdown', max_results to 8; last_updated is auto-detected from fetched URLs when omitted. Responses carry an [OK] or [ERROR] envelope: search/list/get return entries, stats returns counts, delete reports the chunks removed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to fetch and index (for ingest action) | |
| text | No | Raw text to index (for ingest action) | |
| query | No | Search query (for search action) | |
| title | No | Document/report title | |
| action | Yes | Operation to perform | |
| format | No | File format for disk save | markdown |
| save_to | No | Where to save (ingest action): kb, disk, or both. Default kb | |
| key_file | No | Path to a key file (generate_key writes here; rekey reads the target key from here). Never pass the secret itself — only a file path. | |
| operation | No | Sub-operation for the 'encryption' action | |
| collection | No | Collection name | |
| source_url | No | Source URL filter (get/delete action) or identity for ingest | |
| max_results | No | Maximum results to return (1-50, default 8) | |
| source_type | No | Source type (tag on ingest; filter on search/list) | |
| last_updated | No | Source last-updated date to store (e.g. '2026-08-24'); auto-detected from fetched URLs when omitted | |
| freshness_tier | No | Freshness tier tag on ingest (e.g. 'report', 'evergreen') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, but the description adds the specifics annotations cannot: delete is irreversible, and a lost encryption key makes the store unrecoverable by design. It also discloses the [OK]/[ERROR] response envelope. It stops short of covering auth/permission requirements or any concurrency behavior, so a 4 rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, then the destructive warning, then parameter mapping — a sensible priority order with almost no filler. The closing parameter sentence is dense and list-like, slightly harder to parse than the rest, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, 7-action tool with no output schema, the description covers action routing, destructive risk, and return shape ([OK]/[ERROR], entries for search/list/get, counts for stats, chunks removed for delete). An agent has everything needed to select and invoke it without opening the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real cross-field semantics the schema cannot express: which parameters apply to which value of `action`, the source_type/freshness_tier tag-vs-filter duality, and that last_updated is auto-detected from fetched URLs. It does restate a few defaults already present in the schema (save_to, format, max_results), which keeps it out of the top band.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (local knowledge base) and enumerates the concrete operations (search, ingest, list/get/delete, stats, encryption), so the agent knows exactly what the tool covers. It explicitly distinguishes itself from siblings by naming infobroker_search_web and infobroker_fetch_page as the tools for fresh external work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives positive triggers (archive a generated report with source_type 'report' and save_to 'kb', revisit stored content, manage encryption keys) and explicit exclusions with the alternative tool named for each. This is the when/when-not/alternatives pattern at full strength.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_reload_configReload ConfigurationADestructiveIdempotent
Re-read the configuration file and apply provider, rate-limit, and knowledge-base changes without restarting; active connections are preserved. Use when you have edited config.json or config.local.json, or when passing patch to change configuration programmatically. Do NOT use to inspect configuration or provider state (use infobroker_inspect_providers). patch deep-merges a partial config into the user layer (config.local.json), rejecting unknown top-level keys, backing up the previous layer, and validating the merged result before writing; migrate true backs up and applies registered user-layer migrations before reloading. The call re-reads the startup configuration source and reports any user-layer schema drift. If the new configuration is invalid, the previous configuration stays active and an error is returned. Returns an [OK] or [ERROR] JSON envelope.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | Partial configuration deep-merged into the user layer (config.local.json) before reloading; unknown top-level keys are rejected and the previous layer is backed up | |
| migrate | No | Back up and apply registered user-configuration-layer migrations before reloading (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the annotations (destructiveHint=true, idempotentHint=true, openWorldHint=true) by disclosing the write/merge sequence: deep-merge into config.local.json, rejection of unknown top-level keys, backup of the previous layer, validation before writing, migration ordering, and the failure semantics (invalid config leaves the previous config active and returns an error). It also names the response envelope. It does not mention auth requirements or whether the backup file location is exposed, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage routing, then behavioral detail; every sentence conveys a distinct fact and none is filler. It is dense (long semicolon-chained clauses) and slightly heavy for a two-parameter tool, but the mutation risk justifies the depth.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description states the return contract ([OK] or [ERROR] JSON envelope), covers the failure path and rollback behavior, and describes both parameters' effect on disk state. Nothing essential for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics beyond the schema: patch is defined as a deep-merge into the user layer with unknown-key rejection and pre-write validation of the merged result, and migrate is described in terms of backup plus registered migration application ordered 'before reloading'. That ordering and validation detail is not in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (re-read/apply) plus resource (configuration file) and the exact classes of changes applied (provider, rate-limit, knowledge-base) with the key constraint 'without restarting; active connections are preserved'. It also differentiates itself from the sibling infobroker_inspect_providers, so an agent can route without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger ('Use when you have edited config.json or config.local.json, or when passing patch'), an explicit exclusion ('Do NOT use to inspect configuration or provider state') and the named alternative to use instead. Both the positive and negative conditions are stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_search_webWeb SearchARead-onlyIdempotent
Search the web, encyclopedia, academic, and code sources through one interface with automatic provider selection and a fallback chain. Use when you need broad or batched search, query autocomplete (suggest), query expansion (expand), ranked passages from the top pages (deep), or a multi-variant research compile that deep-reads each variant (research). Do NOT use for a URL you already have (use infobroker_fetch_page), for high-stakes claim verification (use infobroker_verify_claims), for stored-report recall (use infobroker_manage_kb search), or for academic citations (use infobroker_get_citations). Caches results in the local knowledge base, enforces per-provider rate limits, and needs no API key for the free providers. Returns a JSON envelope prefixed [OK] or [ERROR] with status, provider, results, and meta.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Read the top results and return each page's ranked passages against the query | |
| page | No | Results page number (default 1) | |
| query | Yes | Search query: a single string, or up to five strings searched in parallel | |
| expand | No | Return query-expansion strings instead of search results | |
| region | No | ISO region/country code (e.g. 'us-en', 'DE') | |
| suggest | No | Return query-autocomplete strings instead of results (default false) | |
| priority | No | Route by intent: speed, quality, privacy, or free_only | |
| provider | No | Provider slug (auto-select if omitted) | |
| research | No | Research compile: fan out to derived query variants and deep-read each, returning passages grouped by variant (default off) | |
| time_range | No | Restrict results to day, week, month, or year | |
| max_results | No | Number of results to return (1-30, default 8) | |
| safe_search | No | Safe-search filtering: on, off, or strict (default on) | on |
| content_type | No | Source kind to search: docs, issue, changelog, blog, spec, or all (default all) | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses caching, rate-limit enforcement, no-API-key behavior, and the [OK]/[ERROR] JSON envelope. These details align with the readOnlyHint, openWorldHint, idempotentHint, and destructiveHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact but information-dense, using clear 'Use when' and 'Do NOT use' sections without redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, it names the return envelope and its top-level fields, describes side effects, and gives enough context for an agent to choose and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 13 parameters are described, with defaults and enums explicitly covered; the query parameter's string-or-array(max 5) behavior is explained, and mode flags like deep/expand/suggest/research are clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the tool searches web, encyclopedia, academic, and code sources through one interface, with a clear 'Use when...' list covering broad/batched search and modes. It also names sibling tools for disambiguation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use it and when not to, pointing to infobroker_fetch_page, verify_claims, manage_kb, and get_citations for exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infobroker_verify_claimsVerify ClaimsA
Verify a contested claim against independent sources and return confidence-scored findings. Use when a claim is high-stakes or contested and you need agreement, contradiction, and gaps surfaced with confidence scores. Do NOT use for simple lookups or broad search (use infobroker_search_web) or citation formatting (use infobroker_get_citations). The loop recalls prior findings from the knowledge base, then searches providers and writes findings back; max_iterations bounds the refinement passes and per-provider rate limits apply. query states the claim plainly; priority routes the pool by intent (speed, quality, privacy, free_only); providers restricts the pool and omitting it uses the full dispatch chain; confidence_threshold sets the bar for confirmation, and findings below it are reported unverified. Returns an [OK] or [ERROR] envelope listing confirmed, contested, and unverified findings with per-source claims and confidence scores.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query | |
| priority | No | Route the corroboration pool by intent | |
| providers | No | Optional array of provider slugs to limit the search to | |
| max_iterations | No | Maximum search-refinement passes (1-10, default 5) | |
| confidence_threshold | No | Minimum confidence to report a finding as confirmed (0-1, default 0.8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag readOnlyHint=false, openWorldHint=true, idempotentHint=false; the description goes well beyond them by disclosing the actual write-back loop ('searches providers and writes findings back'), the KB recall step, per-provider rate limits, and the max_iterations refinement bound. It also explains how below-threshold findings are downgraded to 'unverified', which no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and use/do-not-use guidance are front-loaded, and the operational detail follows in a single dense block. It is long, but nearly every clause contributes routing or behavioral information; a little trimming of overlapping phrases would make it tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return contract — an [OK]/[ERROR] envelope listing confirmed, contested, and unverified findings with per-source claims and confidence scores. Combined with the loop/rate-limit notes, an agent has everything needed to invoke and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real semantics beyond the field docs: providers 'omitting it uses the full dispatch chain', priority routing intents (speed, quality, privacy, free_only), and that confidence_threshold determines which findings are reported unverified. max_iterations and query get less added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'Verify a contested claim against independent sources and return confidence-scored findings' — and immediately separates this tool from infobroker_search_web (broad search) and infobroker_get_citations (citation formatting). An agent can distinguish it from every sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit inclusion condition ('high-stakes or contested') and explicit exclusions with named alternatives ('Do NOT use for simple lookups or broad search (use infobroker_search_web) or citation formatting (use infobroker_get_citations)'). Routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.4- Changed
infobroker_reload_config1 field changed- added
Input schema / properties / patchAdded value: +{ + "additionalProperties": {}, + "description": "Partial configuration deep-merged into the user layer (config.local.json) before reloading; unknown top-level keys are rejected and the previous layer is backed up", + "propertyNames": { + "type": "string" + }, + "type": "object" +}
1 tool update
v0.1.3- Changed
infobroker_reload_config1 field changed- added
Input schema / properties / migrateAdded value: +{ + "default": false, + "description": "Back up and apply registered user-configuration-layer migrations before reloading (default false)", + "type": "boolean" +}
3 tool updates
v0.1.2- Changed
infobroker_fetch_page2 fields changed- added
Input schema / properties / crawlAdded value: +{ + "default": false, + "description": "Bounded same-origin crawl: recursively fetch same-origin pages up to config caps (default off)", + "type": "boolean" +} - added
Input schema / properties / extractAdded value: +{ + "default": false, + "description": "Return structured metadata (JSON-LD, OpenGraph, microdata) alongside the content (default off)", + "type": "boolean" +}
- Added
infobroker_search_web - Removed
infobroker_web_search
9 tool updates
v0.1.1- Removed
infobroker_corroborate - Changed
infobroker_fetch_page9 fields changed- added
Input schema / properties / detect_dateAdded value: +{ + "description": "Detect and report the page's last-updated date (default from config)", + "type": "boolean" +} - added
Input schema / properties / max_length / descriptionAdded value: +"Maximum characters to return (default 50000)" - added
Input schema / properties / max_passagesAdded value: +{ + "description": "Number of passages to return (default from config)", + "type": "number" +} - added
Input schema / properties / passage_sizeAdded value: +{ + "description": "Target words per passage (default from config)", + "type": "number" +} - added
Input schema / properties / questionAdded value: +{ + "description": "Question to extract ranked passages for, instead of returning the whole page", + "type": "string" +} - added
Input schema / properties / renderer / descriptionAdded value: +"Renderer: jina (default), native_fetch, wikipedia, internet_archive, arxiv, or stack_exchange" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "description": "URL to fetch", + "type": "string" + }, + { + "description": "Multiple URLs to fetch in parallel (max 5)", + "items": { + "type": "string" + }, + "maxItems": 5, + "type": "array" + } +] - changed
Input schema / properties / url / descriptionPrevious value: -"URL to fetch"New value: +"URL to fetch: a single URL, or up to five URLs fetched in parallel" - removed
Input schema / properties / url / typeRemoved value: -"string"
- Added
infobroker_get_citations - Added
infobroker_inspect_providers - Removed
infobroker_kb - Added
infobroker_manage_kb - Removed
infobroker_providers - Added
infobroker_verify_claims - Changed
infobroker_web_search12 fields changed- added
Input schema / properties / content_type / descriptionAdded value: +"Source kind to search: docs, issue, changelog, blog, spec, or all (default all)" - added
Input schema / properties / deepAdded value: +{ + "default": false, + "description": "Read the top results and return each page's ranked passages against the query", + "type": "boolean" +} - added
Input schema / properties / expandAdded value: +{ + "default": false, + "description": "Return query-expansion strings instead of search results", + "type": "boolean" +} - added
Input schema / properties / max_results / descriptionAdded value: +"Number of results to return (1-30, default 8)" - added
Input schema / properties / page / descriptionAdded value: +"Results page number (default 1)" - added
Input schema / properties / priority / descriptionAdded value: +"Route by intent: speed, quality, privacy, or free_only" - added
Input schema / properties / query / anyOfAdded value: +[ + { + "description": "Search query", + "type": "string" + }, + { + "description": "Multiple queries to search in parallel (max 5)", + "items": { + "type": "string" + }, + "maxItems": 5, + "type": "array" + } +] - changed
Input schema / properties / query / descriptionPrevious value: -"Search query"New value: +"Search query: a single string, or up to five strings searched in parallel" - removed
Input schema / properties / query / typeRemoved value: -"string" - added
Input schema / properties / safe_search / descriptionAdded value: +"Safe-search filtering: on, off, or strict (default on)" - added
Input schema / properties / suggest / descriptionAdded value: +"Return query-autocomplete strings instead of results (default false)" - added
Input schema / properties / time_range / descriptionAdded value: +"Restrict results to day, week, month, or year"
6 tool updates
v0.1.0- First observed
infobroker_corroborate - First observed
infobroker_fetch_page - First observed
infobroker_kb - First observed
infobroker_providers - First observed
infobroker_reload_config - First observed
infobroker_web_search
TDQS
Scored across 7 tools
Each tool targets a distinct operation (fetch, search, verify, citations, KB management, provider inspection, config reload), and the descriptions contain explicit 'Do NOT use X, use Y instead' cross-references that preempt the few plausible overlaps like fetch_page vs search_web deep mode.
All seven tools share the infobroker_ prefix and a consistent verb_noun pattern: fetch_page, search_web, verify_claims, get_citations, manage_kb, inspect_providers, reload_config.
Seven tools is well-scoped for a research/information-broker agent, with no redundant tools and every capability (search, fetch, verify, citations, KB, config) earning its place.
Core research lifecycle is covered: search, fetch, verify, cite, store/recall in KB, and manage providers/config. Minor gap: no dedicated report generation or export tool despite KB ingest referencing 'report' source types.
Maintenance
Related MCP Connectors
Live AI-native web search with citations. One tool for every MCP client. Flat per-request pricing.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Free web search for AI agents. No API key required. Hosted MCP in active development.
Your memory, everywhere AI goes. Build knowledge once, access it via MCP anywhere.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceZero-auth multi-source research MCP server that enables web search, reading URLs, PDFs, GitHub repos, and querying Hacker News, Stack Overflow, Semantic Scholar, and YouTube transcripts without API keys.11Apache 2.0
- FlicenseAqualityBmaintenanceEnables searching and gathering information from multiple online sources (web, Twitter/X, Telegram, GitHub, Hugging Face, arXiv) through a unified MCP interface.12-
- AlicenseNot gradedqualityFmaintenanceMCP server enabling local-first web search, fetch, extract, and caching with citeable excerpts, no API key required. Supports research workflows for agents and apps.301 npm1MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to perform unified web research through a single MCP server, including search, page fetching, recursive crawling, document parsing, YouTube transcript extraction, and deep multi-query research.3-