open-agent-connector
Summary: open-agent-connector is a zero-dependency MCP server that lets Claude delegate chores (search, page reading, drafting, image generation) to cheaper OpenAI-compatible helper models, keeping Claude in charge while saving its tokens.
ask_agent— hand a self-contained text task (summarising, drafting, translating, brainstorming) to a helper; useautofor failover,allto run every text model in parallel, or an exact model id.web_search— get a short grounded answer plus source URLs for docs, errors, versions or news; falls back to clearly labelled⚠️ NOT A LIVE SEARCHmodel knowledge when no live search is available.web_fetch— read a URL; passquestionand a helper reads the full page and returns only the answer, otherwise you get capped cleaned page text.generate_images— generate images across configured image models (fallback to first success, or all in parallel in judge mode), save candidates to disk, and review them viapaths,judge(vision model pre-ranks), orinline.list_helper_models— inspect configured role/model ids and what the backend advertises.Cross-cutting: any OpenAI-compatible endpoint behind a gateway, token savings via digests and
OAC_MAX_RESULT_CHARScapping, adoctorsmoke-test command, andinitwiring that adds a delegation policy toCLAUDE.md— with helpers never touching your repo.
Connects to a local Ollama instance through its OpenAI-compatible endpoint to serve as the helper model backend for text tasks (e.g. ask_agent, page digests, drafting).
Connects to OpenAI's API (or any OpenAI-compatible endpoint) to delegate work to helper models: chat completions for text tasks and page digests, image generation via /images/generations, web search via the Responses API web_search tool, and model listing via /models. Used for tasks like drafting text, reading web pages, running searches, and generating/ranking candidate images.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@open-agent-connectorsearch the web for the latest Node.js LTS version and summarize it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
open-agent-connector
Keep Claude as the lead engineer and hand the chores to cheaper models, to cut Claude Code token costs.
open-agent-connector is a zero-dependency MCP server. It gives
Claude (Claude Code, Claude Desktop) a small set of tools that hand work to helper models behind
any OpenAI-compatible endpoint: web search, page reading, text drafting and image generation.
Claude decides when to delegate, checks what comes back, and makes every final call. The helpers
spend their tokens and Claude saves its own.
┌──────────────────────────────────────┐
You ──▶ │ Claude (lead: plans, decides, codes) │
└───────────────────┬──────────────────┘
│ MCP over stdio (Claude chooses when to call)
┌───────────────────▼──────────────────┐
│ open-agent-connector (this repo) │ Node.js, zero dependencies
└───────────────────┬──────────────────┘
│ HTTPS + Bearer key, OpenAI-compatible API
┌───────────────────▼──────────────────┐
│ Any gateway / provider │ 9router · LiteLLM · OpenRouter ·
│ /chat/completions /images/... │ OpenAI · Ollama · …
└──────────────────────────────────────┘Why
Claude stays in charge. If you point Claude Code's base URL at another model, Claude is gone. With an MCP server, Claude keeps the controls and only calls helpers when it decides to.
Fewer Claude tokens. Helpers read the long pages, search results and drafts. Claude only gets a short digest, capped by
OAC_MAX_RESULT_CHARS.Image models with fallback or competition. By default the image models are tried in order and the first success is used. Turn on judge mode and they all draw in parallel while a cheap vision model ranks the results.
Not tied to a backend. It talks standard OpenAI endpoints. Models are plain ids from your config, so you can mix vendors behind one gateway.
Nothing to install. Run it with
npx -y github:HTB95/open-agent-connector. It has no dependencies and no build step.
Related MCP server: Nexus MCP
Tools exposed to Claude
Tool | What it does | Backend calls |
| Hands a self-contained text task to a helper. |
|
| Runs a live search and returns a short answer with source URLs. | Native search endpoint (opt-in) → Responses API |
| Reads a URL. If you pass | Native fetch endpoint (opt-in), otherwise a local fetch, then |
| Generates images and saves them to disk. Default: tries |
|
| Shows the configured roles and the model ids the backend advertises. |
|
The review option of generate_images controls how Claude sees the candidates:
paths(default in fallback mode): you get file paths, and Claude opens the ones it wants with its Read tool.judge(default in judge mode): a vision-capable helper scores the candidates first, so Claude reads a short ranking.inline: the images are embedded in the tool result. This uses the most tokens.
Quick start
Requirements: Node.js ≥ 18.17, git, and an OpenAI-compatible endpoint with an API key.
Let Claude do it. Set OAC_BASE_URL and OAC_API_KEY yourself (never paste the key into the
chat), then tell Claude Code:
Set up open-agent-connector in this project by following https://github.com/HTB95/open-agent-connector/blob/main/docs/AGENT_SETUP.md
Claude runs init, lets doctor suggest model ids from your backend, smoke-tests and commits.
Or by hand:
# 1. Point the connector at your backend (example: OpenAI directly)
export OAC_BASE_URL="https://api.openai.com/v1"
export OAC_API_KEY="sk-..."
export OAC_TEXT_MODELS="gpt-5-mini"
export OAC_IMAGE_MODELS="gpt-image-1"
export OAC_SEARCH_MODEL="gpt-5-mini"
# 2. Smoke-test the backend outside Claude (costs no Claude tokens).
# With only BASE_URL and API_KEY set, doctor prints suggested OAC_TEXT_MODELS / OAC_IMAGE_MODELS.
npx -y github:HTB95/open-agent-connector doctor --text --search --fetch --image
# 3. Wire a project (writes .mcp.json, .claude/settings.json, CLAUDE.md, .gitignore)
cd your-project
npx -y github:HTB95/open-agent-connector init
git add -A && git commit -m "Wire Claude helpers"Restart Claude Code and run /mcp. You should see helpers · connected. Then try:
Create a hero image for a coffee-shop landing page with the helpers and copy it to
public/hero.png.
For per-user install, Claude Code on the web (cloud), Claude Desktop and troubleshooting, see docs/SETUP.md.
Configuration
All settings are environment variables. init writes a .mcp.json that forwards them with
${VAR:-} expansion, so the file never contains a secret.
Variable | Default | Description |
| required | OpenAI-compatible base URL, including |
| (empty) | Sent as |
| auto-pick | Comma-separated chat model ids. |
| (empty) | Comma-separated image model ids, in fallback order. |
|
|
|
| first text model | Vision-capable chat model used by |
| (empty) | Provider ids for a native search endpoint ( |
|
| Path of that native search endpoint. |
| (empty) | A model that supports the OpenAI Responses API |
| (empty) | Provider ids for a native fetch endpoint ( |
|
| Path of that native fetch endpoint. |
|
| Where candidate images are saved. |
|
| Maximum characters any tool returns to Claude. |
|
| Larger images are never inlined. |
|
| HTTP timeouts. |
Notes:
The image parameters
quality,backgroundandoutput_formatare sent only to OpenAI-style image models: ids containinggpt-image,gpt-…-imageordall-e. Other models getprompt,sizeandn.response_format: "b64_json"is left out forgpt-image-*models, because OpenAI documents it as unsupported there.Claude Code gives MCP tools a limited time to run. Image generation can take 30–120 s, so set
MCP_TOOL_TIMEOUT=300000in the environment that launches Claude Code.
Backend examples
These are example settings. Check yours with doctor before you rely on them.
OAC_BASE_URL=http://localhost:4000/v1
OAC_API_KEY=sk-litellm-master-or-virtual-key
OAC_TEXT_MODELS=gemini-flash,gpt-mini # model_name values from your config.yaml
OAC_IMAGE_MODELS=gpt-image,imagen
OAC_SEARCH_MODEL=gpt-mini # only if that deployment supports Responses web_searchOAC_BASE_URL=http://localhost:20128/v1 # use the port your dashboard shows
OAC_API_KEY=sk-...
OAC_TEXT_MODELS=ag/gemini-3.5-flash-medium,cx/gpt-5.5
OAC_IMAGE_MODELS=cx/gpt-image-2.5,ag/gemini-3.1-flash-image
OAC_SEARCH_PROVIDERS=antigravity # 9router POST /v1/search
OAC_SEARCH_MODEL=cx/gpt-5.5 # fallback: Responses web_search
OAC_FETCH_PROVIDERS=jina-reader # optional, only if connected in 9routerOAC_BASE_URL=https://api.openai.com/v1
OAC_API_KEY=sk-...
OAC_TEXT_MODELS=gpt-5-mini
OAC_IMAGE_MODELS=gpt-image-1
OAC_SEARCH_MODEL=gpt-5-miniOAC_BASE_URL=https://openrouter.ai/api/v1
OAC_API_KEY=sk-or-...
OAC_TEXT_MODELS=google/gemini-2.5-flash,openai/gpt-5-minigenerate_images needs an OpenAI-style /images/generations endpoint. Leave OAC_IMAGE_MODELS
empty here unless your gateway provides one.
OAC_BASE_URL=http://localhost:11434/v1
OAC_TEXT_MODELS=qwen3:8bIf you use a gateway that pools consumer subscriptions, it is up to you to follow each provider's terms of service.
How the token savings work
Digest, don't dump.
web_searchandweb_fetchwithquestionreturn a short answer, not the raw page.Hard cap. Every tool result is cut at
OAC_MAX_RESULT_CHARS, with a note that it was cut.One image by default, cheap judging when asked. Fallback mode pays for one generation per request. In judge mode,
review: "judge"turns N images into a few lines of text, so Claude doesn't have to view each image.A delegation policy.
initadds a short policy toCLAUDE.mdtelling Claude when to delegate. Without one, Claude rarely calls helper tools on its own.Small tool surface. Five tools with short schemas, because tool definitions are re-sent on every turn.
Who does what
The policy splits work by kind, not by size:
Helpers (information gathering, production) | Claude (code and reasoning) |
Web search, finding and reading docs | Reading and writing code |
Summarising long pages, changelogs, search results | Debugging, design, architecture |
Translating, drafting prose, test-data boilerplate | Reviewing and deciding what to keep |
Generating images | Anything that needs the repo's context |
Delegating still has a cost: Claude writes the prompt and reads the answer. So the one exception is work where the prompt costs more than the job, such as a fact Claude already knows.
No measured before/after numbers are published yet. When the backend reports usage, ask_agent
prints each helper's token in/out next to its answer, so you can compare a delegated task against
doing it in Claude. If you measure real savings, a PR with the numbers and the method is welcome.
How it compares
There are other MCP servers that let Claude ask another model for an opinion, and routers that swap the model behind Claude Code. This project focuses on chores rather than second opinions:
Image generation with model fallback, or several models in parallel with a judge model so Claude only opens the top one or two images.
web_fetchwith aquestion: the helper reads the page and Claude only gets the answer.Model ids per role (
text,image,judge,search) on any OpenAI-compatible gateway.A ready delegation policy that
initwrites intoCLAUDE.md, plus a one-command setup that works the same locally and in Claude Code cloud sessions.
Security
Don't commit keys. Keys belong in your shell profile or in the cloud environment settings.
.mcp.jsononly holds${VAR}references.If your gateway is reachable from the internet, put it behind HTTPS and require an API key. Without TLS, the bearer key travels as plain text.
Helpers never touch your repository. They only see what Claude puts in a prompt, and the delegation policy tells Claude never to put secrets or
.envcontents there.See SECURITY.md to report a vulnerability.
Limitations
Only the stdio transport is supported. The claude.ai web chat (remote MCP connectors) is not. A remote MCP version (for example on a serverless platform) is an idea under consideration and has not been tried. It would keep keys on the server and skip the
npxdownload per session, butgenerate_imageswould have to return image URLs (for example from object storage) instead of local file paths.The automated tests run against a mock gateway (
test/mock-router.mjs), not real providers. Usedoctorto check your own backend.Native search and fetch endpoints are not part of the OpenAI API. They follow the request and response shape used by 9router, and the paths can be changed.
Development
git clone https://github.com/HTB95/open-agent-connector.git
cd open-agent-connector
npm test # node:test, no install needed
npm run doctor -- --text # live probe using your OAC_* envProject layout:
bin/server.mjs CLI entry: MCP server (default), `init`, `doctor`
src/mcp-server.mjs hand-rolled MCP JSON-RPC over stdio
src/router-client.mjs OpenAI-compatible HTTP client (JSON + SSE tolerant)
src/models.mjs role resolution / text-model auto-pick
src/tools/*.mjs ask_agent, web_search, web_fetch, generate_images, list_helper_models
src/init.mjs project wiring for Claude Code
examples/ delegation policy + sample .mcp.json
test/ node:test suites + mock gatewayContributions are welcome. Read CONTRIBUTING.md first.
License
Available Tools
5 toolsask_agentA
Delegate a self-contained chore to a cheaper helper model to save Claude tokens: summarising docs/logs, drafting boilerplate, translating, brainstorming, explaining an API. Give full context in prompt (helpers cannot see the repo). You remain responsible for verifying the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | "auto" = first healthy model in OAC_TEXT_MODELS; "all" = every configured text model in parallel; or an exact model id (see list_helper_models). | auto |
| prompt | Yes | Complete, self-contained task for the helper. | |
| system | No | Optional extra instructions (format, length, persona). | |
| max_tokens | No | Optional output cap for the helper. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose two non-obvious traits: helpers cannot see the repo (context isolation) and the caller remains responsible for verifying output (quality caveat). It does not cover failure modes, latency, or the parallel-call implications of model='all', which the schema only partly addresses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose and cost rationale, then the context requirement and the verification caveat. Every sentence carries actionable information; the example list is compact and aids routing rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with a fully documented schema and no output schema, the description covers purpose, context requirements, and the verification obligation well. It is slightly thin on what the caller receives back (a text answer) and on behavior when model='all' fans out to multiple helpers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that `prompt` must be complete and self-contained and that `system` carries optional extra instructions, but it adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delegate) and resource (a self-contained chore to a cheaper helper model), plus the concrete payoff (save Claude tokens). The examples of delegated work and the contrast with sibling tools like web_search/web_fetch/generate_images make the boundary unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context through enumerated task types (summarising docs/logs, drafting boilerplate, translating, brainstorming, explaining an API) and warns to supply full context. It stops short of naming an explicit alternative or a when-not-to-use condition, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imagesA
Generate images with every configured image model in parallel (OAC_IMAGE_MODELS), save candidates to disk, and return them for you (Claude) to choose the best. review: "paths" (default; open chosen files with Read), "judge" (a helper vision model pre-ranks — cheapest), "inline" (images embedded in this result). Copy the winner into the project yourself afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | auto | |
| models | No | Optional subset of image model ids (default: all of OAC_IMAGE_MODELS). | |
| prompt | Yes | Detailed image brief (subject, style, composition, colours, text). | |
| review | No | paths | |
| out_dir | No | Directory for candidates (default OAC_IMAGE_OUT_DIR). | |
| quality | No | OpenAI-style models only (gpt-image, dall-e). | auto |
| background | No | OpenAI-style models only. | auto |
| n_per_model | No | Images per model. | |
| output_format | No | OpenAI-style models only. | |
| judge_criteria | No | Optional criteria for review=judge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses parallel multi-model execution, on-disk side effects (files written to out_dir), that results are returned for human/model selection, and relative cost of the judge mode. It omits failure/partial-failure behavior and permission or rate-limit considerations, keeping it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and workflow are front-loaded in the first sentence, and the inline enumeration of review modes is dense but earns its space. The quoted mode syntax is slightly awkward to parse but no sentence is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no annotations and no output schema, the description covers the essential lifecycle: generation, persistence, return-for-review, and the manual copy step. Remaining gaps (what the return payload looks like per mode, error handling when a model fails) are modest but real.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the baseline is 3, but the description adds real value by explaining the three 'review' enum values (paths/judge/inline) and their implications, which the schema does not document. It also clarifies the 'models' default (all of OAC_IMAGE_MODELS) and the out_dir default. The size/quality/background knobs are left to the schema, which is acceptable given the high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Generate images') plus the distinctive mechanism: it fans out across every configured model in parallel, saves candidates to disk, and returns them for selection. This is far more informative than a tautological restatement and lets an agent understand the orchestration behavior immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit mode-selection guidance ('paths' default, 'judge' is cheapest, 'inline' embeds images) and a clear post-step ('Copy the winner into the project yourself afterwards'). It stops short of stating when to prefer this tool over the unrelated siblings (web_search/web_fetch), but for a specialized generator that omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_helper_modelsB
Show which helper model ids are configured for each role and which ids the backend advertises.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. "Show" implies a non-destructive read, which is appropriate, and it distinguishes configured ids from backend-advertised ones. However it never confirms read-only behavior or states whether any auth or environment context is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, and the key distinction (configured vs advertised) lands up front. Slightly dense with jargon ("helper model ids", "role") but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter introspection tool with no output schema, the description adequately conveys what is reported: per-role configured ids and backend-advertised ids. Given the trivial input surface, the only real omission is when to use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. There is nothing for the description to clarify about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ("show") and resource ("helper model ids"), and adds scope detail: ids configured per role plus ids the backend advertises. That distinction is informative beyond a restatement of the name. It does not need sibling differentiation since ask_agent, web_search, web_fetch, and generate_images are unrelated operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool, what scenario prompts it (e.g., debugging model routing or verifying configuration), or any prerequisites. The description only says what is shown, never when an agent should reach for it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_fetchA
Read a web page. With question, a helper model reads the full page and returns only the answer — far cheaper than loading the page into Claude. Without it, returns cleaned page text (capped).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | http(s) URL to read. | |
| question | No | What to extract/answer from the page (recommended). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: a helper model processes the page and returns only the answer when `question` is set, and otherwise cleaned (and capped) page text is returned. It omits auth requirements, rate limits, and failure modes such as dynamic or blocked pages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the two modes, then the cost rationale. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does the work of explaining what each mode returns, which is the key unknown for a fetch tool. It stops short of covering error/dynamic-content behavior and what the text cap actually is, but the essentials for calling it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining the mechanical consequence of supplying `question`: a downstream model reads the full page and returns just the answer, versus raw cleaned text without it. This adds real meaning beyond the schema's brief descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a web page'), and the two operating modes are spelled out. It does not explicitly name or contrast with the sibling web_search, so an agent still has to infer the fetch-vs-search distinction, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear mode-selection guidance for the `question` parameter ('far cheaper than loading the page into Claude'), which is effectively a cost-based routing hint. However, it never states when to choose this tool over siblings like web_search or ask_agent, so tool-level usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchA
Live web search done by helpers. Returns a short grounded answer + source URLs instead of raw pages, saving Claude tokens. Use for docs, error messages, versions, news. If the result starts with "NOT A LIVE SEARCH", treat it as possibly outdated.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | Optional: what the answer should concentrate on. | |
| query | Yes | What to find out. Be specific (versions, dates, error text). | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does meaningfully more than the name: it discloses the return contract (grounded answer + source URLs, not pages), a token-saving rationale, and a real edge case ('NOT A LIVE SEARCH' prefix means possibly stale). It omits rate limits, freshness guarantees, and auth needs, which keeps it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the purpose and output contract, then usage triggers, then the staleness caveat. Nothing is redundant, though 'done by helpers' is slightly vague and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description is the sole source of behavioral context, and it covers purpose, return shape, usage triggers, and a notable result-prefix caveat. It leaves out pagination/limits behavior and freshness details, but is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with query and focus already documented in the schema and the query guidance ('be specific — versions, dates, error text') largely duplicated in the description. The description says nothing about 'focus' or the 'max_results' bound (1-10, default 5), so it adds little beyond the structured fields; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Live web search') and immediately defines the output shape ('short grounded answer + source URLs instead of raw pages'), which implicitly separates it from the sibling web_fetch that returns raw pages. An agent can pick between web_search and web_fetch without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use triggers ('docs, error messages, versions, news') and contrasts with the raw-page alternative. It stops short of explicit exclusions (e.g., when to prefer web_fetch or ask_agent instead), but the context is clear enough to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.4.0- First observed
ask_agent - First observed
generate_images - First observed
list_helper_models - First observed
web_fetch - First observed
web_search
TDQS
Scored across 5 tools
Each tool has a distinct primary role: ask_agent delegates text chores, web_search returns grounded search answers, web_fetch extracts page content, generate_images handles image generation, and list_helper_models introspects configuration. There is minor overlap between ask_agent and web_search for documentation/error queries, but the descriptions clarify when to use each.
All tool names use consistent snake_case, which makes them easy to scan. The pattern is mostly verb_noun (ask_agent, generate_images, list_helper_models) but web_search and web_fetch use noun_verb, a minor deviation from full consistency.
Five tools are well-scoped for a connector that delegates to helper models and performs web/image operations. Each tool covers a distinct capability without redundancy or bloat.
The surface covers delegation, live web search, page fetching, image generation, and helper model introspection. Minor gaps exist, such as no explicit local file-reading tool or helper configuration tool, but the core workflows are supported.
Maintenance
Related MCP Connectors
Agent personas for Claude. 16 tools, 13 personas, 3 workflows. Zero extra API cost. Free.
Build agents to automate any background task. Works with your ChatGPT/Claude subscription.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Remote streamable-HTTP MCP server running on a single Cloudflare Worker. Your assistant gets live Airbnb, Amazon, Booking.com, Google Flights, Maps and Reddit data, social search on X, Instagram and TikTok, the Meta Ad Library, and image/video generation without any keys. Connect your own accounts to let it send WhatsApp or Telegram messages, work an IMAP inbox, manage Meta Ads campaigns and publish to X and LinkedIn. OAuth 2.1 with PKCE; stored credentials are AES-256-GCM encrypted.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables Claude to delegate tasks to OpenAI expert models (GPT-5.3-Codex and GPT-5.5) as subagents, with orchestration patterns for safe and effective use.3MIT
- AlicenseAqualityDmaintenanceEnables Claude to orchestrate tasks across 27 AI providers, run multi-agent plans, and conduct multi-model councils for decision-making.15108 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables Claude to offload mechanical, output-heavy tasks like boilerplate, type generation, and summarization to OpenAI while keeping reasoning in-context.-
- AlicenseNot gradedqualityAmaintenanceLets Claude Code, Claude Desktop, or Codex delegate busywork to many cheap AI models at once and receive structured answers as data.1MIT