Skip to main content
Glama
r28ai

Web Research to Docs

by r28ai

Web Research to Docs

Competitor watch, cited research reports, paper alerts and market maps, filed where you read.

An MCP server with 10 workflows across Firecrawl, Tavily, Notion, Slack, Google Docs, Google Sheets, GitHub and Linear. Each workflow is a prompt your agent runs as a slash command, over the 19 tools it needs and no others.

uv tool install https://github.com/r28ai/web-research-to-docs-mcp/releases/download/v0.1.0/web_research_to_docs_mcp-0.1.0-py3-none-any.whl
claude mcp add research -- web-research-to-docs-mcp

It installs with uv from this repository's release, with no git and nothing to build; nothing but Charter and the libraries it uses comes from PyPI. To update, run the install line from the latest release. If a desktop app cannot find web-research-to-docs-mcp, give it the full path from which web-research-to-docs-mcp (where web-research-to-docs-mcp on Windows).

Then ask your agent to connect your apps, or run /mcp__research__setup.

Connect your apps

Ask the agent to connect one ("connect Linear"). It tells you where to get that app's key and the command that stores it, and the next call works, with no restart. The agent never asks for a key in the chat.

Or connect everything this server uses from a terminal:

web-research-to-docs-mcp login            # each app in turn
web-research-to-docs-mcp login firecrawl  # just one
web-research-to-docs-mcp status           # what is connected

Tokens and keys go to your operating system's keychain (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), and are checked with one read-only call to the app's own API before they are kept. Every key, token and OAuth client is yours: we register no app with any of these services, and nothing passes through a server of ours, because there isn't one.

App

How it connects

Or set

Firecrawl

Your own key (get one), entered once.

FIRECRAWL_API_KEY

Tavily

Your own key (get one), entered once.

TAVILY_API_KEY

Notion

Your own key (get one), entered once. Then share the pages it should see with the integration.

NOTION_API_KEY

Slack

Your own key (get one), entered once. A bot token from your own Slack app, which the guide sets up in about three minutes.

SLACK_BOT_TOKEN

Google

Browser sign-in, over your own OAuth client (make one).

GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET

GitHub

Your own key (get one), entered once.

GITHUB_TOKEN

Linear

Your own key (get one), entered once.

LINEAR_API_KEY

A variable set in your client's config always wins over the keychain.

Related MCP server: MCP OSINT Server

Workflows

Workflow

What you get

Apps

Competitive intel digest competitive_intel_digest

Site changes and news per competitor, weekly, in one page.

Firecrawl, Tavily, Notion, Slack

Deep research → shared doc deep_research_to_shared_doc

A cited report in the team's Drive, not in someone's chat history.

Tavily, Google Docs, Slack

Paper watch paper_watch

New papers on your topics land in a reading list with abstracts.

Firecrawl, Notion, Slack

Market map market_map

Players, pricing, funding and positioning, one row each.

Tavily, Firecrawl, Google Sheets

Pricing benchmark memo pricing_benchmark_memo

Ten competitors' pricing pages normalised into one table and a recommendation.

Firecrawl, Google Sheets, Google Docs

Regulatory watch regulatory_watch

A regulator's guidance page changes and legal hears the same day.

Firecrawl, Notion, Slack

Company due diligence company_due_diligence

Public footprint, open-source activity and product surface in one memo.

Tavily, GitHub, Firecrawl, Google Docs

Public complaints → roadmap evidence public_complaints_to_roadmap_evidence

What people complain about in your category, attached to the issues it supports.

Tavily, Firecrawl, Notion, Linear

Open-source landscape open_source_landscape

Who is building what in your space, with momentum, before you build it.

GitHub, Tavily, Notion

Web page → Linear issue web_page_to_linear_issue

A public bug report, forum post or status page becomes a tracked issue.

Firecrawl, Linear

Every prompt takes one optional argument, details: the repo, team, channel, customer or date range you mean, so the agent does not have to ask. In Claude Code, put it in quotes, or only its first word arrives:

/mcp__research__competitive_intel_digest "competitors acme.com and globex.com"

Reads run without asking. Before anything that creates, sends, changes or deletes, the prompt tells the agent to show you the call and wait.

6 of the 10 workflows need no Google or Granola credential.

Other clients

Claude Desktop: install uv if you have not, since Claude Desktop starts the server with it, then open the .mcpb from the latest release. Claude asks for any keys in its own settings and keeps them in your keychain. The first start takes a few seconds longer, while uv installs it.

VS Code (.vscode/mcp.json): VS Code asks for each key the first time the server starts and stores it securely. Leave out any you stored with login.

{
  "inputs": [
    {
      "type": "promptString",
      "id": "firecrawl-api-key",
      "description": "Firecrawl: API key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "tavily-api-key",
      "description": "Tavily: API key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "notion-api-key",
      "description": "Notion: Integration secret (ntn_\u2026)",
      "password": true
    },
    {
      "type": "promptString",
      "id": "slack-bot-token",
      "description": "Slack: Bot token (xoxb-\u2026)",
      "password": true
    },
    {
      "type": "promptString",
      "id": "google-client-secret",
      "description": "Google: OAuth client secret",
      "password": true
    },
    {
      "type": "promptString",
      "id": "github-token",
      "description": "GitHub: Personal access token",
      "password": true
    },
    {
      "type": "promptString",
      "id": "linear-api-key",
      "description": "Linear: Personal API key",
      "password": true
    }
  ],
  "servers": {
    "research": {
      "type": "stdio",
      "command": "web-research-to-docs-mcp",
      "env": {
        "FIRECRAWL_API_KEY": "${input:firecrawl-api-key}",
        "TAVILY_API_KEY": "${input:tavily-api-key}",
        "NOTION_API_KEY": "${input:notion-api-key}",
        "SLACK_BOT_TOKEN": "${input:slack-bot-token}",
        "GOOGLE_CLIENT_SECRET": "${input:google-client-secret}",
        "GITHUB_TOKEN": "${input:github-token}",
        "LINEAR_API_KEY": "${input:linear-api-key}",
        "GOOGLE_CLIENT_ID": ""
      }
    }
  }
}

Cursor (.cursor/mcp.json) starts it the same way:

{
  "mcpServers": {
    "research": {
      "command": "web-research-to-docs-mcp"
    }
  }
}

Codex (~/.codex/config.toml) starts a turn without waiting for a server unless it is required, and then the agent has none of its tools. required = true makes the session wait for it, and startup_readiness = "catalog" waits for its tool list rather than just its connection:

[mcp_servers.research]
command = "web-research-to-docs-mcp"
required = true
startup_readiness = "catalog"
startup_timeout_sec = 30

Name the server research. A host builds each tool's name from that key, and a longer one can push a tool past the 64 characters a function name allows.

Built with Charter

Every tool here is a Charter declaration: a Pydantic schema saying where each field goes on the wire. Charter's runtime builds the request, attaches and refreshes the credential, and trims the response before the model reads it. It runs in your process, with no proxy and no telemetry.

The 19 tool schemas come to 38,591 tokens.

The same tools work in your own agent, without MCP:

from charter.adapters.openai import to_openai_tools
from charter_packs_mcp import FAMILIES

tools = FAMILIES["research"].tools()
definitions = to_openai_tools(tools)   # or charter.adapters.langchain

Need an API that isn't here? Write a pack: your coding agent writes the declarations, and Charter's conformance suite checks them.

  • Firecrawl: firecrawl_monitor_checks_list, firecrawl_research_papers_search, firecrawl_research_paper_get, firecrawl_extract, firecrawl_monitor_create, firecrawl_crawl, firecrawl_scrape

  • Tavily: tavily_search, tavily_research_create, tavily_research_get

  • Notion: notion_pages_create

  • Slack: slack_chat_post_message

  • Google Docs: gdocs_documents_create

  • Google Sheets: gsheets_spreadsheets_values_update

  • GitHub: github_search_repositories, github_repos_list_languages

  • Linear: linear_customer_need_create, linear_search_issues, linear_issue_create

License

Apache 2.0.

Available Tools

21 tools
connectA

Connect one app this server uses. For an app that issues keys, says where to get one and the terminal command that stores it. For Google, once the user's own OAuth client is set, starts the browser sign-in and returns at once: the user approves in the browser and the next call works. To see which apps are connected, call connection_status. Never ask the user for a key in the chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesThe app to connect.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=false, so the description carries real weight: it discloses that key-based apps surface where to get a key plus a terminal command, while Google starts an OAuth browser flow that 'returns at once' and works on the next call. That interactive, deferred-completion behavior is meaningful context absent from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then behavior per app type, then the sibling pointer and constraint. Every sentence carries information, though the Google flow sentence is a touch dense. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-param mutation tool with no output schema, the description covers the key behavioral cases (key entry vs OAuth), the immediate-return semantics of the Google path, and how to verify state via `connection_status`. It omits failure handling, but is otherwise complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum lists all apps, so the baseline is 3. The description adds value beyond the raw enum by grouping apps into 'key-issuing' vs OAuth (Google) categories and explaining what each implies, giving the `app` parameter semantic meaning the schema alone does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Connect one app this server uses'), so the action is unambiguous. It also names the relevant sibling, `connection_status`, so an agent can tell the connect action apart from the status-check action. It stops short of naming the other siblings, but those are app-specific tools unrelated to this operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'To see which apps are connected, call `connection_status`,' and adds a hard constraint, 'Never ask the user for a key in the chat.' It gives clear context for when/how to call, though it doesn't spell out refusal or retry conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connection_statusA
Read-only

See which apps this server is connected to, and how to connect each one that is not. Changes nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered by structured data. 'Changes nothing' restates the readOnly hint rather than adding new behavior; the only incremental value is noting that connect instructions are returned for unconnected apps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the primary purpose and appends the secondary benefit with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only status tool with no output schema, the description covers both what is inspected and the shape of the useful payload (connect guidance). Return format details are absent but minimal given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly implies no input is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('See') and resource ('which apps this server is connected to'), and adds the secondary payload of connect instructions for missing apps. This distinguishes it from the sibling 'connect' tool, which performs the connection rather than reporting status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'how to connect each one that is not' implies this tool is the discovery step before using 'connect', but the sibling is never named and there is no explicit when-to-use/when-not statement. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_crawlC

Recursively crawl a website and scrape each discovered page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe base URL to start crawling from
delayNoDelay in seconds between scrapes. Setting this forces concurrency to 1.
limitNoMaximum number of pages to crawl. The server applies 10000 when this is absent.
promptNoNatural language prompt to generate crawler options from.
sitemapNoSitemap mode when crawling. The server applies 'include' when this is absent.
webhookNoWebhook specification for crawl lifecycle events.
excludePathsNoURL pathname regex patterns that exclude matching URLs from the crawl.
includePathsNoURL pathname regex patterns that include matching URLs in the crawl.
scrapeOptionsNoOptions applied when scraping each crawled page.
maxConcurrencyNoMaximum number of concurrent scrapes for this crawl.
regexOnFullURLNoMatch includePaths and excludePaths against the full URL instead of just the pathname.
allowSubdomainsNoAllow the crawler to follow links to subdomains of the main domain.
ignoreRobotsTxtNoIgnore the website's robots.txt rules. Enterprise only.
robotsUserAgentNoCustom User-Agent string for robots.txt evaluation. Enterprise only.
crawlEntireDomainNoAllow the crawler to follow internal links to sibling or parent URLs, not just child paths.
maxDiscoveryDepthNoMaximum depth to crawl based on discovery order.
zeroDataRetentionNoIf true, this will enable zero data retention for this crawl.
allowExternalLinksNoAllow the crawler to follow links to external websites (one hop only).
ignoreQueryParametersNoDo not re-scrape the same path with different query parameters.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the description need only add context, and it adds none. It omits that this kicks off an asynchronous crawl job (a firecrawl_crawl_status sibling exists for polling), that pages/credits are consumed, and that a limit defaults to 10000.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. It is efficient, though its brevity is partly under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter, open-world, non-read-only crawl tool with no output schema, the description is far too thin. It says nothing about asynchronous job behavior, polling via crawl_status, scope controls (allowSubdomains, includePaths), or cost, all of which an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the 19 parameters are fully documented in the schema. The description contributes no additional parameter meaning, which is the baseline-3 case when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'recursively crawl a website and scrape each discovered page.' The word 'recursively' and 'each discovered page' implicitly distinguish it from the single-page firecrawl_scrape sibling, but no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use guidance, no conditions selecting it over firecrawl_scrape, firecrawl_map, or firecrawl_extract, and no prerequisites. The agent must infer usage purely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_extractA

Extract structured data from one or more URLs using an LLM. Poll results with extract_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesThe URLs to extract data from. URLs should be in glob format.
promptNoPrompt to guide the extraction process.
schemaNoSchema to define the structure of the extracted data. Must conform to JSON Schema.
showSourcesNoWhen true, the sources used to extract the data will be included in the response as `sources`.
ignoreSitemapNoWhen true, sitemap.xml files will be ignored during website scanning.
scrapeOptionsNoOptions applied when scraping pages for extraction.
enableWebSearchNoWhen true, the extraction will use web search to find additional data.
threatProtectionNoPer-request threat protection override. Enterprise feature.
ignoreInvalidURLsNoIf invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent.
includeSubdomainsNoWhen true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true) are minimal and do not reveal the async job lifecycle, so the description correctly surfaces the most important behavioral trait: results must be polled via extract_status. It stops there, omitting what the initial call returns (a job id?), failure behavior, and whether web-search/credit costing applies. Decent added context over annotations, but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste. The core purpose is front-loaded and the polling instruction follows immediately. Nothing to trim and nothing misplaced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no output schema, and an async job model, the description covers the poll step but not the full lifecycle: it does not explain what the initial invocation returns or how to correlate extract_status results back to it. For a complex async tool with no output schema to fall back on, this is a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 10 parameters (urls, prompt, schema, scrapeOptions, enableWebSearch, etc.) are already documented in the schema. The description adds no parameter-level detail (e.g., the glob URL format or JSON Schema requirement) beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Extract structured data from one or more URLs') and adds the key mechanism ('using an LLM'), so the agent understands this is LLM-driven structured extraction rather than raw scraping. It does not, however, distinguish itself from the very close sibling firecrawl_scrape, leaving the boundary between the two to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence ('Poll results with extract_status') is genuine usage guidance: it tells the agent this is an async operation that must be polled. But it gives no guidance on when to choose this tool over firecrawl_scrape or other extraction options, and does not name any exclusion criteria. Usage is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_monitor_checks_listC
Read-onlyIdempotent

List checks for a monitor.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum checks to return. The server applies 25 when this is absent.
offsetNoNumber of checks to skip. The server applies 0 when this is absent.
statusNoFilter checks by status.
monitorIdYesThe monitor ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered externally. The description adds nothing beyond them: no note on ordering, pagination behavior, or what a check record contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero padding. It is efficient, though its brevity borders on under-specification for a tool with paging and status filters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool whose schema fully documents its four parameters and whose annotations cover safety, this is roughly adequate. It omits return-shape or ordering context, but with no output schema and a simple contract, that gap is modest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself documents limit (default 25), offset (default 0), status enum values, and monitorId. The one-line description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (checks) scoped to a monitor, which is enough to distinguish it from firecrawl_monitor_create. It does not, however, explicitly contrast itself with any sibling tool, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites, and no alternatives. An agent must infer from the name alone that this is the read path for monitor check history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_monitor_createB

Create a scheduled monitor for scrape, crawl, or search targets.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoPlain-language goal used to judge whether changed pages are meaningful.
nameYesMonitor name.
targetsYesTargets to run on each check.
webhookNoWebhook destination for monitor events.
scheduleYesSchedule for monitor checks.
judgeEnabledNoWhether to judge changed pages against goal.
notificationNoNotification destinations.
retentionDaysNoHow long to retain monitor history. The server applies 30 when this is absent.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the mutation and external-reach profile is already covered; the description's 'create' is consistent with these and adds no contradiction. It adds only the notion that the created object is scheduled/recurring, but says nothing about lifecycle, cost, or persistence beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, correctly leading with the verb and the resource. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex creation tool with required schedule/targets, nested target union types, webhooks, notifications, retention, and an optional judge/goal mechanism, and no output schema to fall back on. The one-line description does not explain what a monitor does over time or what happens after creation, leaving significant gaps for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters (schedule, targets, webhook, judgeEnabled, notification, retentionDays, goal, name) are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a scheduled monitor') and scopes it to the three target kinds the schema supports (scrape, crawl, search). It implicitly separates this from the one-off firecrawl_scrape/firecrawl_crawl siblings via 'scheduled,' but never names them, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g. that monitors run repeatedly and incur ongoing credit usage), and no pointer to alternatives like firecrawl_crawl or firecrawl_crawl_status for one-off jobs. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_research_paper_getA
Read-onlyIdempotent

Inspect metadata or read passages from a research paper.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoPassage count for read mode. Only valid when query is present. The server applies 4 when this is absent.
idYesPaper reference: a canonical paperId or source-specific primaryId.
queryNoWhen present, returns top matching full-text passages for this question.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds the useful fact that there are two distinct behaviors (metadata vs passage reading), but does not disclose mode-selection rules, limits, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, front-loading the two capabilities. Nothing is wasted and the essential action is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries some burden for explaining returns, yet it only broadly says metadata or passages. The schema compensates for parameter behavior, but the two-mode contract and what each mode returns are only minimally conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents that query returns matching passages, k sets passage count (defaulting to 4), and id is a paperId or primaryId. The description adds no parameter meaning beyond this, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (inspect/read) and resource (research paper) and outlines two modes: metadata inspection and passage reading. It is clear what the tool does, though it does not explicitly differentiate itself from the sibling firecrawl_research_papers_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two operating modes are implied, and the schema hints that passing a query triggers read mode, but the description never states when to use this tool versus firecrawl_research_papers_search or when to prefer metadata vs passage reading. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_scrapeB

Scrape a single URL and optionally extract information. Use when the user wants to read or summarize a specific webpage. Supports markdown, HTML, screenshots, and structured JSON extraction.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape
proxyNoSpecifies the type of proxy to use.
maxAgeNoReturns a cached version of the page if it is younger than this age in milliseconds. The server applies 172800000 (2 days) when this is absent.
minAgeNoWhen set, the request only checks the cache and never triggers a fresh scrape.
mobileNoEmulate scraping from a mobile device.
actionsNoActions to perform on the page before grabbing the content.
formatsNoOutput formats to include in the response. Strings or objects. The server applies markdown when this is absent.
headersNoHeaders to send with the request.
parsersNoControls how files are processed during scraping.
profileNoPersistent browser storage across scrape and interact sessions.
timeoutNoTimeout in milliseconds. The server applies 60000 when this is absent.
waitForNoSpecify a delay in milliseconds before fetching the content. The server applies 0 when this is absent.
blockAdsNoEnables ad-blocking and cookie popup blocking.
locationNoLocation settings for the request.
lockdownNoServe from cache only and never make an outbound request. On miss, returns 404 SCRAPE_LOCKDOWN_CACHE_MISS.
redactPIINoRedact personally identifiable information from returned markdown. Pass true for defaults, or an object to tune it.
excludeTagsNoTags to exclude from the output.
includeTagsNoTags to include in the output.
storeInCacheNoIf true, the page will be stored in the Firecrawl index and cache.
auditMetadataNoUser attribution included with SIEM logging events when SIEM is enabled.
onlyMainContentNoOnly return the main content of the page excluding headers, navs, footers, etc. The server applies true when this is absent.
onlyCleanContentNoBeta. LLM pass over markdown to remove residual boilerplate that onlyMainContent can miss.
threatProtectionNoPer-request threat protection override. Enterprise feature.
zeroDataRetentionNoIf true, this will enable zero data retention for this scrape. To enable this feature, please contact help@firecrawl.dev
removeBase64ImagesNoRemoves all base64 images from the markdown output.
skipTlsVerificationNoSkip TLS certificate verification when making requests.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the agent knows this touches external state. The description adds the supported output formats, which is useful, but it omits behavior implied by the schema — that actions (click/write/executeJavascript) mutate the page, that storeInCache writes to an external index, and that some features cost credits. Nothing contradicts the annotations, but the added behavioral detail is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core purpose front-loaded and no filler. The trailing format list is somewhat redundant with the schema's formats enum, which keeps it short of ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a 26-parameter tool with no output schema and only minimal annotations, yet the description is three sentences long. It omits cost/credit implications, caching/lockdown semantics, the relationship to firecrawl_extract, and any hint of what the response looks like, so an agent invoking it correctly still depends almost entirely on reading the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 26 parameters are already documented in-schema and the baseline is 3. The description echoes the formats dimension ('markdown, HTML, screenshots, structured JSON extraction') but adds no format syntax, precedence, or interaction detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Scrape a single URL') plus an optional outgrowth ('optionally extract information'), so an agent can tell it is a per-URL content fetcher. The phrase 'a single URL' gestures at the multi-URL alternative but never names firecrawl_extract, so sibling differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one clear trigger ('Use when the user wants to read or summarize a specific webpage'), which is real usage guidance. But it never states when NOT to use it, nor does it point to firecrawl_extract for bulk/structured extraction, so the routing decision against the closest sibling is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gdocs_documents_createA

Create a blank document with a title. Only the title is honoured — the document is created empty. To add content, call this and then documents_batch_update with the returned documentId.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoThe document to create. Only the title is honoured.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the behavioral load and does well: it discloses the critical constraint that only the title is honoured and the document is created empty. It also notes the documentId is returned, which matters with no output schema. It stops short of noting auth/permission or quota behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, zero padding, and the most important constraint (empty document, title-only) is front-loaded before the follow-up instructions. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by mentioning the returned documentId and the batch_update follow-up, which is what an agent needs to chain calls. For a single-param mutation tool with annotations covering the safety profile, this is nearly complete; only auth/error behavior is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the title field already states it is the only honoured field, so the description largely restates structured data. Baseline 3 applies; the emphasis on the empty-document effect is useful but not new information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (document), and immediately disambiguates the scope: it produces a blank/empty document, not a content-bearing one. This is a distinct action an agent can tell apart from content-writing operations. The mention of documents_batch_update further fixes its place in the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear sequencing guidance: to add content, create first and then call documents_batch_update with the returned documentId. It does not state explicit exclusions or alternatives (e.g., when to use a copy/template flow instead), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

github_repos_list_languagesA
Read-onlyIdempotent

List a repository's languages with bytes of code each — the quickest way to find out what a repository is written in before reading any of it. Answers with an object keyed by language, so there is nothing to page.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYesThe name of the repository without the `.git` extension. The name is not case sensitive.
ownerYesThe account owner of the repository. The name is not case sensitive.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint, idempotentHint and openWorldHint already covering the safety profile, the description still earns credit by disclosing the return shape (an object keyed by language) and the absence of pagination, which matters since no output schema exists. It does not mention auth requirements or rate limits, so it is not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tightly written sentence that front-loads the action and resource, then appends the two facts an agent needs (bytes per language, no paging). Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only tool with a fully documented schema and no output schema, the description supplies everything missing: what it returns, its keying, and that there is no pagination to handle. Nothing an agent needs to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both required parameters (owner, repo) are documented in the schema, so the baseline is 3. The description adds no formatting or case-sensitivity guidance beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (a repository's languages) and adds the payload detail (bytes of code each). An agent immediately knows this is the language-breakdown read, distinct from the GitHub search tools among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for use — 'the quickest way to find out what a repository is written in before reading any of it' — which tells the agent when this tool is the right first step. It does not, however, name an alternative tool or state exclusions, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

github_search_repositoriesA
Read-onlyIdempotent

Search repositories with qualifiers, e.g. 'topic:cli language:go stars:>500'. Rate limited to 30 requests per minute.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesThe query containing one or more search keywords and qualifiers, e.g. `tetris language:assembly stars:>100`. Qualifiers include `language:`, `stars:`, `forks:`, `topic:`, `org:`, `user:`, `license:`.
pageNoThe page number of the results to fetch. Defaults to 1.
sortNoSorts the results by number of stars, forks, help-wanted issues, or how recently the items were updated. Default: best match.
orderNoDetermines whether the first search result returned is the highest number of matches (`desc`) or lowest (`asc`). Ignored unless `sort` is provided. GitHub uses `desc` when this is absent.
perPageNoThe number of results per page (max 100). Defaults to 30.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, open-world, and idempotent behavior. The description adds a valuable operational constraint: a rate limit of 30 requests per minute, which is not available from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no waste. The purpose and example are front-loaded, followed by the rate limit. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema and annotations are rich, and the description adds a key rate-limit detail. It is nearly complete for a search tool, though it could note that results are paginated or what the response contains if an agent needed that reassurance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters and qualifier syntax. The description's example query is consistent with the schema but does not add meaning beyond it, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: search repositories. The example query with qualifiers makes the scope unambiguous and distinguishes it from sibling tools like github_search_commits or github_repos_list_languages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool name and description, but there is no explicit when-to-use or when-not-to-use guidance, nor any comparison to alternatives such as github_repos_list_languages or tavily_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_updateC

Sets values in a range of a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation of the values to update.
valueRangeYesThe request body contains an instance of ValueRange.
spreadsheetIdYesThe ID of the spreadsheet to update.
valueInputOptionYesHow the input data should be interpreted.
includeValuesInResponseNoDetermines if the update response should include the values of the cells that were updated. By default, responses do not include the updated values. If the range to write was larger than the range actually written, the response includes all values in the requested range (excluding trailing empty rows and columns).
responseValueRenderOptionNoDetermines how values in the response should be rendered. The default render option is FORMATTED_VALUE.
responseDateTimeRenderOptionNoDetermines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the write nature is known. But the description adds nothing beyond the name – it does not disclose that existing cell values are overwritten, how valueInputOption affects interpretation, or what the update returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is tight. But for a 7-parameter mutation tool it is under-specified rather than appropriately sized; conciseness here is closer to omission.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with 7 params and no output schema, the description omits overwrite semantics, auth requirements, and response behavior. An agent could call it, but not safely without reading the schema closely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with 7 well-documented parameters, so the schema carries the meaning. The description adds no parameter detail beyond 'in a range', which is baseline 3 for fully documented schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (sets values in a range of a spreadsheet), which an agent can distinguish from gsheets_spreadsheets_values_get by direction of data flow. However it offers no explicit sibling differentiation and largely restates the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the sibling gsheets_spreadsheets_values_get or when reading vs writing applies. The agent must infer context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

linear_customer_need_createA

Record a customer request, optionally attached to an issue or project. This is the one Linear mutation whose reply carries no object — it answers only with whether it worked.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe request to record.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true. The description adds a genuinely useful behavioral trait that no structured field conveys: the response carries no object and only indicates success/failure, which is important given there is no output schema. It stops short of mentioning permission or side-effect details, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no waste. The core action is front-loaded, and the second sentence efficiently conveys the unusual no-object return without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create mutation with annotations covering the safety profile and no output schema, the description covers the key agent-facing concerns: what it does and what it returns. The main gap is the absence of usage/permission context, which keeps it below a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and every parameter (including the nested input object fields) is fully documented in the schema. The description only lightly gestures at the attachment options ('an issue or project'), adding little beyond what the schema already provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Record') and resource ('customer request'), which cleanly distinguishes it from linear_issue_create and linear_search_issues in the sibling list. It could go further by explicitly naming the alternative tool, but the operation and target object are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named. The phrase 'optionally attached to an issue or project' implies a relationship to those areas but does not help an agent decide between this tool and linear_issue_create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

linear_issue_createA

Create an issue. team_id and title are required; everything else is optional. The UUIDs for team, assignee, state and labels come from teams_list, users_list and workflow_states_list — Linear does not accept names here. Set parent_id to create a sub-issue.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe issue to create.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the mutation/external-scope profile is covered. The description adds genuinely useful behavior: Linear rejects names and only accepts UUIDs resolved via teams_list/users_list/workflow_states_list, and parent_id turns this into a sub-issue creation. It stops short of noting side effects like notifications or returned identity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and the required fields, followed by the ID-resolution constraint and the sub-issue tip. No filler, though the UUID guidance partially duplicates the schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool whose one parameter is a deeply nested object fully documented by the schema, plus annotations covering the safety profile, the description covers the essentials an agent needs. It lacks any mention of what creation returns, which is a minor gap given there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every field including the resolver-tool hints and the sub-issue semantics. The description's parameter notes (team_id/title required, everything else optional) largely restate the schema rather than adding new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an issue'), which is instantly distinguishable from the list-oriented sibling linear_issues_list. It does not explicitly name a sibling it is not, so it falls short of a 5, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives practical invocation guidance (required vs optional fields, which resolver tools supply UUIDs, how to make a sub-issue) but never states when to reach for this tool versus alternatives such as linear_issues_list, nor any exclusion or prerequisite conditions. Usage is implied rather than framed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

linear_search_issuesA
Read-onlyIdempotent

Search issues by text, across titles and descriptions. Set include_comments to search inside comments too. This is full-text search; to filter on fields such as state or assignee, use issues_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesWhat to search for.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so safety is covered. The description adds genuinely useful behavioral context beyond them: the search spans titles and descriptions, and comment text is only included when include_comments is set. No return-format or pagination detail, but the schema's `after`/`first` params cover that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with what the tool does, then the one parameter worth calling out, then the routing rule. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only search wrapper whose input schema documents every field, the description supplies exactly the missing layer: search scope, the include_comments toggle, and when to prefer the sibling. Nothing an agent needs in order to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description still adds real meaning: it defines the search surface (titles and descriptions) and explains the effect of include_comments rather than restating its schema text. It does not, however, reconcile that guidance with the presence of a `filter` object in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (search) plus resource (issues) plus the exact searchable fields (titles and descriptions). It explicitly names the sibling it is not (issues_list) and characterizes itself as full-text, so an agent can separate it from filter-based listing without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the alternative and the selecting condition: use issues_list to filter on fields such as state or assignee. That is clear routing guidance, though it slightly undersells the tool's own `filter` parameter, which the schema shows can narrow results on top of the text match.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notion_pages_createA

Create a page — as a subpage of another page, or as a row of a database by giving its data_source_id as the parent. Content comes as a markdown string Notion parses into blocks, or from a template: one or the other, never both.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesThe page to create.
filterPropertiesNoProperty IDs to return on the page that comes back, instead of all of them. A page that does not have a listed property omits it.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, establishing this as an external write. The description adds the meaningful markdown/template mutual exclusion. It does not disclose auth requirements, rate limits, or the allowAsync async-202 behavior, but with annotations carrying the safety profile a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the verb and the two modes, with the exclusivity constraint phrased crisply. No wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation tool with a fully documented schema, the description covers parent selection and content sourcing adequately. It omits the async task path (allowAsync → 202) and return shape, but with no output schema and 100% schema coverage these are minor gaps, and annotations cover the write semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value: it clarifies that `parent` takes `data_source_id` to make a database row and that template vs. markdown are mutually exclusive — beyond the schema's raw field docs. It doesn't add detail on `properties` or `filterProperties`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a page') and immediately delineates the two creation modes — subpage vs. database row — which is exactly what separates it from siblings like notion_pages_update. An agent can identify the tool's scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance: use a page parent for a subpage, `data_source_id` for a database row, and content comes from `markdown` OR a template, 'never both.' The mutual-exclusion rule is explicit. However, it names no sibling alternatives (e.g., update vs. create routing) and states no prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

slack_chat_post_messageA

Send a message to a Slack channel, private group, or DM. Provide text for a plain message; set thread_ts to reply inside an existing thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoThe main body text of the message. Required unless blocks or attachments are provided. Used as the fallback string for notifications when blocks are provided, so it is worth setting even then.
parseNoChange how messages are treated. Accepts 'none' or 'full'.
blocksNoA JSON-based array of structured Block Kit blocks.
mrkdwnNoDisable Slack markup parsing by setting to false. Defaults to true.
channelYesAn encoded ID or channel name that represents a channel, private group, or IM channel to send the message to. Prefer the encoded ID (e.g. 'C123ABC456').
iconUrlNoURL to an image to use as the icon for this message. Requires the chat:write.customize scope.
metadataNoApplication-specific metadata to attach to the message.
threadTsNoProvide another message's 'ts' value to make this message a reply in that thread. Avoid using a reply's ts value; use the parent's.
usernameNoSet the bot's user name. Requires the chat:write.customize scope.
iconEmojiNoEmoji to use as the icon for this message, e.g. ':chart_with_upwards_trend:'. Requires the chat:write.customize scope.
linkNamesNoFind and link user groups.
attachmentsNoA JSON-based array of structured attachments.
unfurlLinksNoPass true to enable unfurling of primarily text-based content.
unfurlMediaNoPass false to disable unfurling of media content.
markdownTextNoAccepts message text formatted in markdown. Limit this field to 12,000 characters. Cannot be used together with blocks or text.
replyBroadcastNoUsed in conjunction with thread_ts and indicates whether the reply should be made visible to everyone in the channel. Defaults to false.
unfurlAppLinksNoPass true to enable unfurling of links to installed apps.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external nature is covered. The description adds the threading behavior, but does not disclose required scopes, rate limits, message-size limits, or what a successful send returns — and most scope info already lives in the schema parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and then the two most important parameter behaviors. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with no output schema, the description is thin but the schema carries full parameter documentation, so an agent can call it correctly. Missing behavioral context (rate limits, required scopes, response shape) keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 17 parameters are documented in the schema itself. The description's notes on `text` and `thread_ts` largely restate what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Send') and resource ('a message to a Slack channel, private group, or DM'), making the action and destination unambiguous. An agent can distinguish this from siblings like slack_conversations_create without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives implied usage for `text` and `thread_ts`, but there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. chat:write scope), and no reference to alternative messaging tools such as gmail_messages_send. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tavily_research_createA

Create an async research task that searches, analyzes sources, and generates a cited report. Poll results with research_get.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNoAttach up to 5 files as additional sources. Each file may be at most 80,000 words; combined total at most 80,000 words.
inputYesResearch task or question.
modelNoResearch agent model tier. The server applies `auto` when this is absent.
outputLengthNoTarget response size. The server applies `standard` when this is absent.
outputSchemaNoJSON Schema defining structured output shape.
citationFormatNoCitation format in the report. The server applies `numbered` when this is absent.
excludeDomainsNoHard blocklist (max 20). Downward subdomain matching only.
includeDomainsNoSoft source preference (max 20). Host-based subdomain matching.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the safety profile is partly covered. The description adds important behavioral context beyond annotations: the task is asynchronous and results must be polled via research_get. It does not cover potential costs, rate limits, or task persistence, but the async/polling disclosure is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and immediately followed by the essential polling instruction. No filler or repetition; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 8-parameter tool with no output schema, the description covers the key behavioral facts: it creates an async task that produces a cited report, and results are retrieved with research_get. It omits details about return shape or parameter nuances, but those are either in the schema or in the sibling polling tool, making the description largely complete for its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters are fully documented in the schema. The description adds no parameter-specific meaning (e.g., it does not explain input, files, or model tiers). Baseline 3 is appropriate when the schema already carries the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create') and resource ('async research task'), and explains the action chain: searches, analyzes sources, and generates a cited report. It differentiates from the sibling tavily_research_get by pointing to polling, but does not explicitly distinguish itself from tavily_search, leaving a small gap in sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through 'Create an async research task...' and gives the follow-up action 'Poll results with research_get.' However, it does not state when to choose this over alternatives like tavily_search, nor any preconditions or exclusions. Usage is only partially implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tavily_research_getA
Read-onlyIdempotent

Retrieve the status and results of a research task by request_id. HTTP 202 means still running; poll until HTTP 200.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestIdYesResearch task UUID returned by `research_create`.
includeUsageNoInclude credit usage in the response.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, open-world, and idempotent characteristics. The description adds useful behavioral detail beyond those annotations by explaining the HTTP 202 in-progress status and the need to poll until HTTP 200.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no wasted words. The purpose is front-loaded, followed immediately by the key polling behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter polling tool with full schema coverage and no output schema, the description covers the essential purpose and polling semantics. It does not describe the shape of returned results, but with no output schema that omission is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both requestId and includeUsage are already documented in the input schema. The description mentions request_id but adds no syntax, format, or usage detail beyond what the schema provides, making this baseline-level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Retrieve') and resource ('status and results of a research task') with the lookup key ('request_id'). It is clear enough to distinguish from generic search tools, though it does not explicitly name the sibling tavily_research_create as the origin of the request_id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational guidance for polling: HTTP 202 means still running, and the agent should poll until HTTP 200. It does not explicitly say when not to use this tool or name alternatives, but the intended context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv0.1.0
    • First observedconnect
    • First observedconnection_status
    • First observedfirecrawl_crawl
    • First observedfirecrawl_extract
    • First observedfirecrawl_monitor_checks_list
    • First observedfirecrawl_monitor_create
    • First observedfirecrawl_research_paper_get
    • First observedfirecrawl_research_papers_search
    • First observedfirecrawl_scrape
    • First observedgdocs_documents_create
    • First observedgithub_repos_list_languages
    • First observedgithub_search_repositories
    • First observedgsheets_spreadsheets_values_update
    • First observedlinear_customer_need_create
    • First observedlinear_issue_create
    • First observedlinear_search_issues
    • First observednotion_pages_create
    • First observedslack_chat_post_message
    • First observedtavily_research_create
    • First observedtavily_research_get
    • First observedtavily_search

TDQS

B3.2/5.0

Scored across 21 tools

Disambiguation4/5

Most tools target clearly distinct resources (web search, crawl, scrape, Notion pages, Slack messages, Linear issues, GitHub repos), and descriptions provide usage guidance. However, the web research cluster (tavily_search, firecrawl_scrape, firecrawl_extract, tavily_research_create) has overlapping purposes that an agent could confuse without careful reading.

Naming Consistency3/5

All names use snake_case with app prefixes, but the ordering is mixed: some are verb_noun (github_search_repositories, linear_search_issues) while others are noun_verb (notion_pages_create, gdocs_documents_create, linear_issue_create). The server-level tools 'connect' and 'connection_status' also lack the app-prefix pattern, making the set readable but not fully predictable.

Tool Count3/5

21 tools is on the heavy side for a server named 'Web Research to Docs', and the surface sprawls across web search, document creation, spreadsheets, Slack, Linear, GitHub, and connection management. While each tool could earn its place in a broad integration hub, the count feels over-scoped for the stated workflow.

Completeness2/5

Several tools reference operations that are not exposed: gdocs_documents_create points to documents_batch_update for adding content, firecrawl_extract points to extract_status for polling, and linear_issue_create requires teams_list/users_list/workflow_states_list that are absent. Create-only surfaces for Notion, Google Docs, and Sheets lack read/update/delete counterparts, leaving agents with dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to run local-first web research: intent-routed search across independent engines with reranking, a multi-stage fetch/crawl ladder, and document extraction. Results come back as signed-cursor, citation-bearing evidence envelopes, with an optional separately enabled profile for browser click/type actions.
    AGPL 3.0
  • A
    license
    C
    quality
    C
    maintenance
    Enables AI agents to run open-source intelligence workflows such as sanctions screening, prioritized vulnerability briefs, IOC searches, evidence-chain verification, and sun-position chronolocation.
    5
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables freelancers and agencies to run eight back-office workflows across Stripe, Google Drive, Linear, Google Calendar, Gmail, GitHub, Google Docs, Granola, Google Sheets and Firecrawl, covering client onboarding, invoices from calendar and commits, status reports, scope-creep detection, site audits and overdue invoice chasers. Reads run freely, while anything that creates, sends, changes or deletes is shown for approval first, with credentials kept in your own OS keychain and no proxying through any third-party server.
    Apache 2.0
  • A
    license
    B
    quality
    B
    maintenance
    Runs 13 agent slash-command workflows across Linear, GitHub, Google Docs, Slack, Firecrawl, Sheets, Tavily, Forms, Calendar, Gmail, Stripe, Granola and Notion — turning shipped features into blog and social drafts and handling competitor pricing, mention monitoring, SEO gaps, webinars, newsletters and launch-day tracking. Every credential is your own, kept in the OS keychain or client config, and anything that writes, sends or deletes is shown for approval first.
    29
    Apache 2.0