Skip to main content
Glama
praveenc

llmstxt-doc-search

by praveenc

llmstxt-doc-search

Live, ranked search across any number of llms.txt documentation sites - Strands, Kiro, the AWS guides, and whatever you add at runtime.

npm version MCP Registry License: MIT Release

llmstxt-doc-search is a Model Context Protocol (MCP) server that turns the llms.txt index a documentation site publishes into a fast, ranked search tool your agent can call. It indexes titles at startup, ranks queries with BM25, and fetches the full document only when you open a result - so you get current docs with almost no local storage. Built on the search engine from @praveenc/mcp-docs-server, generalized to a runtime registry of sources.

It implements the MCP 2026-07-28 specification over stdio and still works with clients on the 2025 protocol, which open with an initialize handshake. Requires Node.js 20 or later.


Why

An llms.txt file is a curated index of a doc site's pages, published for tools like this one to consume. They can be large - AWS Bedrock's lists roughly a thousand documents - so downloading everything is wasteful and goes stale fast.

This server takes a leaner approach:

  • Title-only index, built lazily. On first search of a source, only the page titles are indexed. That is fast to build and tiny to hold in memory.

  • Ranked with BM25. Queries are scored with BM25 plus Porter stemming, bigrams, and markdown-aware weighting (headers, code, and links count for more). Technical terms like mcp, json, and stdio are preserved rather than stemmed.

  • Content on demand. The full markdown or HTML of a result is fetched only when you call fetch_doc.

The result is a good fit for broad, fast-moving reference material - the opposite tradeoff to snapshotting docs into a local vault.


Related MCP server: MCP Docs Server

Installation

Add the server to your MCP client configuration (Claude Desktop, Kiro, and others). It is downloaded and run on demand via npx - no manual build:

{
  "mcpServers": {
    "llmstxt-doc-search": {
      "command": "npx",
      "args": ["-y", "@praveenc/llmstxt-doc-search"]
    }
  }
}

Global install

npm install -g @praveenc/llmstxt-doc-search

Then point your MCP client at the installed binary:

{
  "mcpServers": {
    "llmstxt-doc-search": {
      "command": "llmstxt-doc-search"
    }
  }
}

Quick start

Once the server is connected, the typical flow is three calls:

  1. docs_home() - orient yourself: see the registered sources and how to search and fetch.

  2. search_docs("prompt caching", "aws-bedrock-userguide") - rank matching docs. Omit the source to search everything.

  3. fetch_doc(url) - read the full content of a result you like.

Add your own source at any time and it is indexed immediately and persisted for future runs:

add_doc_source("langgraph", "https://langchain-ai.github.io/langgraph/llms.txt")

Tools

Tool

Purpose

docs_home()

Orientation: registered sources plus how to search and fetch. Call this first.

list_doc_sources()

List sources with their llms.txt URL and index status, including the last index error for a failing source.

search_docs(query, source?, k?)

BM25 search. Omit source to search all, or scope to one. Returns ranked {source, url, title, score, snippet}, where score (0-1) is comparable across sources. k defaults to 5 (max 50).

fetch_doc(url)

Fetch the full content of a result URL. The URL must be under the llms.txt directory of a registered source, or listed in the llms.txt of a source that has been indexed (by a search, add, or refresh).

add_doc_source(name, llms_txt_url)

Register and index a new llms.txt source at runtime. Persisted. Rejected if that llms.txt is already registered or contains no links.

remove_doc_source(name)

Remove a registered source.

refresh_doc_source(name)

Re-index a source to pick up new or changed docs. A source whose llms.txt fails to index is skipped for 5 minutes; this retries it immediately.

docs_home, list_doc_sources, search_docs, and fetch_doc are annotated read-only, so a client can approve them without prompting. add_doc_source, remove_doc_source, and refresh_doc_source change the persisted registry, and remove_doc_source is annotated destructive. The tool list is fixed, so it is advertised as cacheable for one hour.

Default sources

Seeded into the registry on first run:

strands, kiro, aws-bedrock-userguide, aws-agentic-ai-lens, aws-bedrock-agentcore-devguide, mcp.

The registry is persisted at ~/.config/llmstxt-doc-search/sources.json (override with LLMSTXT_REGISTRY_PATH). Anything you add, remove, or refresh at runtime is saved there.


Configuration

All configuration is via environment variables; none are required.

Variable

Default

Meaning

LLMSTXT_REGISTRY_PATH

~/.config/llmstxt-doc-search/sources.json

Where the source registry is persisted.

LLMSTXT_SNIPPET_HYDRATE_MAX

5

How many top hits to fetch when building result snippets.

LLMSTXT_PAGE_CACHE_MAX

50

Max fetched pages kept in memory per source (LRU); least-recently-used pages are evicted past this. 0 disables the cap.

LLMSTXT_LOG_LEVEL

info

Log verbosity: debug, info, warn, or error. Logs go to stderr only.


Testing with MCP Inspector

npx @modelcontextprotocol/inspector npx -y @praveenc/llmstxt-doc-search

The Inspector can connect in either protocol era; see Protocol eras.


Development

Clone the repository for local work (Node.js 20 or later):

git clone https://github.com/praveenc/llmstxt-doc-search.git
cd llmstxt-doc-search
npm install

Commands

npm run dev         # run from source with tsx (no build)
npm test            # offline unit and protocol tests
npm run typecheck   # type-check without emitting
npm run build       # compile to dist/
npm run inspect:dev # MCP Inspector against the source

Local MCP client config (development)

Point your client at a source checkout instead of the published package:

{
  "mcpServers": {
    "llmstxt-doc-search": {
      "command": "npx",
      "args": ["tsx", "/ABS/PATH/llmstxt-doc-search/src/index.ts"]
    }
  }
}

Or, after npm run build, at the compiled entry point:

{
  "mcpServers": {
    "llmstxt-doc-search": {
      "command": "node",
      "args": ["/ABS/PATH/llmstxt-doc-search/dist/index.js"]
    }
  }
}

Architecture

src/
├── index.ts              # Tool registration and the stdio entry point (serves 2026-07-28 and 2025-era clients)
├── config.ts             # Defaults and environment configuration
├── tools/
│   └── docs.ts           # search_docs, fetch_doc, and source management
└── utils/
    ├── doc-fetcher.ts    # HTTP fetching, redirect handling, HTML parsing
    ├── indexer.ts        # BM25 search index
    ├── registry.ts       # Persisted source registry
    ├── store.ts          # In-memory document store
    ├── text-processor.ts # Tokenization and snippet helpers
    ├── url-validator.ts   # SSRF guard and URL validation
    ├── stopwords.ts      # Stop-word list
    └── logger.ts         # Logging utilities

Search algorithm

Ranking uses BM25 (Best Matching 25) with several enhancements:

  • Porter stemming matches word variants (for example, running and run).

  • Bigrams capture phrase matches (for example, prompt caching).

  • Weighted scoring boosts title matches (3-8x), headers (4x), code blocks (2x), and link text (2x).

  • Domain-term preservation keeps technical terms like mcp, json, and stdio unstemmed so they match exactly.


Security

This server fetches user-supplied URLs at runtime, so its SSRF surface is guarded in depth:

  • Scoped fetches. fetch_doc only retrieves URLs that a registered source's llms.txt lists exactly, or that sit under that source's origin and path prefix (matched on a path boundary rather than a raw string prefix). There is no arbitrary fetch.

  • Scheme allow-list. Non-http(s) schemes are rejected.

  • Range-based address blocking. Private and reserved destinations are blocked using IP range classification (ipaddr.js), covering decimal, octal, and hex IPv4, IPv4-mapped IPv6, loopback, link-local, unique-local, carrier-grade NAT, and other reserved ranges - not just a hostname regex.

  • Connection-time validation. The resolved IP is checked at connection time via a custom DNS lookup, closing DNS-rebinding, and every redirect hop is re-validated.

  • Bounded responses. Response bodies are capped at 10 MB to limit memory and regular-expression (ReDoS) exposure.

Runtime dependencies report zero known vulnerabilities.


License

MIT - Copyright (c) 2026 Praveen Chamarthi


Contributing

Contributions are welcome. If you find a bug or have an idea:

  1. Open an issue describing the problem or proposal.

  2. For code changes, fork the repo and create a feature branch.

  3. Keep changes focused, add or update tests, and make sure npm test, npm run typecheck, and npm run build all pass.

  4. Open a pull request against main with a clear description of what changed and why.

Commit messages follow the Conventional Commits style.


Support

  • Questions and ideas: open a GitHub issue.

  • Bugs: please include your MCP client, the tool call you made, and any relevant logs (set LLMSTXT_LOG_LEVEL=debug for more detail).

  • Security issues: open an issue marked as security-sensitive, or contact the maintainer directly rather than posting exploit details publicly.


Available Tools

7 tools
add_doc_sourceAdd doc sourceA

Register a new llms.txt source at runtime and index it. Persisted for future runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesShort id, e.g. 'langgraph'
llms_txt_urlYesURL of the source's llms.txt (https)

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, open-world, non-idempotent, non-destructive operation, so the safety profile is covered. The description adds two useful facts beyond annotations: it triggers indexing and the registration is persisted for future runs. It does not address duplicate names, whether indexing is immediate or async, or error behavior, which a 4-5 would require.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the action front-loaded and no filler; 'Persisted for future runs' efficiently conveys a durability guarantee. Nothing is wasted, though the second sentence leans slightly on terseness ('persisted' for what, exactly).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter registration tool with no output schema and full schema coverage, the description supplies the essential extras: indexing side effect and cross-run persistence. Only edge-case behavior (duplicate registration, failure modes) is left unstated, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with both required parameters documented (name as a short id, llms_txt_url as an https URI). The description adds nothing about parameter meaning, format, or constraints, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Register a new llms.txt source') plus the side effect of indexing, which clearly separates it from list_doc_sources, remove_doc_source and refresh_doc_source. It stops short of explicitly naming the sibling it complements, so sibling differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'at runtime' hints at when this applies (dynamic source registration vs. pre-configured sources), but there is no explicit when-to-use, when-not-to-use, or pointer to refresh_doc_source for already-registered sources. Usage must be inferred from the tool's name and behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_homeDocs homeA
Read-only

Orientation: registered llms.txt sources + how to search/fetch. Call this first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety and scope are covered. The description adds that the payload is static orientation content (source registry + workflow guidance) rather than live data, which is mildly useful but no richer than the annotations themselves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Eleven words, front-loaded with the tool's identity and immediately followed by its call-to-action. Every fragment earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and annotations covering the safety profile, the description carries almost the whole burden and does state what the agent gets back. A hint about the format of the orientation payload would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema carries no semantics for the description to supplement. Baseline 4 applies; nothing is missing or misleading on this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Orientation' names a specific resource type, and the description enumerates its content: registered llms.txt sources plus search/fetch guidance. It is distinguishable from list_doc_sources, which presumably only lists sources, but the overlap is not explicitly resolved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call this first' supplies an explicit ordering directive that tells the agent when to reach for this tool relative to the six sibling doc tools. No when-not condition or named alternative is offered, so it stops short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_docFetch docA
Read-only

Fetch full content of a doc url. The url must be under the llms.txt directory of a registered source, or listed by a source already searched (use search_docs first). Content is fetched live.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesDocument URL from a search_docs result

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered; the description adds that 'content is fetched live,' telling the agent results are not cached/stale snapshots. It doesn't describe failure modes or whether unreachable URLs raise or return empty, so it is solid but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and followed by the URL constraint; no filler. The second sentence packs two conditions together, which is dense but still readable and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with annotations covering the safety profile and no output schema, the description supplies the key operational constraint (valid URL origins) and the live-fetch behavior. Return-shape detail is not required, though error handling for out-of-scope URLs is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), and the description goes beyond the schema's 'Document URL from a search_docs result' by adding the alternative valid source (llms.txt directory of a registered source). That extra constraint meaningfully narrows what URLs are acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch full content of a doc url') and immediately scopes it to URLs from the llms.txt directory or a prior search_docs result, which cleanly separates it from the sibling search_docs. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the prerequisite flow ('use search_docs first') and the two valid URL origins, which is real routing guidance. It stops short of stating when NOT to use it (e.g., for non-URL lookups) or what happens if the URL is invalid/out of scope beyond the implicit failure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_doc_sourcesList doc sourcesA
Read-only

List the registered documentation sources (name, llms.txt url, index status).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds value by disclosing the three fields returned (name, llms.txt url, index status), but says nothing about ordering, pagination, or whether the list can be empty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, with the returned fields appended compactly. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with no output schema, the description supplies enough (including the shape of the returned data) to call it correctly. Minor gaps only: no sibling routing and no note on list size or ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline there is nothing for the description to disambiguate. The parenthetical field list describes the response, not inputs, which is still useful orientation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (registered documentation sources) and enumerates the fields returned, which clearly separates it from write-oriented siblings like add_doc_source and remove_doc_source. It stops short of naming an alternative to distinguish itself from search_docs or fetch_doc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no preconditions, and no mention of sibling tools. The intended use (enumerate sources before searching or fetching) is only implied by the verb 'List'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_doc_sourceRefresh doc sourceA
Idempotent

Re-index a source from its llms.txt (pick up new/changed docs).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSource name to refresh

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety and repeatability profile is covered. The description usefully adds that content is pulled from llms.txt (explaining the open-world fetch), but says nothing about duration, rate limits, or what happens to previously indexed content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and the clarifying mechanism tucked into a parenthetical. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter re-index tool with no output schema and annotations covering the safety profile, the description supplies enough to act correctly. A note on expected duration or whether the source must already exist would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter exists and schema coverage is 100%, so the schema fully documents 'name' as the source to refresh. The description adds no format, naming-convention, or lookup detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('re-index') and resource ('a source'), with the mechanism ('from its llms.txt') and outcome ('pick up new/changed docs'). This clearly separates it from add_doc_source and remove_doc_source, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'pick up new/changed docs' implies the scenario for using it, but there is no explicit when-to-use-versus-alternatives statement or prerequisite guidance (e.g., after adding a source, or when docs changed upstream). Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_doc_sourceRemove doc sourceC
DestructiveIdempotent

Remove a registered source.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSource name to remove

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered by structured data. The description adds nothing on top of that — it does not say whether previously ingested documents from the source are deleted, cached data invalidated, or how failures are reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero padding, and the verb+resource are front-loaded. It is efficient, though the extreme brevity leaves little room for the qualifications a destructive operation warrants.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with rich annotations and no output schema, the annotations already convey that this is a destructive, idempotent, closed-world mutation. The main unanswered question for an agent — whether removal also purges indexed content derived from that source — is left open, which is a meaningful gap for a destructive tool but not fatal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'name' parameter is documented as 'Source name to remove', so the schema carries the semantics. The description adds no extra information (e.g. whether an ID or display name is expected), which matches the baseline 3 when coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Remove') and resource ('a registered source'), so an agent knows exactly what the tool operates on. However, it offers no explicit differentiation from siblings such as add_doc_source or refresh_doc_source beyond the obvious verb contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus add_doc_source, refresh_doc_source, or list_doc_sources, nor any stated prerequisites (e.g. confirming the source is registered). The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_docsSearch docsA
Read-only

BM25 search across registered llms.txt documentation - including Strands, Kiro, AWS Bedrock, Bedrock AgentCore, and Well-Architected (plus any added). Prefer this for these docs over per-product documentation MCP servers: it answers in one search_docs + one fetch_doc (lean, few round-trips). Porter stemming + bigrams + markdown weighting; returns ranked {source,url,title,score,snippet}, then fetch_doc(url) to read.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoMax results (default 5, max 50)
queryYesSearch query, e.g. 'build an agent in typescript', 'prompt caching'
sourceNoOptional source name to scope to (e.g. 'strands', 'aws-bedrock-userguide'); omit to search all

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover read-only/open-world safety, and the description adds real behavioral depth: the retrieval method (BM25, Porter stemming, bigrams, markdown weighting) and the exact ranked result shape {source,url,title,score,snippet} with no output schema present to convey it. This is decisive context for judging result quality.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is dense but front-loaded with the core action and corpus before the justification and mechanics. Slightly packed into two long sentences, which costs a point on readability, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape and the chaining step to fetch_doc, and annotations cover safety. An agent has everything needed to call this correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so query/k/source are already documented with examples and defaults. The description adds the 'plus any added' scope note for sources but no extra syntax beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific mechanism and resource ('BM25 search across registered llms.txt documentation') and enumerates the covered corpora (Strands, Kiro, Bedrock, AgentCore, Well-Architected). It clearly separates itself from siblings like list_doc_sources and fetch_doc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to prefer this over per-product documentation MCP servers for these docs, and prescribes the intended workflow ('one search_docs + one fetch_doc'). Both the when-to-use and the follow-up alternative (fetch_doc) are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv1.0.0
    • Changedadd_doc_source2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changeddocs_home2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedfetch_doc2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changedlist_doc_sources2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedrefresh_doc_source2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changedremove_doc_source2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changedsearch_docs2 fields changed
      • changedInput schema / $schema
        Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
      • removedInput schema / additionalProperties
        Removed value: -false
  2. 7 tool updatesv0.1.0
    • First observedadd_doc_source
    • First observeddocs_home
    • First observedfetch_doc
    • First observedlist_doc_sources
    • First observedrefresh_doc_source
    • First observedremove_doc_source
    • First observedsearch_docs

TDQS

A3.7/5.0

Scored across 7 tools

Disambiguation4/5

Each tool has a distinct role: search, fetch, list, add, remove, refresh, orientation. The only mild overlap is docs_home vs list_doc_sources, since both touch registered sources, but docs_home is orientation/guidance while list_doc_sources returns actual data.

Naming Consistency4/5

Most tools follow a clear verb_noun pattern (list_doc_sources, search_docs, fetch_doc, add_doc_source, remove_doc_source, refresh_doc_source). docs_home breaks the pattern with a noun-only name, a minor deviation in an otherwise consistent set.

Tool Count5/5

Seven tools is well-scoped for a documentation search server. Each tool earns its place: core search/fetch plus full source lifecycle management and an orientation entry point.

Completeness5/5

The surface covers the full lifecycle: orientation, listing sources, searching, fetching, and add/remove/refresh of sources with persisted indexing. No obvious gaps for a doc-search domain.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables fast, token-efficient access to large documentation files in llms.txt format through semantic search. Solves token limit issues by searching first and retrieving only relevant sections instead of dumping entire documentation.
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Aggregates documentation from multiple sources (llms.txt format or web scraping) and provides semantic search capabilities using vector embeddings and hybrid search for each documentation source.
    142 npm
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables LLM hosts to retrieve live, relevant documentation excerpts from official library docs sites via a search-and-RAG tool, avoiding reliance on training data.
    1
    MIT