llmstxt-doc-search
Search and read llms.txt-based documentation sources live via MCP, with runtime source management.
Orient with
docs_home()- see registered sources and how to search/fetch.List sources with
list_doc_sources()- shows name,llms.txtURL, and index status.Search with
search_docs(query, source?, k?)- BM25-ranked results across all sources or one, returning{source, url, title, score, snippet}(default 5, max 50).Fetch with
fetch_doc(url)- pull the full live content of a result under a registered source.Add a source with
add_doc_source(name, llms_txt_url)- register and index a newllms.txtat runtime, persisted.Remove a source with
remove_doc_source(name).Refresh a source with
refresh_doc_source(name)- re-index to pick up new or changed docs.
Search covers seeded sources like Strands, Kiro, AWS Bedrock, Bedrock AgentCore, and MCP, plus anything you add.
Allows searching and fetching documentation from LangGraph's llms.txt index, providing ranked BM25 search results and on-demand content retrieval for LangGraph developer guides.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@llmstxt-doc-searchsearch for 'MCP transport' across all sources"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
llmstxt-doc-search
Live, ranked search across any number of
llms.txtdocumentation sites - Strands, Kiro, the AWS guides, and whatever you add at runtime.
llmstxt-doc-search is a Model Context Protocol (MCP) server that turns the llms.txt index a documentation site publishes into a fast, ranked search tool your agent can call. It indexes titles at startup, ranks queries with BM25, and fetches the full document only when you open a result - so you get current docs with almost no local storage. Built on the search engine from @praveenc/mcp-docs-server, generalized to a runtime registry of sources.
It implements the MCP 2026-07-28 specification over stdio and still works with clients on the 2025 protocol, which open with an initialize handshake. Requires Node.js 20 or later.
Why
An llms.txt file is a curated index of a doc site's pages, published for tools like this one to consume. They can be large - AWS Bedrock's lists roughly a thousand documents - so downloading everything is wasteful and goes stale fast.
This server takes a leaner approach:
Title-only index, built lazily. On first search of a source, only the page titles are indexed. That is fast to build and tiny to hold in memory.
Ranked with BM25. Queries are scored with BM25 plus Porter stemming, bigrams, and markdown-aware weighting (headers, code, and links count for more). Technical terms like
mcp,json, andstdioare preserved rather than stemmed.Content on demand. The full markdown or HTML of a result is fetched only when you call
fetch_doc.
The result is a good fit for broad, fast-moving reference material - the opposite tradeoff to snapshotting docs into a local vault.
Related MCP server: MCP Docs Server
Installation
Quick start (recommended)
Add the server to your MCP client configuration (Claude Desktop, Kiro, and others). It is downloaded and run on demand via npx - no manual build:
{
"mcpServers": {
"llmstxt-doc-search": {
"command": "npx",
"args": ["-y", "@praveenc/llmstxt-doc-search"]
}
}
}Global install
npm install -g @praveenc/llmstxt-doc-searchThen point your MCP client at the installed binary:
{
"mcpServers": {
"llmstxt-doc-search": {
"command": "llmstxt-doc-search"
}
}
}Quick start
Once the server is connected, the typical flow is three calls:
docs_home()- orient yourself: see the registered sources and how to search and fetch.search_docs("prompt caching", "aws-bedrock-userguide")- rank matching docs. Omit the source to search everything.fetch_doc(url)- read the full content of a result you like.
Add your own source at any time and it is indexed immediately and persisted for future runs:
add_doc_source("langgraph", "https://langchain-ai.github.io/langgraph/llms.txt")Tools
Tool | Purpose |
| Orientation: registered sources plus how to search and fetch. Call this first. |
| List sources with their |
| BM25 search. Omit |
| Fetch the full content of a result URL. The URL must be under the |
| Register and index a new |
| Remove a registered source. |
| Re-index a source to pick up new or changed docs. A source whose |
docs_home, list_doc_sources, search_docs, and fetch_doc are annotated read-only, so a client can approve them without prompting. add_doc_source, remove_doc_source, and refresh_doc_source change the persisted registry, and remove_doc_source is annotated destructive. The tool list is fixed, so it is advertised as cacheable for one hour.
Default sources
Seeded into the registry on first run:
strands, kiro, aws-bedrock-userguide, aws-agentic-ai-lens, aws-bedrock-agentcore-devguide, mcp.
The registry is persisted at ~/.config/llmstxt-doc-search/sources.json (override with LLMSTXT_REGISTRY_PATH). Anything you add, remove, or refresh at runtime is saved there.
Configuration
All configuration is via environment variables; none are required.
Variable | Default | Meaning |
|
| Where the source registry is persisted. |
|
| How many top hits to fetch when building result snippets. |
|
| Max fetched pages kept in memory per source (LRU); least-recently-used pages are evicted past this. |
|
| Log verbosity: |
Testing with MCP Inspector
npx @modelcontextprotocol/inspector npx -y @praveenc/llmstxt-doc-searchThe Inspector can connect in either protocol era; see Protocol eras.
Development
Clone the repository for local work (Node.js 20 or later):
git clone https://github.com/praveenc/llmstxt-doc-search.git
cd llmstxt-doc-search
npm installCommands
npm run dev # run from source with tsx (no build)
npm test # offline unit and protocol tests
npm run typecheck # type-check without emitting
npm run build # compile to dist/
npm run inspect:dev # MCP Inspector against the sourceLocal MCP client config (development)
Point your client at a source checkout instead of the published package:
{
"mcpServers": {
"llmstxt-doc-search": {
"command": "npx",
"args": ["tsx", "/ABS/PATH/llmstxt-doc-search/src/index.ts"]
}
}
}Or, after npm run build, at the compiled entry point:
{
"mcpServers": {
"llmstxt-doc-search": {
"command": "node",
"args": ["/ABS/PATH/llmstxt-doc-search/dist/index.js"]
}
}
}Architecture
src/
├── index.ts # Tool registration and the stdio entry point (serves 2026-07-28 and 2025-era clients)
├── config.ts # Defaults and environment configuration
├── tools/
│ └── docs.ts # search_docs, fetch_doc, and source management
└── utils/
├── doc-fetcher.ts # HTTP fetching, redirect handling, HTML parsing
├── indexer.ts # BM25 search index
├── registry.ts # Persisted source registry
├── store.ts # In-memory document store
├── text-processor.ts # Tokenization and snippet helpers
├── url-validator.ts # SSRF guard and URL validation
├── stopwords.ts # Stop-word list
└── logger.ts # Logging utilitiesSearch algorithm
Ranking uses BM25 (Best Matching 25) with several enhancements:
Porter stemming matches word variants (for example,
runningandrun).Bigrams capture phrase matches (for example,
prompt caching).Weighted scoring boosts title matches (3-8x), headers (4x), code blocks (2x), and link text (2x).
Domain-term preservation keeps technical terms like
mcp,json, andstdiounstemmed so they match exactly.
Security
This server fetches user-supplied URLs at runtime, so its SSRF surface is guarded in depth:
Scoped fetches.
fetch_doconly retrieves URLs that a registered source'sllms.txtlists exactly, or that sit under that source's origin and path prefix (matched on a path boundary rather than a raw string prefix). There is no arbitrary fetch.Scheme allow-list. Non-
http(s)schemes are rejected.Range-based address blocking. Private and reserved destinations are blocked using IP range classification (
ipaddr.js), covering decimal, octal, and hex IPv4, IPv4-mapped IPv6, loopback, link-local, unique-local, carrier-grade NAT, and other reserved ranges - not just a hostname regex.Connection-time validation. The resolved IP is checked at connection time via a custom DNS lookup, closing DNS-rebinding, and every redirect hop is re-validated.
Bounded responses. Response bodies are capped at 10 MB to limit memory and regular-expression (ReDoS) exposure.
Runtime dependencies report zero known vulnerabilities.
License
MIT - Copyright (c) 2026 Praveen Chamarthi
Contributing
Contributions are welcome. If you find a bug or have an idea:
Open an issue describing the problem or proposal.
For code changes, fork the repo and create a feature branch.
Keep changes focused, add or update tests, and make sure
npm test,npm run typecheck, andnpm run buildall pass.Open a pull request against
mainwith a clear description of what changed and why.
Commit messages follow the Conventional Commits style.
Support
Questions and ideas: open a GitHub issue.
Bugs: please include your MCP client, the tool call you made, and any relevant logs (set
LLMSTXT_LOG_LEVEL=debugfor more detail).Security issues: open an issue marked as security-sensitive, or contact the maintainer directly rather than posting exploit details publicly.
Available Tools
7 toolsadd_doc_sourceAdd doc sourceA
Register a new llms.txt source at runtime and index it. Persisted for future runs.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Short id, e.g. 'langgraph' | |
| llms_txt_url | Yes | URL of the source's llms.txt (https) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, open-world, non-idempotent, non-destructive operation, so the safety profile is covered. The description adds two useful facts beyond annotations: it triggers indexing and the registration is persisted for future runs. It does not address duplicate names, whether indexing is immediate or async, or error behavior, which a 4-5 would require.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action front-loaded and no filler; 'Persisted for future runs' efficiently conveys a durability guarantee. Nothing is wasted, though the second sentence leans slightly on terseness ('persisted' for what, exactly).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter registration tool with no output schema and full schema coverage, the description supplies the essential extras: indexing side effect and cross-run persistence. Only edge-case behavior (duplicate registration, failure modes) is left unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with both required parameters documented (name as a short id, llms_txt_url as an https URI). The description adds nothing about parameter meaning, format, or constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Register a new llms.txt source') plus the side effect of indexing, which clearly separates it from list_doc_sources, remove_doc_source and refresh_doc_source. It stops short of explicitly naming the sibling it complements, so sibling differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'at runtime' hints at when this applies (dynamic source registration vs. pre-configured sources), but there is no explicit when-to-use, when-not-to-use, or pointer to refresh_doc_source for already-registered sources. Usage must be inferred from the tool's name and behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
docs_homeDocs homeARead-only
Orientation: registered llms.txt sources + how to search/fetch. Call this first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so safety and scope are covered. The description adds that the payload is static orientation content (source registry + workflow guidance) rather than live data, which is mildly useful but no richer than the annotations themselves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Eleven words, front-loaded with the tool's identity and immediately followed by its call-to-action. Every fragment earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and annotations covering the safety profile, the description carries almost the whole burden and does state what the agent gets back. A hint about the format of the orientation payload would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema carries no semantics for the description to supplement. Baseline 4 applies; nothing is missing or misleading on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Orientation' names a specific resource type, and the description enumerates its content: registered llms.txt sources plus search/fetch guidance. It is distinguishable from list_doc_sources, which presumably only lists sources, but the overlap is not explicitly resolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this first' supplies an explicit ordering directive that tells the agent when to reach for this tool relative to the six sibling doc tools. No when-not condition or named alternative is offered, so it stops short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_docFetch docARead-only
Fetch full content of a doc url. The url must be under the llms.txt directory of a registered source, or listed by a source already searched (use search_docs first). Content is fetched live.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Document URL from a search_docs result |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered; the description adds that 'content is fetched live,' telling the agent results are not cached/stale snapshots. It doesn't describe failure modes or whether unreachable URLs raise or return empty, so it is solid but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and followed by the URL constraint; no filler. The second sentence packs two conditions together, which is dense but still readable and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with annotations covering the safety profile and no output schema, the description supplies the key operational constraint (valid URL origins) and the live-fetch behavior. Return-shape detail is not required, though error handling for out-of-scope URLs is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), and the description goes beyond the schema's 'Document URL from a search_docs result' by adding the alternative valid source (llms.txt directory of a registered source). That extra constraint meaningfully narrows what URLs are acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch full content of a doc url') and immediately scopes it to URLs from the llms.txt directory or a prior search_docs result, which cleanly separates it from the sibling search_docs. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisite flow ('use search_docs first') and the two valid URL origins, which is real routing guidance. It stops short of stating when NOT to use it (e.g., for non-URL lookups) or what happens if the URL is invalid/out of scope beyond the implicit failure.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_doc_sourcesList doc sourcesARead-only
List the registered documentation sources (name, llms.txt url, index status).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds value by disclosing the three fields returned (name, llms.txt url, index status), but says nothing about ordering, pagination, or whether the list can be empty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with the returned fields appended compactly. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with no output schema, the description supplies enough (including the shape of the returned data) to call it correctly. Minor gaps only: no sibling routing and no note on list size or ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline there is nothing for the description to disambiguate. The parenthetical field list describes the response, not inputs, which is still useful orientation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (registered documentation sources) and enumerates the fields returned, which clearly separates it from write-oriented siblings like add_doc_source and remove_doc_source. It stops short of naming an alternative to distinguish itself from search_docs or fetch_doc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no preconditions, and no mention of sibling tools. The intended use (enumerate sources before searching or fetching) is only implied by the verb 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_doc_sourceRefresh doc sourceAIdempotent
Re-index a source from its llms.txt (pick up new/changed docs).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Source name to refresh |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety and repeatability profile is covered. The description usefully adds that content is pulled from llms.txt (explaining the open-world fetch), but says nothing about duration, rate limits, or what happens to previously indexed content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and the clarifying mechanism tucked into a parenthetical. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter re-index tool with no output schema and annotations covering the safety profile, the description supplies enough to act correctly. A note on expected duration or whether the source must already exist would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema coverage is 100%, so the schema fully documents 'name' as the source to refresh. The description adds no format, naming-convention, or lookup detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('re-index') and resource ('a source'), with the mechanism ('from its llms.txt') and outcome ('pick up new/changed docs'). This clearly separates it from add_doc_source and remove_doc_source, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'pick up new/changed docs' implies the scenario for using it, but there is no explicit when-to-use-versus-alternatives statement or prerequisite guidance (e.g., after adding a source, or when docs changed upstream). Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_doc_sourceRemove doc sourceCDestructiveIdempotent
Remove a registered source.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Source name to remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered by structured data. The description adds nothing on top of that — it does not say whether previously ingested documents from the source are deleted, cached data invalidated, or how failures are reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero padding, and the verb+resource are front-loaded. It is efficient, though the extreme brevity leaves little room for the qualifications a destructive operation warrants.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with rich annotations and no output schema, the annotations already convey that this is a destructive, idempotent, closed-world mutation. The main unanswered question for an agent — whether removal also purges indexed content derived from that source — is left open, which is a meaningful gap for a destructive tool but not fatal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'name' parameter is documented as 'Source name to remove', so the schema carries the semantics. The description adds no extra information (e.g. whether an ID or display name is expected), which matches the baseline 3 when coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Remove') and resource ('a registered source'), so an agent knows exactly what the tool operates on. However, it offers no explicit differentiation from siblings such as add_doc_source or refresh_doc_source beyond the obvious verb contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus add_doc_source, refresh_doc_source, or list_doc_sources, nor any stated prerequisites (e.g. confirming the source is registered). The agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsSearch docsARead-only
BM25 search across registered llms.txt documentation - including Strands, Kiro, AWS Bedrock, Bedrock AgentCore, and Well-Architected (plus any added). Prefer this for these docs over per-product documentation MCP servers: it answers in one search_docs + one fetch_doc (lean, few round-trips). Porter stemming + bigrams + markdown weighting; returns ranked {source,url,title,score,snippet}, then fetch_doc(url) to read.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Max results (default 5, max 50) | |
| query | Yes | Search query, e.g. 'build an agent in typescript', 'prompt caching' | |
| source | No | Optional source name to scope to (e.g. 'strands', 'aws-bedrock-userguide'); omit to search all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover read-only/open-world safety, and the description adds real behavioral depth: the retrieval method (BM25, Porter stemming, bigrams, markdown weighting) and the exact ranked result shape {source,url,title,score,snippet} with no output schema present to convey it. This is decisive context for judging result quality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is dense but front-loaded with the core action and corpus before the justification and mechanics. Slightly packed into two long sentences, which costs a point on readability, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape and the chaining step to fetch_doc, and annotations cover safety. An agent has everything needed to call this correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so query/k/source are already documented with examples and defaults. The description adds the 'plus any added' scope note for sources but no extra syntax beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific mechanism and resource ('BM25 search across registered llms.txt documentation') and enumerates the covered corpora (Strands, Kiro, Bedrock, AgentCore, Well-Architected). It clearly separates itself from siblings like list_doc_sources and fetch_doc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to prefer this over per-product documentation MCP servers for these docs, and prescribes the intended workflow ('one search_docs + one fetch_doc'). Both the when-to-use and the follow-up alternative (fetch_doc) are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- Changed
add_doc_source2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
docs_home2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false
- Changed
fetch_doc2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
list_doc_sources2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false
- Changed
refresh_doc_source2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
remove_doc_source2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
search_docs2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
7 tool updates
v0.1.0- First observed
add_doc_source - First observed
docs_home - First observed
fetch_doc - First observed
list_doc_sources - First observed
refresh_doc_source - First observed
remove_doc_source - First observed
search_docs
TDQS
Scored across 7 tools
Each tool has a distinct role: search, fetch, list, add, remove, refresh, orientation. The only mild overlap is docs_home vs list_doc_sources, since both touch registered sources, but docs_home is orientation/guidance while list_doc_sources returns actual data.
Most tools follow a clear verb_noun pattern (list_doc_sources, search_docs, fetch_doc, add_doc_source, remove_doc_source, refresh_doc_source). docs_home breaks the pattern with a noun-only name, a minor deviation in an otherwise consistent set.
Seven tools is well-scoped for a documentation search server. Each tool earns its place: core search/fetch plus full source lifecycle management and an orientation entry point.
The surface covers the full lifecycle: orientation, listing sources, searching, fetching, and add/remove/refresh of sources with persisted indexing. No obvious gaps for a doc-search domain.
Maintenance
Related MCP Connectors
Any site's llms.txt: find the covering index, read its linked docs as markdown, search sections.
Search and read the Lium GPU rental docs: pods, CLI, SDK, REST API and agent guides.
Read-only search and discovery for the international llms.txt directory maintained by llmsmap.me.
Search and read Vector Panda docs: API operations, pricing, storage tiers, measured benchmarks.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables fast, token-efficient access to large documentation files in llms.txt format through semantic search. Solves token limit issues by searching first and retrieving only relevant sections instead of dumping entire documentation.3MIT
- AlicenseNot gradedqualityDmaintenanceAggregates documentation from multiple sources (llms.txt format or web scraping) and provides semantic search capabilities using vector embeddings and hybrid search for each documentation source.142 npmMIT
- AlicenseAqualityAmaintenanceWeb search (embedded SearXNG), content extraction, and library docs indexing with hybrid search. No API keys required.61,488 PyPI18Apache 2.0
- AlicenseAqualityCmaintenanceEnables LLM hosts to retrieve live, relevant documentation excerpts from official library docs sites via a search-and-RAG tool, avoiding reliance on training data.1MIT