js-ts-mcp
Fetches live MDN Web Docs pages by slug or resolved identifier, converts them to clean markdown, supports section filtering, and provides ranked search over a local MDN documentation index.
Retrieves npm package registry metadata and README, including resolved version, description, license, homepage, repository URL, keywords, engines, and dependencies, optionally pinned to a specific version.
Fetches TypeScript handbook pages live from typescriptlang.org and includes them in the local search index, enabling lookup of handbook topics such as intro or generics.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@js-ts-mcppull the live MDN docs for Array.prototype.flat"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
js-ts-mcp
An MCP (Model Context Protocol) server that gives AI coding agents live MDN Web Docs pages, the TypeScript handbook and npm package information instead of whatever the model happens to remember. Pages are fetched at request time, cached locally, and returned as clean markdown — so generated JavaScript and TypeScript target real, current APIs rather than deprecated or invented ones.
MDN Web Docs — any docs page fetched live by its slug (e.g.
Web/JavaScript/Reference/Global_Objects/Array) and converted to markdownTypeScript handbook — all 117 handbook pages (e.g.
intro,2/generics) scraped live from typescriptlang.orgnpm packages — registry metadata + README from the npm registry, optionally pinned to a version (see npm source)
Local search index over 14,708 MDN docs (from the en-US sitemap) + 117 TypeScript handbook pages (name + path recorded, so
Promiseresolves to the right page on its own), rebuilt at most every 7 days with a stale fallback when the network failsSQLite TTL cache so repeated lookups are instant
Politeness layer — robots.txt (RFC 9309), per-host throttle,
Retry-After, conditional GET, request budgets and a host allowlist (details)
The five tools
Tool | What it does |
| Resolve one identifier ( |
| Ranked multi-token search over the local index of 14,825 MDN + TypeScript entries (name and path). |
| npm package metadata + README: resolved version, description, license, homepage, repository URL, keywords, engines, dependencies. |
| Real health check: index size/age, cache stats, live probes of all three hosts, and politeness counters. |
| Server liveness and version. |
Every tool returns a plain dict. Failures come back as {"ok": false, "error": …, "suggestion": …} — a bad lookup never raises into the MCP layer, and no tool delegates to another tool.
Related MCP server: docs-mcp
Requirements
Python 3.10 or newer (developed and tested on 3.12)
uvfor the one-command install below (uvxships with it)The MCP Python SDK 1.x is pinned (
mcp>=1.2.0,<2.0); this package does not support the mcp 2.x rename ofFastMCPhttpxis the only HTTP transport. There is nocurlfallback and no browser user-agent masking — why
Install & run
One command — no checkout, no token
uvx --from git+https://github.com/KEEPEE/js-ts-mcp.git js-ts-docsThis builds the package in an isolated environment and starts the stdio MCP server. It prints nothing on purpose: stdout is the protocol channel. Stop it with Ctrl-C.
After the repository is updated, force uv to re-resolve the commit:
uvx --refresh --from git+https://github.com/KEEPEE/js-ts-mcp.git js-ts-docsLocal checkout — fastest startup, editable while developing
git clone https://github.com/KEEPEE/js-ts-mcp.git
cd js-ts-mcp
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/js-ts-docs # stdio MCP serverMCP client configuration
Generic stdio client (Claude Desktop, Cursor, Cline, …)
{
"mcpServers": {
"js-ts": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/KEEPEE/js-ts-mcp.git",
"js-ts-docs"
]
}
}
}With an explicit cache directory (any env you set is passed straight through to the server):
{
"mcpServers": {
"js-ts": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/KEEPEE/js-ts-mcp.git",
"js-ts-docs"
],
"env": {
"JS_TS_MCP_CACHE_DIR": "/tmp/js-ts-cache"
}
}
}
}DeepSeek Harness (cordis-style plugin list)
- id: mcp-js-ts
name: '@deepseek-ai/dsh-mcp-client'
config:
serverName: js-ts
transport: stdio
command: uvx
args:
[
'--from',
'git+https://github.com/KEEPEE/js-ts-mcp.git',
'js-ts-docs'
]To skip the build at client start, point command at the console script of a checkout instead — command: /path/to/js-ts-mcp/.venv/bin/js-ts-docs with args: [].
Tools reference
js_docs(identifier, topic=None, max_tokens=8000)
Resolves an identifier to one docs page. Identifier forms, tried in this order:
Form | Meaning |
| fetched directly from that MDN slug |
| fetched directly as a TypeScript handbook page |
| resolved through the local index: only the top-score candidates are considered; a single distinct path wins, ties prefer MDN |
topic keeps only the section whose heading matches it (Methods, Examples, …) plus the page title and description; if nothing matches, the full page is returned with a note listing the available headings. max_tokens is a rough budget (1 token ≈ 4 characters): the markdown is cut at a line boundary and truncated is set.
Returns {"ok", "identifier", "url", "source", "title", "markdown", "truncated", "cached"}, plus an optional note.
// js_docs({ "identifier": "mdn:Web/JavaScript/Reference/Global_Objects/Array", "max_tokens": 200 })
{
"ok": true,
"identifier": "mdn:Web/JavaScript/Reference/Global_Objects/Array",
"url": "https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array",
"source": "mdn",
"markdown": "# Array\n\nThe **`Array`** object, as with arrays in other programming languages, enables [storing a collection of multiple items under a single variable name](…), and has members for [performing common array operations](#examples). ## [Description](#description) In JavaScript, arrays aren't [primitives](…) but are instead `Array` objects with the following core characteristics: … [truncated: showing ~130 of ~16043 estimated tokens]",
"truncated": true,
"cached": false,
"title": "Array - JavaScript"
}A bare name is resolved through the index — no guessing:
// js_docs({ "identifier": "Promise", "max_tokens": 160 })
{
"ok": true,
"identifier": "Promise",
"url": "https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Promise",
"source": "mdn",
"title": "Promise - JavaScript",
"markdown": "# Promise\n\nThe **`Promise`** object represents the eventual completion (or failure) of an asynchronous operation and its resulting value. … [truncated: showing ~89 of ~8695 estimated tokens]",
"truncated": true,
"cached": false
}A TypeScript handbook page:
// js_docs({ "identifier": "ts:intro", "max_tokens": 160 })
{
"ok": true,
"identifier": "ts:intro",
"url": "https://www.typescriptlang.org/docs/handbook/intro.html",
"source": "typescript",
"title": "The TypeScript Handbook",
"markdown": "# The TypeScript Handbook\n\n## About this Handbook\n … [truncated: showing ~12 of ~1347 estimated tokens]",
"truncated": true,
"cached": false
}A topic that matches nothing tells you what the page does have:
// js_docs({ "identifier": "ts:intro", "topic": "nonexistent-section-xyz", "max_tokens": 120 })
{
"ok": true,
"…": "…",
"note": "no section matching 'nonexistent-section-xyz'; available headings: The TypeScript Handbook, About this Handbook, How is this Handbook Structured, Non-Goals, Get Started"
}A name that is not in the index is an error dict, not an empty page:
// js_docs({ "identifier": "nope-not-a-real-page-xyz" })
{
"ok": false,
"error": "no MDN/TypeScript page matching 'nope-not-a-real-page-xyz' found in the search index",
"suggestion": "use js_search to find the exact page, or use the mdn:/ts: prefix"
}js_search(query, limit=8)
Multi-token ranked search over the local index: every token must match, and both the entry name and its path are searched. Returns {"ok", "query", "results": [{source, name, path, score}], "index_count", "stale"}. Score tiers: exact name 10 > name prefix 8 > substring in name 6 > substring in path 3, summed per token.
// js_search({ "query": "regular expression", "limit": 3 })
{
"ok": true,
"query": "regular expression",
"results": [
{ "source": "mdn", "name": "Regular expression", "path": "Glossary/Regular_expression", "score": 14.0 },
{ "source": "mdn", "name": "Regular expressions", "path": "Web/JavaScript/Guide/Regular_expressions", "score": 14.0 },
{ "source": "mdn", "name": "Regular expressions", "path": "Web/JavaScript/Reference/Regular_expressions", "score": 14.0 }
],
"index_count": 14825,
"stale": false
}The index spans both sources, so a TypeScript term comes back from the handbook:
// js_search({ "query": "generics", "limit": 3 })
{ "results": [ { "source": "typescript", "name": "generics", "path": "2/generics", "score": 10.0 } ], "index_count": 14825, "…": "…" }An ambiguous name returns every candidate and lets you pick — this is why the mdn:/ts: prefix exists:
// js_search({ "query": "promise", "limit": 3 })
{
"results": [
{ "source": "mdn", "name": "Promise", "path": "Glossary/Promise", "score": 10.0 },
{ "source": "mdn", "name": "promise", "path": "Web/API/PromiseRejectionEvent/promise", "score": 10.0 },
{ "source": "mdn", "name": "Promise", "path": "Web/JavaScript/Reference/Global_Objects/Promise", "score": 10.0 }
], "…": "…"
}js_search costs zero requests while the cached index is fresh (measured warm: 0 requests, 0 budget units). If the index is stale the call rebuilds it inside the same budget — measured cold: 4 units.
npm_package(name, version=None)
npm metadata + README. name is the package name; version pins one release. Returns {"ok", "name", "version" (resolved), "description", "license", "homepage", "repository_url", "keywords", "engines", "dependencies", "readme_markdown", "url"}. Cached 1 day, keyed by name + version.
// npm_package({ "name": "left-pad" })
{
"ok": true,
"name": "left-pad",
"version": "1.3.0",
"description": "String left pad",
"license": "WTFPL",
"homepage": "https://github.com/stevemao/left-pad#readme",
"repository_url": "git+ssh://git@github.com/stevemao/left-pad.git",
"keywords": ["leftpad", "left", "pad", "padding", "string", "repeat"],
"engines": null,
"dependencies": {},
"readme_markdown": "## left-pad\n\nString left pad\n…\n```js\nconst leftPad = require('left-pad')\n\nleftPad('foo', 5)\n// => \" foo\"\n``` …",
"url": "https://www.npmjs.com/package/left-pad",
"cached": false
}A pinned version resolves to exactly that release. Note that the registry does not always carry a README for a non-latest release — readme_markdown is then null rather than a guess:
// npm_package({ "name": "typescript", "version": "5.6.3" })
{
"ok": true,
"name": "typescript",
"version": "5.6.3",
"description": "TypeScript is a language for application scale JavaScript development",
"license": "Apache-2.0",
"homepage": "https://www.typescriptlang.org/",
"repository_url": "git+https://github.com/microsoft/TypeScript.git",
"engines": { "node": ">=14.17" },
"readme_markdown": null,
"url": "https://www.npmjs.com/package/typescript/5.6.3",
"…": "…"
}An unknown name is an error dict:
// npm_package({ "name": "nope-not-a-real-package-xyz" })
{
"ok": false,
"error": "npm lookup failed: package not found on npm",
"suggestion": "check the package name"
}url is the human-facing www.npmjs.com page. This server reports that URL but never requests it — the request goes to registry.npmjs.org, which is the only npm host in the allowlist (why).
js_status()
Probes developer.mozilla.org, www.typescriptlang.org and registry.npmjs.org with a light GET (all on robots-allowed paths) and reports the index and cache state. overall is ok, degraded or error. The politeness block is diagnostics only and never changes overall.
// js_status() — cold cache, live
{
"server": "js-ts-mcp",
"version": "0.2.0",
"checks": {
"search_index": { "status": "ok", "entries": 14825, "mdn_count": 14708, "ts_count": 117, "built_at": "2026-10-07T08:24:04.455645+00:00", "stale": false },
"cache": { "status": "ok", "entries": 3, "expired": 0 },
"mdn_org": { "status": "ok", "http_status": 200 },
"typescriptlang_org": { "status": "ok", "http_status": 200 },
"npm_registry": { "status": "ok", "http_status": 200 }
},
"politeness": {
"status": "ok", "disabled": false,
"requests": 5, "robots_requests": 3, "robots_rows": 3, "robots_fetches": 3, "robots_cache_hits": 2,
"robots_negative": 1, "blocked_by_robots": 0,
"throttle_waits": 4, "throttle_sleep_s": 2.341,
"host_delays": { "developer.mozilla.org": 0.0, "www.typescriptlang.org": 0.0, "registry.npmjs.org": 0.0 },
"budgets": { "tool:js_status": [8, 12] }, "budget_denied": 0,
"conditional": 0, "conditional_skipped": 0, "revalidated_304": 0,
"retries_429": 0, "retries_transport": 0, "stalls": 0,
"cached_bodies": 3, "max_cached_bytes": 2097152,
"robots_db": "/tmp/js-ts-cache/robots.db"
},
"overall": "ok"
}robots_negative: 1 on a cold run is the TypeScript handbook 404 — see two more robots facts.
health_check()
{"status": "ok", "server": "js-ts-mcp", "version": "0.2.0"} — no network, no cache; safe as a liveness probe.
npm metadata source (and its robots.txt trap)
registry.npmjs.org/robots.txt is not a robots file
This is the single strangest thing about crawling npm, and this server is explicit about it. Verified live on 2026-10-07:
GET https://registry.npmjs.org/robots.txt
status=200 content-type=application/json bytes=7462
first 80 bytes: {"_id":"robots.txt","_rev":"24-5090ef0ecb2d83c69ed67cf442a0e74d","name":"robots.The registry answers HTTP 200 with JSON — the packument of an npm package literally named robots.txt. There is no robots policy there at all.
Naive code that fetches /robots.txt and feeds whatever comes back into a robots parser is one step away from inventing rules out of package metadata. Worse, a package whose README happens to contain robots-looking lines could make a self-blocking rule appear. The layer runs the real captured body through its RFC 9309 parser and gets zero groups and zero rules, so registry.npmjs.org is treated as a host with no rules and the request proceeds:
parse_robots(<the 7,462-byte body>) -> []
can_fetch("https://registry.npmjs.org/react") -> TrueHow the test suite pins this (tests/test_politeness_wiring.py, all four run offline against the captured file tests/fixtures/npm_robots_txt_packument.json):
test | what it asserts |
| the captured body really is JSON ( |
| a synthetic packument whose |
| end-to-end through |
| the trap body is stored like any other robots record, so the registry costs one robots request per 7 days, not one per call |
Two more robots facts, both negatively cached
https://www.npmjs.com/robots.txtis a Cloudflare 403. That host is deliberately not in the allowlist: this server only reportswww.npmjs.comURLs (theurlfield ofnpm_package), it never requests them. With the production allowlist the layer refuses such a request outright and makes no network call at all — measured live,layer.get("https://www.npmjs.com/robots.txt")returnserror='host www.npmjs.com is not in ALLOWED_HOSTS'with zero bytes on the wire. The 403 itself is still negatively cached by the layer (one robots request per host per 7 days, not one per call), which is asserted offline bytest_www_npmjs_com_403_robots_is_negatively_cachedandtest_www_npmjs_com_is_not_in_the_allowlist.https://www.typescriptlang.org/robots.txtis a 404 (GitHub Pages serves an HTML 404 page, measured 9,379 bytes). Also negatively cached:robots_negative: 1injs_status().politenesson a cold run, and one robots request per week thereafter — asserted bytest_typescriptlang_robots_404_is_negatively_cached.
MDN's robots.txt is the only one of the three that publishes real rules — 119 bytes, one User-agent: * group, three Disallow lines: /api/, /*/files/ and /media. /api/ matters here: this server used to call developer.mozilla.org/api/v1/docs, which is both dead and disallowed. The layer refuses it with no request made — verified live: can_fetch("https://developer.mozilla.org/api/v1/docs/Array") → False, can_fetch("https://developer.mozilla.org/en-US/files/x") → False, while the docs paths this server actually uses return True.
One transport: no curl fallback, no browser user-agent masking
Both of these were in earlier versions of this package. Both are gone, deliberately, and a test keeps them gone.
No curl subprocess fallback. An earlier js_status shelled out to curl when httpx failed. That second transport bypassed robots.txt, the throttle, the request budget and the stored ETag/Last-Modified validators — so a fallback request was by construction an unpolite request. It also doubled the worst case, and the premise it was written for (that some CDN edges throttle Python's TLS fingerprint) did not reproduce under measurement. test_no_subprocess_or_curl_fallback_left_in_the_sources checks this on the AST of fetchers.py, server.py and search.py: no subprocess import, no "curl" argv string. The docstrings still explain why it is gone; the code does not contain it. test_fetchers_build_exactly_one_http_client pins the other half — fetchers.py constructs exactly one httpx.Client and never calls client.get() directly, so every request goes _client() → _get() → the politeness layer, and there is no second path that could skip it.
No browser user-agent spoofing. The module used to send a Chrome 126 UA. It buys nothing: these three sites do not gate on it, and it actively hurts politeness, because a robots file can only match a product token it can see — a User-agent: *-specific rule, a Crawl-delay aimed at a named crawler, or a site's own abuse counter all key off the UA. A spoof also hides who is knocking. The layer now sends one honest, contactable UA on every request, robots.txt included:
js-ts-mcp/0.2 (+https://github.com/KEEPEE/js-ts-mcp)That string is defined once, in src/js_ts_mcp/fetchers.py, and every test refers to it through the constant (tests/conftest.py, tests/test_politeness_wiring.py). test_honest_user_agent_is_what_reaches_the_wire asserts it on the wire: every request the repo makes carries exactly fetchers_mod.USER_AGENT, the string starts with js-ts-mcp/, contains github.com/KEEPEE/js-ts-mcp, contains no Mozilla//Chrome//Safari//AppleWebKit token, and the robots.txt fetch itself carries it — because that is the token a robots file matches on.
Caching & politeness
Cache
Everything lives in one directory: ~/.cache/js-ts-mcp by default, overridable with JS_TS_MCP_CACHE_DIR.
File | Contents | TTL |
| docs pages (parsed result + raw body + | docs 7 days, search index 7 days, npm metadata 1 day |
| the politeness layer's robots.txt cache | 7 days per host |
One env var moves both. Delete the files to force a full refresh. A cache problem is never a tool failure: if the directory is unwritable the server runs without a cache and says so in js_status().checks.cache — covered by test_npm_package_and_status_survive_an_unwritable_cache_dir.
Politeness
Every outbound request — js_docs, js_search (its index build), npm_package and the js_status probes — goes through one small internal module, src/js_ts_mcp/politeness.py (stdlib + httpx, no extra dependency):
robots.txt is read and obeyed. Rules are matched per RFC 9309 (
*, trailing$,%2A/%24literals, mergedUser-agentgroups, most-specific match wins,AllowbeatsDisallowon a tie).urllib.robotparseris deliberately not used: on Python < 3.14 it reads*as a literal character, which is exactly how this server once requesteddeveloper.mozilla.org/api/v1/docs— a path MDN's ownrobots.txtdisallows. Files are cached 7 days per host inrobots.db, including negative results (404 / 403 / 5xx / unparseable body), so a host costs one robots request per week — which is what makes the npm and TypeScript handbook quirks above cheap instead of per-call.Per-host throttle, including
Crawl-delay. Sequential requests to one host are spaced 0.35–0.9 s by default, or by the site's ownCrawl-delaywhen it declares one — none of the three hosts declares aCrawl-delaytoday (checked against all three live robots responses), so the default window applies. Therobots.txtfetch queues in the same per-host line as a page request, and aCrawl-delaylearned from a robots fetch gates the request that triggered the fetch, not only the next one. Measured coldjs_docs("Promise"):throttle_sleep_s: 1.382across 2 waits.429/503/504are handled.Retry-Afteris honoured to the second; without it the per-host delay is escalated and the request retried once. A response that stalls past 10 s escalates the delay instead of being retried blindly.Conditional GET.
ETag/Last-Modifiedand the raw body are stored next to the cached entry, so refreshing an expired page costs a304instead of re-downloading it. Verified against live headers: all three hosts send bothETagandLast-Modified(MDN"2f13939d…", typescriptlang.orgW/"6ac39afa-2fe88", the npm registryW/"c7ff9479…"). Measured live: a coldjs_docs("ts:intro")after the index build had already fetched that same page revalidated it with one304and 0 bytes (conditional: 1,revalidated_304: 1). No conditional headers are ever sent to a host that sent no validators (test_host_without_validators_never_sends_conditional_headers).Request budgets — one unit is one request on the wire. One tool call makes at most 7 requests; the robots.txt fetch of a cold host, a
429retry, a transport retry and every redirect hop each pay a unit. Measured live with a cold cache:js_docs("mdn:…Array")= 2 (robots + page),js_docs("Promise")= 5 (MDN robots + sitemap + TypeScript robots + handbook page + the page itself — the index rebuild is what makes a plain name expensive),js_docs("ts:intro")= 2, an unknown identifier = 4 (the index build, then the index answers),js_search(…)= 4 cold and 0 warm,npm_package(…)= 2.js_statushas its own cap of 12 (measured cold 8, warm 3) so its probes plus a stale-index rebuild always fit — a probe refused by the budget would reporterrorand dragoverallto"degraded"for a purely internal reason. The search-index build shares the tool call's budget; a build that hits its cap stops early and returns a partial index ("partial": true,"partial_reason": …) instead of raising, and a partial index is never cached.Redirects are resolved by the layer, not by httpx. Requests go out with
follow_redirects=False, every 3xx target is re-checked against that host's robots rules before its body is used, the chain is capped at 5 hops and reported as an error (never an exception), and theurla tool returns is the final URL of the chain.Body-size cap: 2 MiB. Raw bodies are only kept for revalidation when they are under
MAX_CACHED_BODY_BYTES = 2 * 1024 * 1024. Why that number.Host allowlist. Only
developer.mozilla.org,www.typescriptlang.organdregistry.npmjs.orgcan be contacted, so a malformed identifier or an unexpected redirect cannot turn a docs lookup into a request somewhere else.
The layer never raises and never changes a tool's return shape; js_status reports its counters under the top-level politeness key.
Why 2 MiB for the body cap
The cap is a measurement, not a taste. The politeness layer stores the raw body next to the cached entry so an expired entry can be revalidated with If-None-Match and answered with a 304 and zero bytes. That only pays if the body is worth holding in memory and on disk — and the two biggest artifacts this server touches are four orders of magnitude apart in value:
artifact | measured size | kept? |
MDN en-US sitemap ( | 127,027 B on the wire today (126,725 B in the captured fixture), 1,940,079 B decoded today — 1,936,390 B for the fixture | yes — it is the whole search index, and revalidating it with a |
npm | 7,016,423 B in the audit, 7,021,259 B measured 2026-10-07 | no — it is parsed into a few hundred bytes of metadata and then has no further use |
npm | 15,762,037 B measured 2026-10-07 | no |
So the cap sits deliberately between the two: big enough that the decoded sitemap (1.94 MB) still fits and can be revalidated, small enough that packuments — which grow with every release ever published, and are already 7 MB and 15 MB for two ordinary packages — are never held. test_cap_sits_between_the_two_real_artifacts asserts exactly that relationship against the real fixture sizes, so the number cannot drift away from the measurement that justifies it.
Worth saying plainly: the sitemap is the tight side of that margin. At 1,940,079 bytes decoded it is within 8% of the 2 MiB cap, so if MDN's en-US index keeps growing the sitemap will eventually stop being revalidated and the index rebuild will re-download it every week. That is a deliberate trade — a weekly re-download is a cost, an unbounded in-memory body is a worse one — and the test is what makes the margin visible rather than folklore.
An oversized body is simply not remembered: the request still succeeds, the caller still gets the whole body, and the next call re-downloads it. test_oversized_body_is_not_cached_but_the_request_still_works asserts the full body is returned, cached_bodies == 0, conditional == 0, and that no revalidation row was written. test_production_layer_caps_cached_bodies_at_2mb asserts the production layer actually passes the cap (the library default is higher).
Opt-out
export JS_TS_MCP_POLITENESS_DISABLED=1This single switch turns off robots.txt, the throttle, retries, conditional GET and the host allowlist at once. Use it at your own risk: you take over responsibility for respecting each site's crawling rules, and you are far more likely to be rate-limited or blocked. There is no partial opt-out.
Development
python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest -q # 168 offline tests; fixtures in tests/fixtures/
.venv/bin/python scripts/politeness_smoke.py /tmp/js-ts-smoke-cache
.venv/bin/python scripts/e2e_mcp_test.pypytestis fully offline: HTTP is simulated withhttpx.MockTransportand time/jitter are injected, so the suite is deterministic and runs in seconds.tests/fixtures/holds real captured artifacts: an MDN page (245 KB), the MDN en-US sitemap (127 KB gzipped), MDN'srobots.txt(119 B), a TypeScript handbook page (196 KB), two npm registry JSON payloads and the 7,462-byterobots.txtpackument trap.scripts/politeness_smoke.py <cache-dir>is a live measurement, not a test: it drives the real tools against the real docs sites and prints what the politeness layer did (requests, robots, throttle, conditional GET, budgets).scripts/e2e_mcp_test.pyspawns the installedjs-ts-docsconsole script and speaks newline-delimited JSON-RPC to it (initialize→tools/list→ every tool → a negative case →health_check). Exit code 0 means every check passed. It resolves the server command asJS_TS_MCP_E2E_CMD(env override) → this checkout's.venv/bin/js-ts-docs→python -m js_ts_mcp.server.src/js_ts_mcp/politeness.pyandtests/test_politeness.pyare shared copies, kept byte-comparable across the four repos in this family; the repo-specific parts are the env-var names, the cache paths and the import path.
Troubleshooting
The first js_docs call is slower than the rest. That is the cold search-index build: MDN's robots.txt + the en-US sitemap + typescriptlang.org's robots.txt + the handbook page, all spaced by the per-host throttle, followed by the page you actually asked for. It happens once every 7 days; later calls hit the cached index. Delete cache.db and you pay for it again.
request budget exhausted for scope 'tool:js_docs'. One tool call hit its cap of 7 requests. js_search first — it is offline and free once the index is warm — and then pass the exact mdn: or ts: path instead of a bare name. js_status().politeness.budgets shows how much each scope spent.
A bare name costs five requests and an explicit mdn: slug costs two. A plain name has to resolve through the index, and on a cold cache that means building it (sitemap + handbook page + two robots fetches) before fetching your page. Once the index is warm the difference is one robots-free page fetch either way.
"partial": true, or fewer search results than expected. An index build ran out of its budget and stopped early, returning a partial index with a partial_reason. A partial index is not cached, so the next run rebuilds it. js_search still works — it just knows fewer pages — and js_status().checks.search_index.entries tells you how many it has (a complete index today is 14,825 entries: 14,708 MDN + 117 TypeScript).
"stale": true in search results. The rebuild failed (offline, DNS failure, 5xx) and the server fell back to the older cached index instead of failing the lookup.
Nothing is cached and overall is degraded. Read js_status().checks.cache.error — an unwritable JS_TS_MCP_CACHE_DIR (read-only mount, missing permission) is the usual cause. Tools keep working without a cache; they are just slower and noisier on the network.
A lookup returns blocked by robots.txt. The site's rules disallow that path for this user agent, and the request was not sent. On MDN that normally means /api/, /*/files/ or /media. Fetch the page yourself, or accept the consequences of the opt-out switch above.
npm_package returns package not found on npm. The registry answered 404 for that name. Check the name (npm names are case-sensitive for scoped packages, e.g. @types/node) and, if you pinned a version, that the version exists.
readme_markdown is null for a pinned version. Normal: the registry does not always carry a README for a non-latest release. Drop the pin to read the latest README.
A www.npmjs.com lookup is refused. By design: that host is not in the allowlist, so the layer returns host www.npmjs.com is not in ALLOWED_HOSTS and makes no request. Use npm_package — it fetches registry.npmjs.org and reports the www.npmjs.com URL to you.
License & attribution
MIT — see LICENSE. Copyright (c) 2026 Michal Gaspierik.
The design of the politeness layer was inspired by Crawl4AI (© Unclecode, Apache-2.0): a TTL-cached robots store, the wildcard rule translation, and a per-domain rate limiter with escalating backoff. The implementation is a clean-room rewrite in this project's own synchronous stdlib-plus-httpx style — no line was transcribed, translated or mechanically adapted from Crawl4AI, and Crawl4AI is not a dependency of this package. Three defects of the original design are fixed (robots fetched_at refresh, negative-result caching, Crawl-delay support), and the RFC 9309 matcher is what makes MDN's wildcard rules actually bind — robotparser would have let the dead /api/ calls through. The GPL-3.0 part of Crawl4AI — its vendored html2text fork — is deliberately excluded: no code, data or dependency from that tree is used or shipped here; HTML→markdown stays markdownify's job. The full statement, including the 37-character clean-room measurement and the reading pointers to the design, is in NOTICE.
Same clean-room architecture as flutter-mcp, java-spring-mcp and python-docs-mcp: fetchers → pure parsers → TTL cache → search index → tools, with no circular delegation, a pinned 1.x MCP SDK, error-dict failures and an offline test suite.
This server cannot be deployed
Maintenance
Related MCP Connectors
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Web data tools for AI agents: pages as markdown, search, maps, commerce, jobs, AI answers.
Web scraping for AI agents: scrape, search, crawl, map any website to markdown + JSON. No browser.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI agents to search local Markdown documents using natural language, with automatic indexing and section-level retrieval.108 npm1MIT
- AlicenseNot gradedqualityDmaintenanceGives AI agents full-text search over any Markdown/MDX documentation folder.14 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to search, fetch, and clean live documentation from LangChain, LlamaIndex, and OpenAI via Model Context Protocol, with token-optimized extraction.-
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to run local multi-engine web searches, fetch pages through anti-bot fallback chains, and query site-specific sources such as GitHub, npm, PyPI, Docker Hub, and Twitter/X for structured results.MIT