Skip to main content
Glama

Babel's Hoard

What does the version on disk actually say?

A local, version-exact API index of the packages installed in your projects, and a checker that catches hallucinated or misused APIs in code a model just wrote - without ever running that code.

Español · Quick start · Connect to Faustus · MCP reference · Portfolio

Search: "send a get request" returns httpx.Client.get with the signature of the installed httpx 0.28.1 Actual application, demo data: Babel's own interpreter with httpx, pydantic, fastapi, griffe and the json stdlib module indexed.

Why

A coding model's training data is a blur of library versions. It writes df.append(...) for pandas 2.x, passes a keyword that was renamed, imports a name that moved, or calls response.jsonify() because it sounds right. Web search is slow and not specific to the version you have. The truth is already on disk - the project's own .venv and node_modules. Babel reads them statically, answers with the real signature, and checks a snippet against them.

The checker is built to be trusted by a model that cannot double-check it: it only reports an error when it can prove it. A name missing from a module that star-imports a compiled extension (os, socket), fills its namespace through globals() (re, hashlib), or resolves names in __getattr__ is never reported as an error; code guarded by try/except ImportError or hasattr is left alone. What it cannot prove is counted as unchecked.

Check code: response.jsonify() and client.post(..., retry=3) flagged against httpx 0.28.1, with the offending lines Actual application, demo data. The snippet is parsed, never executed; jsonify is caught through the return type of client.get and the with ... as client binding.

Related MCP server: Carto MCP Server

Use cases

Each one was tried for real, in the browser and over MCP, on a FastAPI + React project with 228 installed packages (use cases, usability report):

  • Register a big project - paste its folder: the .venv and frontend/node_modules are detected in under a second, the installed packages are listed with exact versions, and runtime dependencies are indexed in the background with visible progress (dev tools skipped).

  • Look up an exact signature - type sqlalchemy.ext.asyncio.async_sessionmaker on Search and press Look up: the constructor parameters of the installed SQLAlchemy, even though its first index pass stopped at the size cap.

  • Check a snippet a model wrote - Python (httpx.AsyncClient(retries=3), a deprecated item.dict() on your own pydantic model) and TypeScript imports (MagicWand from lucide-react), each under its line.

  • "Faustus, write it and prove it exists" - the agent looks up httpx.AsyncClient.stream, checks its endpoint, fixes the two errors from the suggestions and gets a clean check in five calls.

  • Review a whole module - a correct 370-line service gives zero findings; three introduced mistakes come back as errors with their lines.

  • Search your own docs - index the project's docs/ folder from Docsets and find "how do I start the backend" next to the API index.

What is implemented

Area

Available now

Boundary

Python indexing

Static indexing (griffe over source and .pyi stubs, project code never runs) of any registered interpreter's distributions and stdlib modules, on first need and cached per environment + package + version. Breadth-first, so the public API is indexed first; every object expanded once with its other public paths linked to it; bases and re-exports from other packages resolved; completeness recorded per module/class (dynamic namespaces, caps, compiled modules, namespace sub-packages, stub names declared only with @overload); a namespace the cap left unexpanded is indexed the first time a lookup or check reaches it; every import name of a distribution indexed (pytest: py and pytest)

Caps per package for the first pass: 15,000 entries, 1,500 parsed modules (test suites skipped), 600 modules from other packages; libraries over a cap are marked partial and their incomplete namespaces never produce errors. Compiled extensions without stubs are recorded as such, with no members. The project's own (not installed) modules are not indexed

Version tracking

The interpreter probe is cached until its site-packages changes; after an upgrade the next lookup indexes the new version and marks the old index superseded (kept only as "other versions")

Detection relies on site-packages folder timestamps; editable installs that change code in place keep the indexed version until re-indexed

Code check (api_check_code)

Unknown modules/attributes, unexpected keywords, positional-only passed by keyword, too many positional, missing required, deprecated, with suggestions. Follows values through imports, assignments, annotations, return annotations, Self-returning methods, with/async with (including @contextmanager functions such as client.stream(...)), await and classes defined in the snippet (through their indexed bases); receivers (instance/class/static/unbound) and constructors handled; overloads checked against the ones that fit the call; PEP 702 @deprecated. Without env, Python code is checked against the newest project with an interpreter and TypeScript against the newest with node_modules

Errors only for statically complete namespaces; __getattr__/setattr(self, name) classes and lazy modules produce warnings; unknown types, custom metaclasses, __new__, unknown decorators and guarded code are unchecked. TypeScript: named imports/re-exports only (v1)

Lookup (api_lookup)

Signature, parameters (type, default, required, kind, description), return type, summary, first 1,500 characters of the docstring, members, source file:line, library version; found: false with real neighbouring names and whether absence is certain

Python and npm packages with type declarations

JS/TS indexing

TypeScript compiler API over a package's declarations (types/typings, exports[...].types, index.d.ts, @types/<name>): exports, one level of members, JSDoc, @deprecated; indexed on first use

Needs Node.js; types and docs for the first 500 exports of a package, names only for the rest (up to 50,000)

Search

SQLite FTS5 with bm25 weights (name, qualname, signature, summary, doc), identifier-aware tokens, filters, one hit per definition under its shortest path, re-ranked so callables and classes whose name matches come before attributes and constants, trigram "did you mean"; a dotted name offers an exact lookup that indexes the package on the spot

Lexical, no embeddings; only what is indexed

Offline docsets

DevDocs catalogue and install (HTML converted to Markdown, split by anchors) as a background job

Needs the network, only when the user or the model asks; converter is small, not a full HTML renderer

Markdown folders

.md/.mdx/.rst/.txt split by heading (fenced code aware, rst underlines)

Dependency/VCS/build folders skipped; 2,000 files, 2 MB per file

Assistant audit

Every /api/agent/* call (tool, argument summary, duration, result) in "Assistant activity"; the UI uses its own endpoints so only the model's calls appear

Local only

Ask the docs

Search screen: a question is answered from the top search hits by the shared llm model, with citations [id] that link back to the exact entry (and a Sources list; an id the model invented is shown struck through). When the shared embeddings model resolves, the top 50 lexical hits are re-ranked hybrid (reciprocal-rank fusion of the lexical and embedding orders, "semantic re-rank on" badge); otherwise lexical order

Answers only from what is already indexed; a UI feature, not an MCP tool; disabled with the reason shown when no model resolves

Shared model backend

Settings -> Models: which server/model is used for llm/embeddings right now and why, a Re-check button, manual overrides (Faustus URL/token, per-capability URL/model)

Read Shared models below

UI

Search, symbol panel, Libraries (installed packages with index status, filling in while the dependency job runs; default environment per language), Check code (line numbers, findings under each line), Docsets (with Index a folder), Assistant activity (which environment answered each call), Settings; light/dark, English/Spanish

Single user, local browser; backend messages (findings, notes) are in English

Quick start

git clone https://github.com/Luissalet/BabelsHoard.git
cd BabelsHoard

Windows

Double-click Iniciar Babel's Hoard.cmd. On first run scripts/start.ps1 finds Python 3.13 (py launcher, C:\Python313, then PATH; 3.11+ accepted), creates .venv, installs requirements-lock.txt, builds the web UI and the TypeScript probe if Node.js is installed, starts the app from the repository folder and opens http://127.0.0.1:8811 once /api/health answers. Later runs reinstall only when the lock file changed. Detener Babel's Hoard.cmd stops it.

Manual steps:

python -m venv .venv
.venv\Scripts\python -m pip install -r requirements-lock.txt
cd frontend; npm ci; npm run build; cd ..
cd babels_hoard\probes; npm ci; cd ..\..
.venv\Scripts\python -m babels_hoard

Linux / macOS

Python 3.11 or newer; Node.js 22 for the web UI and TypeScript indexing:

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements-lock.txt
(cd frontend && npm ci && npm run build)
(cd babels_hoard/probes && npm ci)
.venv/bin/python -m babels_hoard --demo --no-browser

Then open http://127.0.0.1:8811 (curl http://127.0.0.1:8811/api/health answers "service": "babels-hoard").

Options: --port <p> (default 8811), --data-dir <path> (default data/, or BABEL_DATA_DIR), --no-browser, and --demo, which uses data-demo/ and indexes a few packages of Babel's own interpreter so the UI can be tried without touching your projects.

Connect it to Faustus

Babel's Hoard is a plugin for Faustus, the local AI workspace, and declares itself with faustus-plugin.json. Start Babel's Hoard, then in Faustus: Connectors → Nearby apps → Add. Faustus launches the stdio adapter babels_hoard/mcp_server.py, which talks to the app over loopback.

MCP tool

What

Read-only

docs_libraries

Indexed libraries with versions, environments (and which is default), recent background jobs

yes

docs_search

Ranked search over APIs, docsets and markdown

yes

api_lookup

Exact installed signature, parameters and doc of one symbol

yes (writes only the local index cache)

api_check_code

Hallucinated/removed/misused APIs in a snippet

yes (writes only the local index cache)

docs_read

Read a docstring or docs section by id, in chunks

yes

docs_add_environment

Register a project folder, interpreter or node_modules

no

docs_catalog

Downloadable offline docsets (network)

yes

docs_install_docset

Download and index a docset as a background job (network)

no

docs_index_folder

Index a folder of markdown docs

no

Every tool description ends with English and Spanish keywords for tool retrieval. Argument and result shapes, limits and examples: docs/MCP.md. A skill for the agent is in skills/version-exact-apis/SKILL.md.

Any MCP client can use it over stdio:

{
  "mcpServers": {
    "babels-hoard": {
      "command": "C:\\path\\to\\Babel's Hoard\\.venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\Babel's Hoard\\babels_hoard\\mcp_server.py"],
      "env": { "BABEL_URL": "http://127.0.0.1:8811" }
    }
  }
}

Babel's Hoard never loads its own model. For "Ask the docs" it uses two capabilities from HoardLink (vendored in babels_hoard/hoard_link/), the shared model backend every Faustus plugin uses: llm to answer, embeddings to semantically re-rank search hits first (hybrid search) when one is available. Resolution order in one line: an explicit override in Settings or backend.json, then a running Faustus instance's own model registry, then a shared server already listening on loopback (llama.cpp, Ollama, any OpenAI-compatible server) - the app works fully without any model connected; the "Ask the docs" button is simply disabled with an honest reason ("No language model is connected...", plus the probe details as a tooltip) and every other screen is unaffected. A broken, hand-edited backend.json never stops the app: it falls back to auto-detection and Settings -> Models says why the file was ignored.

Architecture

FastAPI + SQLite (WAL, one connection per thread, FTS5), a background job worker, griffe for static Python analysis, the TypeScript compiler API for declarations, a React 19 + Vite UI and a standalone stdio MCP adapter.

flowchart LR
  UI["React UI"] -->|"/api/*"| API["FastAPI app<br/>127.0.0.1:8811"]
  MCP["MCP stdio adapter"] -->|"/api/agent/*"| API
  API --> DB[("SQLite + FTS5 index")]
  API --> JOBS["background jobs"]
  API --> IDX["griffe (Python)<br/>TypeScript probe (JS/TS)"]
  JOBS --> IDX
  IDX -. "reads, never imports" .-> ENV[".venv / node_modules"]
  API --> LINK["HoardLink"] -. "optional" .-> MODELS["shared llm / embeddings"]

Modules, data model and the decisions behind the checker: docs/ARCHITECTURE.md.

Privacy and security

Binds 127.0.0.1 only; requests with a foreign Host, cross-site writes and cross-site API reads are refused. No telemetry. Only the docset catalogue and docset installs use the network (the public DevDocs mirror; docset content keeps its original licences). Babel runs the interpreters you register, only to execute its stdlib-only probe script, and only files named like a Python interpreter; it reads installed source and declaration files and never imports or executes project code or snippets. Every call the assistant makes is listed under Assistant activity (tool, argument summary, duration, result, and which environment answered).

Development

.venv\Scripts\python -m pip install pytest pytest-asyncio
.venv\Scripts\python -m pytest -q
cd frontend; npm run build

On Linux/macOS the same with .venv/bin/python. 159 tests, about a minute on a shared 2-CPU Linux machine, with no network, no GPU and no model downloads. They cover: the browser guard and static-file confinement (path traversal attempts), error shapes, per-thread connections and a health check that answers while a long tool runs; indexing of real installed packages (httpx, pydantic, fastapi, PyJWT, python-dotenv, the stdlib) and of two fixture packages - fakelib 1.x/2.x (an API removed between versions) and edgelib (every dynamic-namespace, decorator, metaclass, overload and context-manager case the checker must get right); a simulated upgrade inside a throwaway venv; the stdlib names that used to be false positives (os.getcwd, socket.AF_INET, re.IGNORECASE, hashlib.sha256, sqlite3.connect); search, docsets, markdown and JS/TS indexing (needs Node.js); the MCP adapter spawned over the real stdio protocol against a live app, including tool keywords, annotations, id round-trips and error pass-through; and the shared model backend (/api/backend*, config persistence that never leaks the Faustus token, cleared fields, invalid values rejected before they are saved, a broken backend.json at startup, and "Ask the docs" - disabled path, empty query, no matching entries, citation filtering, hybrid re-rank, and a failed model call - all against a fake Link, offline); and one regression test per fixed finding of the usability report (tests/test_usability_fixes.py: per-language default environments, multi-package distributions, compiled submodules and namespace packages, overload-only stubs, on-demand expansion past the cap, context-manager values, snippet classes, fitting overloads, search ranking, large JS export lists, and that no parsed package stays in memory after indexing).

False-positive harness: scripts/check_corpus.py runs the checker over the source of installed packages, which works, so any error it reports is suspect. On 300 files sampled from starlette, fastapi, httpx, uvicorn, mcp, anyio, pydantic-settings, click, jsonschema, griffe, pandas and requests it reports no errors and no warnings (the pandas/requests sample: 4,954 verified checks, 11,551 left unchecked). On 110 files from SQLAlchemy, pydantic-settings, sse-starlette, tenacity, fastembed, psutil, chromadb, aiosqlite, alembic, PyJWT, pytest, attrs and numpy (2,880 verified checks) it reports 3 errors, all genuine: chromadb's distributed-mode code imports *_pb2 modules its wheel does not ship. Seven false-positive classes it found are fixed and covered by tests.

npm run build in frontend/ passes with zero TypeScript errors. The first-run path of scripts/start.ps1 (venv, lock install, UI build, start, health wait, second-run detection) was run under PowerShell 7 on Linux. CI runs the tests on Ubuntu and Windows with Python 3.11, 3.12 and 3.13, builds the UI, and on windows-latest starts the app with start.ps1 and stops it with stop.ps1.

Roadmap / known limits

  • Modules that build their names at import time from platform code (psutil) stay unverifiable rather than risk false errors, so a made-up psutil.gpu_percent() is not caught.

  • pydantic v1 configuration idioms (class Config: orm_mode = True) are out of reach of a static name check.

  • An on-demand lookup waits for the library the background dependency job is indexing at that moment before it runs.

  • TypeScript checks cover named imports and re-exports only.

  • Finding messages, job progress and library notes come from the backend in English, even in the Spanish interface.

  • Not tried yet: the Windows launchers on a real Windows machine, "Ask the docs" with a real model, docset downloads from the UI and dark-mode screenshots.

The full list of findings and how each was found is in docs/USABILITY_REPORT.md.

License

MIT - see LICENSE.

Available Tools

9 tools
api_check_codeA
Read-onlyIdempotent

Check code for hallucinated or misused APIs of installed versions (check code, comprobar código, validar)

Statically checks a snippet against the libraries installed in the environment: unknown modules/attributes (hallucinated or removed APIs), unexpected keyword arguments, too many positional arguments, missing required arguments, deprecated calls. Never runs the code. Run it on code you wrote before showing it; fix errors, re-check.

Args: code: the source (a whole file or a snippet; imports must be included). env: environment id, project folder path, or omitted for the default environment. language: "python" (full checks) or "typescript" (v1: named imports only).

Returns: {ok, findings: [{line, col, severity, code, symbol, message, suggestion}], checked, unchecked, libraries}. ok is false only for errors. severity "warning" = not a declared member but the class/module creates names dynamically, or deprecated. Babel stays silent on anything it cannot prove (dynamic code, unknown types, try/except ImportError, hasattr guards); those count in 'unchecked'. Keywords: check code, validate code, hallucinated api, does this exist, wrong arguments, verify snippet, comprobar codigo, revisar codigo, validar, existe, alucinacion

ParametersJSON Schema
NameRequiredDescriptionDefault
envNo
codeYes
languageNopython

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantial non-obvious behavior: "Never runs the code," the silent-failure policy ("Babel stays silent on anything it cannot prove" — dynamic code, hasattr guards, try/except ImportError), and the fact that those go into 'unchecked' rather than findings. That is exactly the kind of context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core description is front-loaded and information-dense, with the key constraint (never runs code) early. The trailing keyword block is bulky and partially redundant with the opening line, which costs a point, though it aids retrieval.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description reproduces the full return shape ({ok, findings:[{line, col, severity, code, symbol, message, suggestion}], checked, unchecked, libraries}) and defines the semantics of 'ok' and 'warning' severity. An agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it does: 'code' requires imports to be included, 'env' accepts an environment id, a project folder path, or omission for the default, and 'language' enumerates "python" (full checks) vs "typescript" (v1: named imports only). Each parameter's accepted values and limitations are explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) plus the precise resource (code against installed library APIs) and enumerates the failure classes it detects: unknown modules/attributes, unexpected kwargs, arg-count errors, deprecated calls. This clearly separates it from the docs_* and api_lookup siblings, which retrieve information rather than validate a snippet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit temporal guidance: "Run it on code you wrote before showing it; fix errors, re-check." That is a clear when-to-use rule. It does not name a sibling alternative or state when-not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

api_lookupA
Read-onlyIdempotent

Exact signature of an installed API (signature, firma, parámetros, documentación, versión instalada)

Parameters (with defaults and which are required), return type, docstring and source location of one dotted symbol, from the version installed in the environment - e.g. "pandas.DataFrame.merge", "httpx.Client", "json.dumps", or "react.useState" for a node_modules package. Use it before calling any API you are not sure about. Indexes the package on first use (local, no network; can take seconds).

Args: symbol: fully dotted path starting with the import name. env: environment id, project folder path, or omitted for the default environment. library: optional library name to disambiguate.

Returns: found=true -> {id, qualname, kind, signature, params: [{name, kind, annotation, default, required, description}], returns, summary, doc (first 1500 chars), doc_truncated, members (for modules/classes), library, source}. Use docs_read(id) for the rest of a long doc. found=false -> {certain, suggestions: [real names], message}; certain=false means the name may still exist at runtime. Keywords: signature, parameters, arguments, api lookup, what does it take, return type, firma, parametros, argumentos, que recibe, que devuelve, existe esta funcion

ParametersJSON Schema
NameRequiredDescriptionDefault
envNo
symbolYes
libraryNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, closed-world), and the description adds genuinely new behavioral context: first-use indexing is local with no network but 'can take seconds', plus the failure semantics of found=false and the meaning of certain=false. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the Args/Returns/Keywords layout is scannable, with the keyword line serving search recall. Slightly long and the first parenthetical overlaps with the Returns section, but every block earns its place given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description supplies the return shape itself, including the params sub-object and the doc_truncated flag with a pointer to docs_read. Together with the documented parameters, the first-use indexing caveat, and the ambiguity caveat, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does so for all three parameters: symbol as 'fully dotted path starting with the import name', env as 'environment id, project folder path, or omitted for the default environment', and library as 'optional library name to disambiguate'. Required vs optional is also implied by the defaults discussion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Exact signature of an installed API') and immediately scopes it with parenthetical detail (signature, parameters, docs, installed version). Concrete example symbols ('pandas.DataFrame.merge', 'json.dumps', 'react.useState') make it trivially distinguishable from siblings like docs_search or docs_read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use it before calling any API you are not sure about') and routes follow-up work to a sibling ('Use docs_read(id) for the rest of a long doc'). It also notes the indexing precondition on first use. No explicit when-not-this-tool statement, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_add_environmentA
Idempotent

Register a project folder or venv so its packages can be checked (register project, registrar proyecto)

Accepts a project folder (its .venv/venv/env and node_modules, also in frontend/, web/, client/ or ui/, are detected), a python/python.exe interpreter, or a node_modules folder. The most recently registered project becomes the default env (per language: Python checks never default to a node_modules-only env). Registering again is harmless and makes it the default again.

Args: path: absolute path to the project folder, interpreter or node_modules. index_dependencies: also index the project's direct dependencies (from pyproject.toml / requirements*.txt) in a background job.

Returns: {environment: {id, label, python, ...}, dependency_job_id, dependencies, message}. Keywords: register project, add environment, use this project, index project, registrar proyecto, anadir entorno, indexar proyecto, usar este proyecto

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
index_dependenciesNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-readOnly, idempotent, non-destructive, closed-world. The description adds real context beyond them: the per-language default-env rule, that re-registering is harmless and re-promotes the default, and that dependency indexing runs as a background job. It stops short of stating auth requirements or what gets mutated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and args are cleanly separated, but the bilingual duplication ('register project, registrar proyecto') and the trailing Keywords list repeat the same information in two languages, adding length without adding selection value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description documents the return shape ({environment, dependency_job_id, dependencies, message}), the accepted inputs, and the default-env consequence. Only the relationship to sibling tools remains unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage the description carries the burden, and it does: 'path' is explained as an absolute path to a project folder, interpreter, or node_modules (with auto-detected venv/node_modules locations), and 'index_dependencies' is explained as indexing pyproject.toml/requirements*.txt in a background job. Both parameters get meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (register) and resource (project folder/venv/interpreter) plus the goal (so its packages can be checked). However, it never distinguishes itself from the sibling docs_index_folder, which sounds similarly scoped, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies when to use it ('so its packages can be checked') and notes re-registration is harmless and re-sets the default. But it names no alternative and gives no explicit when-not guidance versus docs_index_folder or docs_libraries, leaving the sibling choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_catalogA
Read-onlyIdempotent

List downloadable offline docsets; uses internet (docset catalogue, catálogo, documentación offline)

Language references and frameworks: python, javascript, react, css, rust... Reaches the internet (the public DevDocs catalogue). Use before docs_install_docset to get the exact slug.

Args: query: filter by name or slug, e.g. "python" or "react". limit: default 10, max 50.

Returns: {results: [{slug, name, version, db_size_kb}], count, total, truncated}. Keywords: docset catalogue, offline docs, available documentation, download docs, catalogo de documentacion, documentacion sin conexion, que documentacion hay

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/openWorld/idempotent/non-destructive, so the bar is lower. The description still adds value by disclosing the network dependency ('Reaches the internet (the public DevDocs catalogue)') and the return payload shape, which the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and well-sectioned with Args/Returns, but the trailing multilingual keyword block and the partial sentence fragment 'Reaches the internet (the public DevDocs catalogue)' before the Args section add noise that does not earn its place for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description spells out the return shape ({results: [{slug, name, version, db_size_kb}], count, total, truncated}), so an agent knows exactly what it gets back before calling. Nothing material is missing for a two-param read-only listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does: query is 'filter by name or slug' with concrete examples, and limit documents both the default (10) and the max (50), neither of which appears in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List downloadable offline docsets') and immediately names the sibling relationship ('Use before docs_install_docset'), letting an agent distinguish it from docs_search, docs_libraries, and docs_read without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'Use before docs_install_docset to get the exact slug.' This names the downstream tool and the precondition, which is exactly the routing guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_index_folderA
Idempotent

Index a folder of Markdown docs for search (index folder, indexar carpeta, documentación del proyecto)

Indexes .md/.mdx/.rst/.txt files (a project's docs/ folder, a cloned docs repository) by heading, so docs_search and docs_read can use it. Dependency, VCS and build folders are skipped; at most 2000 files.

Args: path: absolute path to the folder. name: library name to show and filter by; defaults to the folder name.

Returns: {id, name, status, entry_count, note, hint}. Keywords: index folder, index markdown, project docs, indexar carpeta, indexar documentacion, documentacion del proyecto

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pathYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the write/idempotent/non-destructive profile, and the description adds real behavior beyond them: accepted extensions (.md/.mdx/.rst/.txt), automatic skipping of dependency/VCS/build folders, and a hard cap of 2000 files. What it omits is what happens on re-index of an already-indexed path or how errors surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core constraints (file types, skipped folders, 2000-file cap) are front-loaded in the first paragraph after the summary line. It is somewhat padded by duplicated multilingual keyword lines at the top and bottom, which repeat the same intent rather than adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter indexing tool with no output schema, the description supplies inputs, indexing scope, limits, and the returned object shape ({id, name, status, entry_count, note, hint}). Nothing essential for a correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden and does: it states that path must be an absolute path to the folder and that name is a display/filter label defaulting to the folder name. Both parameters gain meaning the bare schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (index a folder of Markdown docs) and explicitly ties it to the downstream consumers docs_search and docs_read, which separates it from those siblings. An agent can tell it builds a searchable index rather than performing a search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'so docs_search and docs_read can use it' gives clear context for when this tool is needed, plus example sources (a project's docs/ folder, a cloned docs repo). No explicit when-not or exclusion guidance is offered, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_install_docsetA
Idempotent

Download and index one offline docset in background (install docs, instalar documentación, descargar)

Takes a slug from docs_catalog (e.g. "python~3.13"). Reaches the internet. Runs as a background job and returns at once; follow it with docs_libraries (jobs) and search it with docs_search when done.

Args: slug: the exact slug from docs_catalog.

Returns: {job_id, status, slug, message}. Keywords: install docset, download documentation, offline docs, instalar documentacion, descargar documentacion

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new context the annotations lack: it reaches the internet, runs asynchronously in the background, and returns immediately rather than blocking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is front-loaded with the scope and the sync/async behavior, and the Args/Returns blocks are compact. The bilingual keyword lines restate content already in the first sentence, which is mild padding but may aid retrieval.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape {job_id, status, slug, message} and the async job-tracking workflow. An agent has everything needed to invoke it and know what to do next.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'slug' parameter, so the description must compensate. It does: the slug is defined as the exact identifier from docs_catalog with a concrete example ('python~3.13'), which is enough to call the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Download and index one offline docset in background.' The word 'offline' plus 'background' distinguishes it from siblings like docs_index_folder and docs_catalog at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow: slug comes from docs_catalog, follow progress with docs_libraries (jobs), then search with docs_search. It names the prerequisite source and the follow-up tools, though it does not explicitly exclude docs_index_folder as an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_librariesA
Read-onlyIdempotent

List indexed libraries, exact versions and project environments (libraries, librerías, versión, entornos)

What Babel has indexed (exact versions), the registered environments (project interpreters / node_modules folders; default_for says which one answers Python or TypeScript when env is omitted) and recent background jobs. Use it to find an env id, or to follow a docset install / dependency indexing job. Installed packages that are not listed yet are indexed automatically the first time api_lookup or api_check_code needs them.

Args: ecosystem: "python", "js", "docset" or "markdown". env: environment id or project path; limits libraries to that environment. limit: libraries per page (default 15, max 200). offset: for the next page.

Returns: {libraries: [{id, ecosystem, name, version, env, status, entries}], total, has_more, next_offset, environments: [{id, label, python, default, ...}], jobs: [{id, kind, status, progress, message}]}. Keywords: list libraries, what is installed, which version, environments, job progress, listar librerias, que hay instalado, que version tengo, entornos, progreso

ParametersJSON Schema
NameRequiredDescriptionDefault
envNo
limitNo
offsetNo
ecosystemNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds real behavioral context: auto-indexing of unlisted packages, the meaning of default_for when env is omitted, and pagination defaults. It does not cover auth requirements or rate limits, which is a minor gap for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then scope, usage, args and returns in a clean structure. The bilingual keyword line is retrievability padding but is short and clearly separated, and the Returns block earns its space since no output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% param documentation in the schema, the description supplies both the argument semantics and a full return shape (libraries, environments, jobs with their fields). An agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it enumerates valid ecosystem values, explains env as an id or project path that limits results, and gives limit's default (15) and max (200) plus offset's paging role. Those are exactly the details absent from the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List indexed libraries, exact versions and project environments') and enumerates the returned entity classes (libraries, environments, jobs), so an agent knows exactly what this tool surfaces. It does not explicitly contrast itself with docs_catalog, but the scope is unambiguous enough on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases: 'Use it to find an env id, or to follow a docset install / dependency indexing job.' It also clarifies the fallback path — unlisted packages are auto-indexed when api_lookup or api_check_code need them — which steers the agent away from unnecessary calls. No explicit 'do not use for X' exclusion is given, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_readA
Read-onlyIdempotent

Read the full text of one doc entry by id, in chunks (read docs, leer documentación, seguir leyendo)

One entry is a docset page section, a markdown section or a long docstring. Ids come from docs_search results and from api_lookup (field id).

Args: id: the entry id, passed back exactly. offset: character offset to continue from (use next_offset). max_chars: chunk size, default 4000, max 20000.

Returns: {found, qualname, library, text, offset, total_chars, has_more, next_offset}. Keywords: read doc, full documentation, read more, continue, leer documentacion, seguir leyendo, documentacion completa

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
offsetNo
max_charsNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds genuinely useful behavioral context beyond that: the chunked-read model, the default and maximum chunk sizes, and how continuation works via next_offset. It does not mention any rate limits or failure modes for bad ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then structures Args and Returns explicitly, which is easy to parse. The multilingual keyword tail ('leer documentacion, seguir leyendo...') is redundant for a model already reading the English description, but it is conventional for retrieval-style tools and costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-output-schema tool, the description supplies everything needed: purpose, id provenance, pagination mechanics, size limits, and even the exact return object keys ({found, qualname, library, text, offset, total_chars, has_more, next_offset}). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does so: id is 'passed back exactly', offset is a character offset to continue from using next_offset, and max_chars has a stated default of 4000 and a hard max of 20000. Every parameter is documented with semantics the bare schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the full text of one doc entry by id') plus the delivery mode ('in chunks'), which cleanly separates it from docs_search and docs_catalog. It also defines what an 'entry' is (docset page section, markdown section, long docstring), removing ambiguity about the unit being read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete provenance for the required id ('Ids come from docs_search results and from api_lookup (field id)') and tells the agent how to continue reading via next_offset. It stops short of an explicit when-not rule (e.g. 'use docs_search first to find ids'), but the routing context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observedapi_check_code
    • First observedapi_lookup
    • First observeddocs_add_environment
    • First observeddocs_catalog
    • First observeddocs_index_folder
    • First observeddocs_install_docset
    • First observeddocs_libraries
    • First observeddocs_read
    • First observeddocs_search

TDQS

A4.3/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct role in the doc/API lifecycle: catalog vs. indexed libraries, fuzzy search vs. exact lookup, snippet check vs. full read, and environment registration vs. docset install/index. Descriptions explicitly chain tools (e.g., search → lookup → read), leaving little chance of misselection.

Naming Consistency4/5

All names use snake_case and a domain prefix (docs_ for doc/environment management, api_ for API inspection), but the prefix pattern mixes noun phrases (docs_libraries, docs_catalog) with verb phrases (docs_add_environment, docs_install_docset). This is readable and mostly predictable, yet not fully uniform.

Tool Count5/5

9 tools cover the server's scope without obvious redundancy or missing core operations; each tool earns its place in the lifecycle.

Completeness4/5

Core workflows—discovering, installing, indexing, searching, reading, and validating against docs/APIs—are covered. Minor gaps exist for cleanup/removal of docsets or environments and for updating stale indexes, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A graph-powered code intelligence engine that indexes codebases into a structural knowledge graph to provide AI agents with deep context on function calls, types, and execution flows. It offers local, zero-dependency tools for hybrid search, impact analysis, and dead code detection across Python, JavaScript, and TypeScript projects.
    956 PyPI
    814
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides real-time access to Python package documentation, source code, and symbol search to prevent AI hallucinations.
    110 PyPI
    6
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    Provides codebase indexing and retrieval tools that give AI agents token-efficient, query-relevant context packages (symbols, imports, and dependencies) instead of scanning entire repositories.
    6
    -