Skip to main content
Glama

mathlib-mcp

Make GPU math libraries LLM-friendly, and measure whether it helps.

LLMs write plausible but wrong code for GPU math libraries. They call functions that do not exist, use keyword arguments from an older SciPy, and miss behaviours that differ from NumPy, such as CuPy returning NaN for a singular matrix instead of raising. This project turns the APIs of CuPy 14.2 and nvmath-python 1.0 into an LLM-first representation, serves it to coding agents over MCP, and includes a GPU benchmark to measure the effect on generated code. Those two libraries are the Python front-ends to cuBLAS, cuFFT, cuSOLVER, cuSPARSE, cuDSS and cuTENSOR.

Part

What it is

API cards

2,741 public symbols and 275 API cards extracted from the library sources (no GPU or import needed). 43 cards are curated with the CUDA-X backend, pitfalls checked against the source and on a GPU, and an example verified on a GPU.

MCP server

search_api, get_api_card, get_example, check_snippet, plus an llms.txt resource.

check_snippet

Static checker that finds hallucinated APIs, wrong or positional-only keywords, too many positional arguments and missing arguments, and suggests the real name.

Benchmark

26 GPU tasks across six CUDA-X libraries with NumPy/SciPy references; 17 have a "naive" solution that falls into a documented trap.

Results

Two Claude models wrote solutions to the 26 benchmark tasks, 3 samples per task, under three conditions: no help; the full API reference (llms-full.txt) in the system prompt; or the mathlib-mcp tools. Every solution was run on a Tesla T4 and checked against NumPy/SciPy references (full report, all attempts and transcripts).

No help

Docs in prompt

MCP tools

Claude Haiku 4.5: tasks solved

51%

87%

92%

final code with a hallucinated API or wrong arguments¹

36%

1%

4%

input tokens per task

0.4k

101.6k

15.3k

cost per task

$0.003

$0.014

$0.018

Claude Sonnet 5.5: tasks solved

94%

99%

99%

final code with a hallucinated API or wrong arguments¹

3%

1%

1%

input tokens per task

0.5k

134.4k

3.5k

cost per task

$0.006

$0.035

$0.008

¹ Scored with the fixed check_snippet (see the bug note below). The run itself used the version before the fix, which accepted the three cupy.cublas attempts.

What the numbers say:

  • Haiku 4.5 goes from coin-flip to reliable. Tasks solved rise by +41 points (95% CI +26 to +58), and hallucinated APIs fall from 36% of attempts to 4%. The remaining 4% are the three cupy.cublas attempts described below. The biggest gain is on nvmath-python (6% → 83% solved), a library newer than the model.

  • On-demand knowledge beats putting it all in the prompt. Both conditions contain the same curated knowledge. MCP solved slightly more for Haiku (+5 points, CI +1 to +10). Sonnet matched the docs prompt with 38× fewer input tokens, at a quarter of the cost.

  • Sonnet 5.5 already knows most of these APIs (94% without help), so it calls the tools only when unsure (0.6 calls per task). Its one remaining MCP failure is a JAX keyword (cupy.repeat(..., total_repeat_length=)) in code it did not run through check_snippet.

  • The evaluation found a bug in the tool itself. In all three failed Haiku + MCP attempts at batched_solve_vectors, Haiku called cupy.cublas after only import cupy. That submodule is not loaded by import cupy, but check_snippet accepted the code. The cause: the symbol index treated every submodule as an attribute of its package. Fixed after v1. The extractor now works out which modules import X actually loads, by following module-level imports through the source. Cards now say when an extra import is needed, and two new GPU-checked facts confirm the model (import cupy does not load cupy.cublas; import cupyx.scipy.sparse does not load .linalg). Re-checked against all 468 v1 attempts, the fixed checker raises no false alarms on the 407 attempts that passed on the GPU. It flags 36 of the 41 that failed with an API-type error. The five it misses are beyond a static check: two implicit NumPy conversions, a key inside an options dict, a cuDSS memory-layout requirement, and a call into an internal Cython module. The solved rates above measure the tool before the fix; re-running the MCP condition would measure it, optimistically, since the task set is the same.

  • Knowledge alone is not enough when the semantics are counter-intuitive. Haiku read the pitfall that cuBLASLt's bias epilog is per row, and still added the bias per column (0/3 in every condition). Sonnet got it right every time. Since v1, the matmul card spells out the transposed recipe for a per-column bias.

Re-run after the fix (v1.1, MCP condition)

After fixing the checker, the MCP condition was re-run with both models: 156 attempts, $2.12 (report).

Tasks solved with the MCP tools

v1

v1.1

Claude Haiku 4.5

92%

96%

Claude Sonnet 5.5

99%

99%

  • Both targeted tasks were fixed. Haiku went from 0/3 to 3/3 on batched_solve_vectors (it now imports cupy.cublas) and on fused_linear_relu (it now uses the transposed bias recipe).

  • The overall change is within noise. Two other tasks dropped, so Haiku's change against v1 MCP is +4 points (CI −6 to +17). Against no help it is now +45 points (CI +27 to +63).

  • The re-run exposed a wrong claim in the curated knowledge. The errstate card said failed factorizations raise under linalg='raise'. Haiku followed that advice with cupyx.scipy.linalg.lu_factor, which never raises: on a singular matrix it only issues a RuntimeWarning. Two of three solve_or_raise attempts failed that way. The cards now say so, and a new GPU-checked fact confirms the behaviour.

  • These numbers are optimistic. The fixes came from failures on the same tasks. The docs condition was not re-run, although llms-full.txt received the same fixes.

Caveats:

  • Same author. The tasks and the curated pitfalls were written by the same person. The docs and MCP conditions therefore measure how well known pitfalls are delivered, not whether unknown ones are discovered. An independently written held-out task set is the next step.

  • Wide intervals. With 26 tasks × 3 samples, the 95% intervals are ±10–15 points. Sonnet is near the ceiling, so its differences are within noise.

  • Pipeline controls. Every reference solution passed (26/26), every naive trap solution failed (17/17), and every null solution failed (26/26). There were 0 infrastructure errors. Generation cost $6.58 in total.

Related MCP server: cudaq-docs-mcp

What we learned

  1. Curated, verified knowledge does most of the work; MCP is one way to deliver it. For Haiku, putting the whole reference in the prompt already gave most of the gain (51% → 87%). The MCP tools added a little more (92–96%). The value is in the content: exact signatures, pitfalls and recipes checked on a GPU.

  2. The benefit depends on the model. Haiku 4.5 gained +41 to +45 points, and its hallucinated APIs almost disappeared. Sonnet 5.5 was already at 94% and gained 5 points. The help matters most for smaller models and for APIs newer than the model: on nvmath-python, Haiku went from 6% to 83%.

  3. On-demand delivery scales; a stuffed context does not. The same knowledge took 3.5k–15k input tokens through the tools, against 101k–134k in the prompt (38× fewer for Sonnet, at a quarter of the cost). Two libraries already take half of Haiku's 200k-token context window, so a whole CUDA-X stack would not fit. The cost is extra turns and latency (11.5 s versus 4.5 s per task for Haiku).

  4. Verification is the most distinctive piece, but the model has to use it. check_snippet catches 36 of 41 runtime API failures with no false alarms. Haiku ran it on 99% of attempts; Sonnet on 12%. Sonnet's one remaining failure was in code it never checked.

  5. Knowing is not applying. A warning ("the bias is per row") did not help Haiku. A concrete recipe did (0/3 → 3/3). Pitfalls work best when they show the correct code.

  6. The knowledge layer needs its own tests. Both regressions came from the tool, not the model: the checker accepted a missing submodule import, and a card over-generalised errstate. The evaluation found both. Facts checked on a GPU, plus re-running the evaluation after each change, catch errors that reading the source does not.

Future work

  • Held-out tasks. Tasks written independently of the curated pitfalls, so the results measure generalisation rather than delivery of known pitfalls.

  • Harder tasks. Sonnet 5.5 is near the ceiling on this set, so harder tasks are needed to measure stronger models.

  • An Agent Skill as another delivery mechanism. A Claude Agent Skill (SKILL.md with the API cards and the checker script bundled) would load a procedure only when relevant: look up the card, write the code, run the checker. That could close the gap in point 4 (strong models skipping verification). It would only work in Claude products, whereas MCP works in any client. The two also combine: a Skill can tell Claude when to use the MCP tools. Measure it as a fourth condition in the existing harness.

  • Evaluation inside real coding tools. Run the same tasks through Claude Code (headless) and other MCP clients, not only through the API.

  • Re-run the docs condition so the v1.1 tool is compared like for like (llms-full.txt received the same fixes).

  • Checker coverage. Check keys inside options={...} dicts against the options dataclass, the one statically detectable miss among the five.

Example

$ check_snippet
import cupy as cp
import cupyx.scipy.sparse.linalg as spla
from cupy.linalg import lu_factor
import nvmath
x, info = spla.cg(A, b, tol=1e-8)
w = spla.eigsh(A, 3, 'SA')
e = nvmath.linalg.advanced.MatmulEpilog.RELU_BIAS_GELU

FAIL: 4 problem(s) found
- line 3 [error: unknown-symbol] `cupy.linalg.lu_factor` does not exist (no `lu_factor` in module `cupy.linalg`). Did you mean: `cupyx.scipy.linalg.lu_factor`?
- line 5 [error: unknown-keyword] `cupyx.scipy.sparse.linalg.cg` has no parameter `tol`. Signature: `cg(A, b, x0=None, *, rtol=1e-05, atol=0.0, maxiter=None, M=None, callback=None)`. Did you mean: `rtol`, `atol`?
- line 6 [error: too-many-positional] `cupyx.scipy.sparse.linalg.eigsh` takes at most 2 positional argument(s) but 3 were given. Signature: `eigsh(a, k=6, *, which='LM', v0=None, ncv=None, maxiter=None, tol=0, return_eigenvectors=True)`.
- line 7 [error: unknown-attribute] `nvmath.linalg.advanced.MatmulEpilog` has no attribute `RELU_BIAS_GELU`. Did you mean: `nvmath.linalg.advanced.MatmulEpilog.RELU_BIAS`, `nvmath.linalg.advanced.MatmulEpilog.GELU_BIAS`, `nvmath.linalg.advanced.MatmulEpilog.RELU_AUX_BIAS`?
Checked 6 API reference(s): ...

A curated card, as an agent sees it (abridged):

### `cupyx.scipy.sparse.linalg.cg`
cg(A, b, x0=None, *, rtol=1e-05, atol=0.0, maxiter=None, M=None, callback=None)
cupy 14.2.0 · function · backend: cuSPARSE (SpMV), cuBLAS · SciPy equivalent: scipy.sparse.linalg.cg

Pitfalls
- The tolerance keyword is `rtol` (keyword-only, matching SciPy >= 1.12). The old `tol=` raises `TypeError`.
- Returns `(x, info)`. `info == 0` means converged and `info > 0` is the iteration count reached without converging. Always check `info`.
...

Use it with an agent

# Claude Code
claude mcp add mathlib -- uvx --from git+https://github.com/jarski/mathlib-mcp mathlib-mcp

Any MCP client works the same way: the server speaks stdio and has no runtime dependencies besides mcp. Semantic search is optional: pip install mathlib-mcp[embeddings] fuses a small embedding model into the BM25 ranking.

How it works

library sources ──► static extractor ──► symbols.json  (existence + signatures, for check_snippet)
 (sdist / wheel,     ast + .pyi + .pyx     cards.json    (API cards)
  never imported)         │                llms.txt / llms-full.txt
                          ▼
              curation/cards/*.toml  (backend, pitfalls, examples; build fails if
                                      a curated API disappears from the source)
                          │
                          ▼
                 MCP server ◄──── coding agent
  • Static extraction. CuPy and nvmath-python cannot be imported without CUDA, so the extractor reads their source with ast. It follows re-exports, __all__ star imports, aliases (MatmulEpilog = cublasLt.Epilogue), lazy module __getattr__, .pyi stubs, Cython .pyx files, docstring templates ({a} placeholders filled by nvmath's docstring_decorator) and functools.partial wrappers. The output is reproducible from pinned versions: scripts/fetch_sources.py, then mathlib-mcp-build.

  • Curation is checked. Every pitfall was checked against the 14.2.0 / 1.0.0 sources (for example, cupy.linalg.solve rejects a stacked (B, M) right-hand side, and cuBLASLt's bias epilog is per row). Behavioural claims are also executable facts that run on a GPU. A card's example is labelled "verified on <GPU, CUDA, versions>" only after it passes on a GPU in its current form.

  • Honest uncertainty. Namespaces the index cannot list completely (Cython modules, excluded packages) give unverified notes instead of false errors. A test makes sure no curated example is ever flagged.

Benchmark

Family

Tasks

Examples of traps

cuBLAS / cuBLASLt

6

in-place gemm(out=, beta=); per-row bias epilog; 24 GB broadcasting

cuFFT

6

irfft length for odd sizes; nvmath.fft.fft rejects real input and transforms all axes; nvmath inverse FFTs are unnormalized

cuSOLVER

8

batched solve needs (B, M, 1); silent NaN unless cupyx.errstate(linalg='raise'); svd full matrices on a 60000×50 input

cuSPARSE

4

cg(rtol=) not tol=; eigsh has no 'SM' and keyword-only which

cuDSS

1

column-major right-hand sides

cuTENSOR

1

nvmath.tensor contraction API

Each task has a prompt, a NumPy/SciPy reference, a correctness check and a canonical GPU solution. Candidates run in a separate subprocess and are classified as pass, wrong_answer, exception, not_on_gpu, expected_error_not_raised, timeout or crash. Static scoring adds the API-error count from check_snippet and checks for the required API.

Running the evaluation

uv run python -m bench.harness.generate --out bench/results/<run> --reps 3   # Claude API
uv run python bench/kaggle/make_kernel.py eval --user <kaggle-user> \
    --candidates bench/results/<run>/generations.jsonl
uvx --from kaggle kaggle kernels push -p bench/kaggle/build/eval             # Kaggle T4
uvx --from kaggle kaggle kernels output <kaggle-user>/mathlib-mcp-eval -p out
cp out/execution*.json* bench/results/<run>/
uv run python -m bench.harness.report bench/results/<run>
uv run python -m bench.harness.chart bench/results/<run> --out docs

The harness gets the tool definitions and handlers from the real MCP server, calls them in-process, and checks on every call that the responding model is the one requested. It keeps infrastructure errors out of the scores, and stores the full transcript of every attempt.

GPU validation

Before any model is scored, everything is checked on a real GPU: the canonical solutions must pass, the naive trap solutions must fail, every curated example must run, and every behavioural pitfall must hold. Latest run (report):

Tesla T4 (sm_75, 15 GB) · CUDA 12.9 · CuPy 14.2.0 · nvmath-python 1.0.0

Result

Canonical task solutions pass

26 / 26

Naive trap solutions fail as designed

17 / 17

Curated card examples pass

39 / 39

Pitfall facts hold

26 / 26

The traps fail for the intended reasons. Examples: OutOfMemoryError: ... 28,800,000,000 bytes (svd with full matrices), ValueError: The M dimension of the bias vector (48) must match the M dimension of A (per-row bias epilog), and TypeError: ... must be a matrix or vector with col-major layout (cuDSS). The first run also found something the docstrings do not say: nvmath-python's inverse FFTs are unnormalized (ifft(fft(x)) == n * x, unlike NumPy and CuPy). That finding is now a checked fact, a pitfall on four cards, and a benchmark task.

uv run python bench/kaggle/make_kernel.py validate --user <kaggle-user>   # self-contained script
uvx --from kaggle kaggle kernels push -p bench/kaggle/build/validate      # runs on a Kaggle T4
uvx --from kaggle kaggle kernels output <kaggle-user>/mathlib-mcp-gpu-validation -p out
uv run python -m bench.summarize_validation out/gpu_validation.json
uv run python -m bench.record_verification out/gpu_validation.json && uv run mathlib-mcp-build

Development

uv sync
uv run python scripts/fetch_sources.py   # pinned CuPy / nvmath-python sources -> .sources/
uv run mathlib-mcp-build                 # regenerate src/mathlib_mcp/data/
uv run pytest                            # includes "packaged data is up to date"
uv run ruff check .

MIT licensed.

Available Tools

4 tools
check_snippetA
Read-only

Statically check Python code that uses CuPy / nvmath-python without running it: reports functions, modules and enum members that do not exist (with the closest real names), unknown or positional-only keyword arguments, too many positional arguments and missing required arguments.

    Args:
        code: Python source code.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered; the description adds the stronger, non-obvious guarantee that the code is never executed and that diagnostics include suggested real names for misspelled symbols. It does not describe the shape of the returned report, but the static-analysis behavior is otherwise well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first clause and the diagnostic list is informative rather than padded. The trailing 'Args:' block restates a single obvious parameter, which is mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description partially carries the burden by enumerating what the check reports (unknown functions/modules/enum members, bad keyword or positional arguments, arity errors), which is effectively the return content. It stops short of describing the report format or whether results are structured, but it is adequate for a one-parameter, read-only analysis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter carries only the title 'Code'. The description compensates minimally with 'code: Python source code,' which confirms the input is the snippet to analyze but adds no constraints such as completeness requirements, size limits, or whether imports must be present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Statically check Python code that uses CuPy / nvmath-python') and enumerates the exact classes of diagnostics it reports, so the agent knows precisely what the tool produces. It never references the sibling retrieval tools (search_api, get_api_card, get_example), so sibling differentiation is absent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'without running it' suggests this is the way to validate a snippet before execution or before suggesting it to a user. There is no explicit when-to-use, no when-not-to-use, and no mention of the sibling tools as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_api_cardA
Read-only

Get the exact signature, parameters, return value, pitfalls and a GPU-verified example for one function or class, e.g. "cupyx.scipy.sparse.linalg.cg".

    Args:
        name: Fully qualified dotted name.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description goes beyond them by enumerating the payload (exact signature, parameters, return value, pitfalls, GPU-verified example), which tells the agent what a call yields. It does not cover failure behavior for an unknown or ambiguous name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The returning-content clause is front-loaded and the whole description is two short sentences plus an args line. The 'Args: name:' block partially restates the schema, a minor redundancy, but it carries the format detail that makes it worthwhile.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly enumerates what comes back, and the single parameter's expected form is given. Missing only disambiguation from sibling tools and what happens when the name is not found or is ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% – the schema only says name is a string. The description compensates by specifying the required format ('Fully qualified dotted name') and supplying a concrete example ('cupyx.scipy.sparse.linalg.cg'), which is exactly the guidance an agent needs to avoid passing a bare symbol.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: it retrieves an API card containing signature, parameters, return value, pitfalls and a verified example for a single named callable, with a concrete example name. It does not explicitly contrast itself with the siblings search_api/get_example/check_snippet, so an agent must infer that this is the exact-name lookup rather than a fuzzy search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the example dotted name and the word 'exact' suggest it should be used when the fully qualified symbol is already known. There is no explicit statement of when to use this instead of search_api (unknown name) or get_example (want only an example), and no error/precondition guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_exampleA
Read-only

Get runnable, GPU-verified example code (with pitfalls) for the curated APIs most relevant to a task, e.g. "matmul with bias and relu epilog".

    Args:
        task: Short description of what the code should do.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered structurally. The description adds genuine value by disclosing output qualities: the code is runnable, GPU-verified, and includes pitfalls, which is information the annotations do not carry. It doesn't state anything about result size, freshness, or whether the example is unique per call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core capability followed by a concrete example and a compact Args block. The Args restatement of the single parameter is slightly redundant given the schema, but it is short and harmless rather than verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter retrieval tool with no output schema and readOnly annotations, the description covers what is returned (verified example code with pitfalls) and what input is expected, which is sufficient to invoke correctly. Missing only sibling routing guidance, which is a usage-guideline concern rather than a completeness gap in the return contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden: it defines task as 'Short description of what the code should do' and supplies a worked example, which is more than the schema's bare string type. The definition is adequate for a single required parameter, though it doesn't state how detailed or keyword-oriented the task string should be.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Get) and a concrete resource (runnable, GPU-verified example code) plus the scope (curated APIs most relevant to a task), with a concrete example. It does not explicitly differentiate itself from the closest sibling check_snippet, which also sounds like code retrieval, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the example ('matmul with bias and relu epilog') and the phrase 'for a task', which tells the agent the input shape expected. However, there is no explicit when-to-use vs. when-not guidance and no mention of the alternatives search_api, get_api_card, or check_snippet, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_apiA
Read-only

Search CuPy / nvmath-python APIs (cuBLAS, cuFFT, cuSOLVER, cuSPARSE, cuDSS, cuTENSOR) by task description or name, e.g. "batched least squares" or "rfft".

    Args:
        query: What the code needs to do, or part of a function name.
        library: Optional prefix filter such as "cupy", "cupyx.scipy.sparse" or "nvmath".
        limit: Maximum number of results (1-20).
    
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
libraryNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the query semantics and the 1-20 limit bound, but says nothing about ranking behavior, whether results are paginated, or what a result contains, which matters for a search tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence with examples, and the Args block is compact. The Args section partially duplicates the schema property names, but given 0% schema description coverage it earns its place as the compensating documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter discovery tool with no output schema, the description covers inputs well but leaves the result shape unspecified (names only? signatures? snippets?) and does not indicate the follow-up workflow to get_api_card. It is callable as-is but not fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load, and it does document all three parameters: query as 'what the code needs to do, or part of a function name', library as an optional prefix filter with concrete prefix examples ('cupy', 'cupyx.scipy.sparse', 'nvmath'), and limit as max results 1-20 (a bound the schema does not state). It stops short of describing the matching/ranking semantics that would fully disambiguate vague queries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Search CuPy / nvmath-python APIs') and scopes the covered libraries (cuBLAS, cuFFT, cuSOLVER, cuSPARSE, cuDSS, cuTENSOR) with concrete query examples. An agent can distinguish it from get_api_card/get_example by the 'search by task description' framing, though it never explicitly contrasts itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through examples ('batched least squares', 'rfft') and the library-prefix filter, which show the intended input style. However, there is no explicit when-to-use/when-not guidance, e.g. that this is the discovery step before get_api_card, or when to prefer an exact-name lookup over a task-description search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcheck_snippet
    • First observedget_api_card
    • First observedget_example
    • First observedsearch_api

TDQS

A3.9/5.0

Scored across 4 tools

Disambiguation4/5

Each tool has a fairly distinct role: search_api discovers APIs by task/name, get_api_card details one exact name, get_example returns task-based example code, and check_snippet validates user code. The main overlap is between search_api and get_example, since both accept a task description and surface relevant APIs/examples, but the distinction (discovery vs. runnable example) is workable.

Naming Consistency5/5

All four tools follow a clean, predictable verb_noun snake_case pattern (search_api, get_api_card, get_example, check_snippet). No mixed conventions or vague verbs; names clearly signal the action.

Tool Count5/5

Four tools is well-scoped for a focused API-assistant server, covering the natural workflow of discover, inspect, exemplify, and validate. Each tool earns its place with no redundant surface.

Completeness4/5

The surface covers discovery, detail, examples, and static validation, which handles the core code-writing lifecycle for CuPy/nvmath-python. Minor gaps exist, such as no way to browse/list a module's full API set or check multiple snippets, but agents can work around these via search.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to search code by meaning, explore codebase structure, store and query knowledge with temporal facts, and read source code through a set of MCP tools.
    301 npm
    7
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Provides AI agents with version-pinned NVIDIA CUDA-Q documentation, API reference, and runnable examples via MCP, using offline SQLite full-text search. It resolves docs to match the installed cudaq package to avoid version skew.
    5
    27 PyPI
    1
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI coding agents to search and retrieve verified, reusable engineering solutions from a shared knowledge network through MCP tools.
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI assistants to access up-to-date library documentation through MCP tools for resolving library IDs and querying docs.
    55 npm
    MIT