agmind-mcp
# agmind-mcp
MCP server for measured local-LLM benchmarks. It exposes the [AGmind Systems Lab](https://agmind.ai) claim registry, currently 40 published claims measured on AMD Strix Halo hardware (Ryzen AI Max+ 395, Radeon 8060S, 128 GB unified memory) running llama.cpp on Vulkan and ROCm backends, as three read-only [Model Context Protocol](https://modelcontextprotocol.io) tools. Two NVIDIA DGX Spark (GB10) nodes are on the same lab bench; their claims enter the registry as runs are published. The lab has separately published vLLM work on DGX Spark; registry claims for it follow the same pipeline.
Every claim is a specific measured number: time to first token, inter-token latency, task success rate, answerless-response rate, long-context needle success, endurance drift. Each carries the exact hardware, runtime build, model revision and quantization, a frozen workload scope, stated limitations, an evidence level, links to the raw run records, and a ready-made citation string. Values are re-derived from raw runs on every CI build of the registry, so the numbers a model quotes through this server match the published evidence.
## Quickstart
Requires Node.js 18 or newer. No install step is needed; `npx` fetches the server from GitHub. Pin the release tag (`#v1.0.1`) so what runs on your machine is a fixed, reviewable commit rather than whatever `main` holds at the moment you start it. The package is **not published on npm**: only the `github:` form below is ours, and a package named `agmind-mcp` on the npm registry would not be.
**Claude Code**
```sh
claude mcp add agmind -- npx -y github:botAGI/agmind-mcp#v1.0.1
```
**Claude Desktop** (`claude_desktop_config.json`) and other MCP clients that take the standard config shape:
```json
{
"mcpServers": {
"agmind": {
"command": "npx",
"args": ["-y", "github:botAGI/agmind-mcp#v1.0.1"]
}
}
}
```
Releases and their commits: <https://github.com/botAGI/agmind-mcp/releases>.
From a local clone:
```sh
npm install
node server.mjs # speaks MCP over stdio
npm test # spawns the server and drives a real MCP session
```
## Tools
All three tools are read-only. Results are JSON in a text content block, and every claim in every result carries its `cite` string and `permalink` so agents can attribute what they quote.
### `search_claims`
Keyword search over headline, metric, system, model, runtime, scope, and id. Case-insensitive; every whitespace-separated term must match.
```
search_claims({ "query": "ttft 32k" })
```
Returns `{id, headline, value, unit, evidence_level, permalink, cite}` per match. Useful queries: `decode`, `answerless`, `ttft cache`, `rocm`, `task-success`, `endurance`.
### `get_claim`
One claim in full by id: the complete answer paragraph, measured value and unit, workload scope, aggregation, limitations, evidence level, raw run ids with GitHub links, the derivation SQL, permalink, and citation string.
```
get_claim({ "id": "strix.qwen36.docsession.c1.ttft-q2-32k-cache" })
```
An unknown id returns an error listing the closest matching ids.
### `list_measured`
The distinct system × model × runtime combinations that have published claims, with claim counts and example ids. Call this first to see what has actually been measured.
```
list_measured({})
```
## Data, license, attribution
- Server code: Apache-2.0.
- Claim data: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/), attribution **AGmind Systems Lab (agmind.ai)**. Each tool result includes a per-claim `cite` string ready to paste; reuse of the numbers should keep the claim permalink.
- Registry source: <https://agmind.ai/claims.json>. Raw run records and derivation SQL: [botAGI/agmind-lab](https://github.com/botAGI/agmind-lab). Benchmark harness and corpora: [botAGI/agmind-bench](https://github.com/botAGI/agmind-bench).
- Methodology, evidence levels, and errata: [agmind.ai/methodology](https://agmind.ai/methodology/), [agmind.ai/errata](https://agmind.ai/errata/).
## Behavior notes
- Read-only. The server never writes anything anywhere.
- No telemetry, no analytics, no accounts. The only network call is fetching the registry from agmind.ai.
- The registry is fetched at startup and cached in memory for one hour; a failed refetch falls back to the cached copy. The fetch times out after fifteen seconds and the response is shape-checked (a `claims` array of records with string ids) before anything is served to a model. Set `AGMIND_CLAIMS_URL` to point at a mirror of the registry if needed.
- Tool inputs are used only for case-insensitive substring matching over the cached registry; the server builds no regular expressions from them and runs nothing.
TDQS
Scored across 3 tools
Each tool serves a clearly distinct purpose: get_claim retrieves a single record by ID, list_measured provides an overview of measured combinations, and search_claims performs keyword queries. There is no overlap in functionality or potential for misselection.
All tool names follow a consistent snake_case verb_noun pattern (get_claim, list_measured, search_claims), which is predictable and easy to remember. No mixed conventions or ambiguous verbs.
With 3 tools, the set is compact yet sufficient for the server's purpose of querying a claim registry. Each tool earns its place, covering retrieval, overview, and search without unnecessary bloat.
For a read-only registry, the surface is complete: users can search, list overall combinations, and fetch full details by ID. There are no missing lifecycle operations (create/update/delete) that would apply to this domain, so no dead ends remain.