agmind-mcp
agmind-mcp
用于已测量的本地 LLM 基准测试的 MCP 服务器。它公开了 AGmind Systems Lab 的声明注册表,目前有 40 项已发布的声明,在 AMD Strix Halo 硬件(Ryzen AI Max+ 395、Radeon 8060S、128 GB 统一内存)上运行 llama.cpp 的 Vulkan 和 ROCm 后端进行测量,以三个只读的 Model Context Protocol 工具的形式提供。两个 NVIDIA DGX Spark (GB10) 节点位于同一实验台上;它们的声明在运行发布后进入注册表。该实验室还单独发布了 DGX Spark 上的 vLLM 工作;其注册表声明遵循相同的流程。
每项声明都是一个具体的测量数值:首个 token 时间、token 间延迟、任务成功率、无答案响应率、长上下文针测试成功率、持久性漂移。每项都包含确切的硬件、运行时构建、模型版本和量化、冻结的工作负载范围、声明的限制、证据级别、原始运行记录的链接,以及现成的引用字符串。值在注册表的每次 CI 构建中从原始运行重新推导,因此模型通过此服务器引用的数字与已发布的证据一致。
快速开始
需要 Node.js 18 或更高版本。无需安装步骤;npx 从 GitHub 获取服务器。
Claude Code
claude mcp add agmind -- npx -y github:botAGI/agmind-mcpClaude Desktop(claude_desktop_config.json)以及其他接受标准配置格式的 MCP 客户端:
{
"mcpServers": {
"agmind": {
"command": "npx",
"args": ["-y", "github:botAGI/agmind-mcp"]
}
}
}从本地克隆:
npm install
node server.mjs # speaks MCP over stdio
npm test # spawns the server and drives a real MCP sessionRelated MCP server: Local AI MCP
工具
所有三个工具都是只读的。结果是以文本内容块形式呈现的 JSON,每个结果中的每项声明都带有 cite 字符串和 permalink,以便代理可以标注其引用的内容。
search_claims
对 headline、metric、system、model、runtime、scope 和 id 进行关键字搜索。不区分大小写;每个以空格分隔的术语都必须匹配。
search_claims({ "query": "ttft 32k" })每个匹配返回 {id, headline, value, unit, evidence_level, permalink, cite}。有用的查询:decode、answerless、ttft cache、rocm、task-success、endurance。
get_claim
按 id 获取一项完整声明:完整的答案段落、实测的 value 和 unit、workload scope、aggregation、limitations、evidence level、带 GitHub 链接的 raw run ids、derivation SQL、permalink 和 citation string。
get_claim({ "id": "strix.qwen36.docsession.c1.ttft-q2-32k-cache" })未知 id 返回错误,列出最接近的匹配 id。
list_measured
已发布声明的不同 system × model × runtime 组合,包含声明计数和示例 id。先调用此工具以查看实际测量了什么。
list_measured({})数据、许可证、归属
服务器代码:Apache-2.0。
声明数据:CC BY 4.0,归属 AGmind Systems Lab (agmind.ai)。每个工具结果都包含可粘贴的每项声明
cite字符串;数字的复用应保留声明 permalink。注册表源:https://agmind.ai/claims.json。原始运行记录和推导 SQL:botAGI/agmind-lab。基准测试工具和语料库:botAGI/agmind-bench。
方法论、证据级别和勘误:agmind.ai/methodology、agmind.ai/errata。
行为说明
只读。服务器从不写入任何地方。
无遥测、无分析、无账户。唯一的网络调用是从 agmind.ai 获取注册表。
注册表在启动时获取并在内存中缓存一小时;失败的重新获取会回退到缓存副本。如果需要,设置
AGMIND_CLAIMS_URL指向注册表的镜像。
Available Tools
3 toolsget_claimGet one claim in fullARead-only
Fetch a single claim from the AGmind registry by id: full statement, measured value, unit, scope, limitations, evidence level, raw run ids and links, permalink, and the ready-made citation string. Ids look like "strix.qwen36.interactive2.c1.ttfa-nothink".
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Claim id, e.g. "strix.qwen36.longctx.c1.ttft-32k-en" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description does not need to emphasize safety. It adds value by disclosing the full set of returned fields, including raw run ids, links, permalink, and citation string, which goes beyond the schema. No contradiction is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core action ('Fetch a single claim'), followed by a list of contents and an id example. Every sentence adds value, and the format is ideal for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-id lookup tool with one parameter, complete schema coverage, and a readOnly annotation, the description fully covers what the agent needs to know: what it returns, what the id looks like, and the tool's non-mutating nature. No output schema exists, but the description lists the return fields explicitly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description reinforces the id format with an example ('strix.qwen36.interactive2.c1.ttfa-nothink') and mentions the id pattern. This adds practical guidance beyond the schema's basic type description, though the schema already documents the parameter adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a single claim by id and enumerates the exact fields returned, distinguishing it from siblings like list_measured and search_claims. The verb 'Fetch' and resource 'single claim from the AGmind registry' are specific, and the inclusion of an example id format adds clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a specific claim id is known and full detail is needed, but it does not explicitly state when not to use it or mention alternatives. Sibling tools exist (list_measured, search_claims), so more explicit guidance could help, but the context is clear enough for a targeted lookup tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_measuredList measured configurationsARead-only
List the distinct system × model × runtime combinations that have published, measured claims in the AGmind registry, with claim counts and example claim ids (with permalinks). Start here to see what hardware and models the lab has qualified.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so the description doesn't need to state it's a read operation. The description adds useful context about the output: includes claim counts and example claim ids with permalinks, which is beyond the annotations. However, it doesn't disclose any other behavioral traits like pagination or result ordering, which could be relevant for a list operation. The description adds some value but not extensive behavioral context beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. It states the core function in the first sentence, then adds a helpful guidance sentence about starting here. Every word earns its place; there is no fluff or repetition of schema/annotation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description does a good job explaining what the tool returns (combinations, counts, example IDs) and gives a clear use case. It could potentially mention the format of permalinks or how to interpret claim counts, but for a simple list tool with good guidance, it is complete enough. The sibling tools (get_claim, search_claims) provide further context for follow-up actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool has zero parameters, so the schema provides no parameter information. The description doesn't need to explain parameters, but it does clarify what the output represents (distinct combinations, claim counts, example IDs). With 0 parameters, baseline is 4 per the rubric, and the description adds sufficient context about the tool's purpose and output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists distinct system × model × runtime combinations with measured claims, including claim counts and example claim IDs with permalinks. It distinguishes from siblings by mentioning 'measured claims' and 'published' in the AGmind registry, which sets it apart from get_claim (retrieving a single claim) and search_claims (searching claims).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Start here to see what hardware and models the lab has qualified.' This explicitly guides the agent to use this tool as an entry point for exploring measured configurations. However, it does not explicitly state when not to use it or mention alternatives like search_claims, so it misses the when-not/alternatives aspect for a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_claimsSearch measured benchmark claimsARead-only
Keyword search over the AGmind claim registry of measured local-LLM benchmarks (hardware, runtime, model, metric, workload scope). Case-insensitive; every whitespace-separated term must match. Returns claim summaries with value, unit, evidence level, permalink, and a ready-made citation string. Use get_claim for the full record including limitations and raw run ids.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Keywords, e.g. "decode", "ttft 32k", "qwen vulkan", "answerless" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only and closed-world behavior. The description adds genuinely useful behavioral details beyond that: case-insensitive matching, the requirement that every whitespace-separated term matches, and the specific summary fields returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences cover the resource, matching behavior, return contents, and the main alternative. Every sentence earns its place with no wasted words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter search tool with no output schema, the description is thorough: it states the registry scope, search semantics, return fields, and the natural follow-up tool for deeper detail. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single query parameter with examples. The description adds extra meaning by explaining the matching rule ('every whitespace-separated term must match') and the scope of keywords across hardware, runtime, model, metric, and workload.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Keyword search over the AGmind claim registry of measured local-LLM benchmarks.' It also clearly distinguishes from get_claim by directing users to that tool for full records, and the keyword-search framing differentiates it from list_measured.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool and explicitly points to get_claim as the alternative for full records. However, it does not explicitly address when list_measured would be preferred over this keyword search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
get_claim - First observed
list_measured - First observed
search_claims
TDQS
Scored across 3 tools
Each tool serves a clearly distinct purpose: get_claim retrieves a single record by ID, list_measured provides an overview of measured combinations, and search_claims performs keyword queries. There is no overlap in functionality or potential for misselection.
All tool names follow a consistent snake_case verb_noun pattern (get_claim, list_measured, search_claims), which is predictable and easy to remember. No mixed conventions or ambiguous verbs.
With 3 tools, the set is compact yet sufficient for the server's purpose of querying a claim registry. Each tool earns its place, covering retrieval, overview, and search without unnecessary bloat.
For a read-only registry, the surface is complete: users can search, list overall combinations, and fetch full details by ID. There are no missing lifecycle operations (create/update/delete) that would apply to this domain, so no dead ends remain.
Maintenance
Related MCP Connectors
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Related MCP Servers
- AlicenseAqualityBmaintenanceAn MCP server that enables LLMs to pull-based search through Clawket's RAG repository for exploratory and conditional queries. It provides read-only access to search artifacts, tasks, and decisions via HTTP API.510 npmMIT

Local AI MCPofficial
AlicenseAqualityCmaintenanceUnified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.1619 npmCreative Commons Attribution Non Commercial No Derivatives 4.0 International- AlicenseCqualityAmaintenanceRead-only MCP server that exposes public TokenLab model catalog tools for agents to discover models, inspect request contracts, and compare pricing.32876 npmMIT
- AlicenseNot gradedqualityCmaintenanceA read-only MCP server for operator-grade release inspection and benchmark browsing.35 npm1MIT