agmind-mcp
agmind-mcp
MCP-сервер для измеренных бенчмарков локальных LLM. Он предоставляет реестр утверждений AGmind Systems Lab — в настоящее время 40 опубликованных утверждений, измеренных на оборудовании AMD Strix Halo (Ryzen AI Max+ 395, Radeon 8060S, 128 ГБ унифицированной памяти) с запуском llama.cpp на бэкендах Vulkan и ROCm, — в виде трёх инструментов только для чтения Model Context Protocol. Два узла NVIDIA DGX Spark (GB10) находятся на том же лабораторном стенде; их утверждения попадают в реестр по мере публикации прогонов. Лаборатория отдельно опубликовала работы по vLLM на DGX Spark; утверждения реестра для них следуют тому же конвейеру.
Каждое утверждение — это конкретное измеренное число: время до первого токена, межтокенная задержка, доля успешных задач, доля ответов без ответа, успех поиска иголки в длинном контексте, дрейф выносливости. Каждое содержит точное оборудование, сборку рантайма, ревизию модели и квантизацию, зафиксированную область нагрузки, указанные ограничения, уровень доказательности, ссылки на исходные записи прогонов и готовую строку цитирования. Значения пересчитываются из исходных прогонов при каждой CI-сборке реестра, поэтому числа, которые модель цитирует через этот сервер, соответствуют опубликованным доказательствам.
Быстрый старт
Требуется Node.js 18 или новее. Установка не требуется; npx загружает сервер с GitHub.
Claude Code
claude mcp add agmind -- npx -y github:botAGI/agmind-mcpClaude Desktop (claude_desktop_config.json) и другие MCP-клиенты, которые принимают стандартную форму конфигурации:
{
"mcpServers": {
"agmind": {
"command": "npx",
"args": ["-y", "github:botAGI/agmind-mcp"]
}
}
}Из локального клона:
npm install
node server.mjs # speaks MCP over stdio
npm test # spawns the server and drives a real MCP sessionRelated MCP server: Local AI MCP
Инструменты
Все три инструмента доступны только для чтения. Результаты — это JSON в текстовом блоке содержимого, и каждое утверждение в каждом результате содержит строку cite и permalink, чтобы агенты могли указать источник цитируемого.
search_claims
Поиск по ключевым словам по заголовку, метрике, системе, модели, рантайму, области и id. Регистр не учитывается; каждый разделённый пробелами термин должен совпадать.
search_claims({ "query": "ttft 32k" })Возвращает {id, headline, value, unit, evidence_level, permalink, cite} для каждого совпадения. Полезные запросы: decode, answerless, ttft cache, rocm, task-success, endurance.
get_claim
Одно утверждение полностью по id: полный абзац ответа, измеренное значение и единица измерения, область нагрузки, агрегация, ограничения, уровень доказательности, идентификаторы исходных прогонов со ссылками на GitHub, SQL-вывод, permalink и строка цитирования.
get_claim({ "id": "strix.qwen36.docsession.c1.ttft-q2-32k-cache" })Неизвестный id возвращает ошибку со списком наиболее близких совпадающих id.
list_measured
Уникальные комбинации система × модель × рантайм, для которых опубликованы утверждения, с количеством утверждений и примерами id. Сначала вызовите этот инструмент, чтобы увидеть, что реально было измерено.
list_measured({})Данные, лицензия, атрибуция
Код сервера: Apache-2.0.
Данные утверждений: CC BY 4.0, атрибуция AGmind Systems Lab (agmind.ai). Каждый результат инструмента включает строку
citeдля каждого утверждения, готовую к вставке; при повторном использовании чисел следует сохранять permalink утверждения.Источник реестра: https://agmind.ai/claims.json. Исходные записи прогонов и SQL-вывод: botAGI/agmind-lab. Бенчмарк-харнесс и корпуса: botAGI/agmind-bench.
Методология, уровни доказательности и эррата: agmind.ai/methodology, agmind.ai/errata.
Примечания по поведению
Только чтение. Сервер никогда ничего никуда не записывает.
Никакой телеметрии, никакой аналитики, никаких аккаунтов. Единственный сетевой вызов — получение реестра с agmind.ai.
Реестр загружается при запуске и кэшируется в памяти на один час; при неудачной повторной загрузке используется кэшированная копия. Установите
AGMIND_CLAIMS_URL, чтобы указать на зеркало реестра, если это необходимо.
Available Tools
3 toolsget_claimGet one claim in fullARead-only
Fetch a single claim from the AGmind registry by id: full statement, measured value, unit, scope, limitations, evidence level, raw run ids and links, permalink, and the ready-made citation string. Ids look like "strix.qwen36.interactive2.c1.ttfa-nothink".
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Claim id, e.g. "strix.qwen36.longctx.c1.ttft-32k-en" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description does not need to emphasize safety. It adds value by disclosing the full set of returned fields, including raw run ids, links, permalink, and citation string, which goes beyond the schema. No contradiction is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core action ('Fetch a single claim'), followed by a list of contents and an id example. Every sentence adds value, and the format is ideal for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-id lookup tool with one parameter, complete schema coverage, and a readOnly annotation, the description fully covers what the agent needs to know: what it returns, what the id looks like, and the tool's non-mutating nature. No output schema exists, but the description lists the return fields explicitly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description reinforces the id format with an example ('strix.qwen36.interactive2.c1.ttfa-nothink') and mentions the id pattern. This adds practical guidance beyond the schema's basic type description, though the schema already documents the parameter adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a single claim by id and enumerates the exact fields returned, distinguishing it from siblings like list_measured and search_claims. The verb 'Fetch' and resource 'single claim from the AGmind registry' are specific, and the inclusion of an example id format adds clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a specific claim id is known and full detail is needed, but it does not explicitly state when not to use it or mention alternatives. Sibling tools exist (list_measured, search_claims), so more explicit guidance could help, but the context is clear enough for a targeted lookup tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_measuredList measured configurationsARead-only
List the distinct system × model × runtime combinations that have published, measured claims in the AGmind registry, with claim counts and example claim ids (with permalinks). Start here to see what hardware and models the lab has qualified.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so the description doesn't need to state it's a read operation. The description adds useful context about the output: includes claim counts and example claim ids with permalinks, which is beyond the annotations. However, it doesn't disclose any other behavioral traits like pagination or result ordering, which could be relevant for a list operation. The description adds some value but not extensive behavioral context beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. It states the core function in the first sentence, then adds a helpful guidance sentence about starting here. Every word earns its place; there is no fluff or repetition of schema/annotation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description does a good job explaining what the tool returns (combinations, counts, example IDs) and gives a clear use case. It could potentially mention the format of permalinks or how to interpret claim counts, but for a simple list tool with good guidance, it is complete enough. The sibling tools (get_claim, search_claims) provide further context for follow-up actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool has zero parameters, so the schema provides no parameter information. The description doesn't need to explain parameters, but it does clarify what the output represents (distinct combinations, claim counts, example IDs). With 0 parameters, baseline is 4 per the rubric, and the description adds sufficient context about the tool's purpose and output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists distinct system × model × runtime combinations with measured claims, including claim counts and example claim IDs with permalinks. It distinguishes from siblings by mentioning 'measured claims' and 'published' in the AGmind registry, which sets it apart from get_claim (retrieving a single claim) and search_claims (searching claims).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Start here to see what hardware and models the lab has qualified.' This explicitly guides the agent to use this tool as an entry point for exploring measured configurations. However, it does not explicitly state when not to use it or mention alternatives like search_claims, so it misses the when-not/alternatives aspect for a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_claimsSearch measured benchmark claimsARead-only
Keyword search over the AGmind claim registry of measured local-LLM benchmarks (hardware, runtime, model, metric, workload scope). Case-insensitive; every whitespace-separated term must match. Returns claim summaries with value, unit, evidence level, permalink, and a ready-made citation string. Use get_claim for the full record including limitations and raw run ids.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Keywords, e.g. "decode", "ttft 32k", "qwen vulkan", "answerless" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only and closed-world behavior. The description adds genuinely useful behavioral details beyond that: case-insensitive matching, the requirement that every whitespace-separated term matches, and the specific summary fields returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences cover the resource, matching behavior, return contents, and the main alternative. Every sentence earns its place with no wasted words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter search tool with no output schema, the description is thorough: it states the registry scope, search semantics, return fields, and the natural follow-up tool for deeper detail. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single query parameter with examples. The description adds extra meaning by explaining the matching rule ('every whitespace-separated term must match') and the scope of keywords across hardware, runtime, model, metric, and workload.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Keyword search over the AGmind claim registry of measured local-LLM benchmarks.' It also clearly distinguishes from get_claim by directing users to that tool for full records, and the keyword-search framing differentiates it from list_measured.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool and explicitly points to get_claim as the alternative for full records. However, it does not explicitly address when list_measured would be preferred over this keyword search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
get_claim - First observed
list_measured - First observed
search_claims
TDQS
Scored across 3 tools
Each tool serves a clearly distinct purpose: get_claim retrieves a single record by ID, list_measured provides an overview of measured combinations, and search_claims performs keyword queries. There is no overlap in functionality or potential for misselection.
All tool names follow a consistent snake_case verb_noun pattern (get_claim, list_measured, search_claims), which is predictable and easy to remember. No mixed conventions or ambiguous verbs.
With 3 tools, the set is compact yet sufficient for the server's purpose of querying a claim registry. Each tool earns its place, covering retrieval, overview, and search without unnecessary bloat.
For a read-only registry, the surface is complete: users can search, list overall combinations, and fetch full details by ID. There are no missing lifecycle operations (create/update/delete) that would apply to this domain, so no dead ends remain.
Maintenance
Related MCP Connectors
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Related MCP Servers
- AlicenseAqualityBmaintenanceAn MCP server that enables LLMs to pull-based search through Clawket's RAG repository for exploratory and conditional queries. It provides read-only access to search artifacts, tasks, and decisions via HTTP API.510 npmMIT

Local AI MCPofficial
AlicenseAqualityCmaintenanceUnified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.1619 npmCreative Commons Attribution Non Commercial No Derivatives 4.0 International- AlicenseCqualityAmaintenanceRead-only MCP server that exposes public TokenLab model catalog tools for agents to discover models, inspect request contracts, and compare pricing.32876 npmMIT
- AlicenseNot gradedqualityCmaintenanceA read-only MCP server for operator-grade release inspection and benchmark browsing.35 npm1MIT