agmind-mcp
agmind-mcp
測定されたローカルLLMベンチマークのためのMCPサーバーです。AGmind Systems Labのクレームレジストリを公開しており、現在40件の公開クレームがAMD Strix Haloハードウェア(Ryzen AI Max+ 395、Radeon 8060S、128 GBユニファイドメモリ)上で、VulkanおよびROCmバックエンドでllama.cppを実行して測定されています。これらは3つの読み取り専用のModel Context Protocolツールとして提供されます。同じラボベンチには2台のNVIDIA DGX Spark(GB10)ノードがあり、実行が公開されるにつれてクレームがレジストリに追加されます。ラボはDGX Spark上でのvLLMの作業を別途公開しています。そのクレームも同じパイプラインに従います。
すべてのクレームは具体的な測定値です:最初のトークンまでの時間、トークン間レイテンシ、タスク成功率、無回答率、長文脈ニードル成功率、耐久性ドリフト。それぞれに、正確なハードウェア、ランタイムビルド、モデルリビジョンと量子化、固定されたワークロードスコープ、明記された制限、エビデンスレベル、生の実行記録へのリンク、およびすぐに使える引用文字列が含まれます。値はレジストリのCIビルドごとに生の実行から再導出されるため、このサーバーを通じてモデルが引用する数値は公開されたエビデンスと一致します。
クイックスタート
Node.js 18以降が必要です。インストール手順は不要です。npxがGitHubからサーバーを取得します。
Claude Code
claude mcp add agmind -- npx -y github:botAGI/agmind-mcpClaude Desktop(claude_desktop_config.json)および標準の設定形式を受け入れる他のMCPクライアント:
{
"mcpServers": {
"agmind": {
"command": "npx",
"args": ["-y", "github:botAGI/agmind-mcp"]
}
}
}ローカルクローンから:
npm install
node server.mjs # speaks MCP over stdio
npm test # spawns the server and drives a real MCP sessionRelated MCP server: Local AI MCP
ツール
3つのツールはすべて読み取り専用です。結果はテキストコンテンツブロック内のJSONであり、各結果のすべてのクレームにはcite文字列とpermalinkが含まれているため、エージェントは引用した内容の出典を明示できます。
search_claims
見出し、メトリック、システム、モデル、ランタイム、スコープ、IDに対するキーワード検索。大文字と小文字を区別しません。空白で区切られたすべての用語が一致する必要があります。
search_claims({ "query": "ttft 32k" })一致ごとに{id, headline, value, unit, evidence_level, permalink, cite}を返します。便利なクエリ:decode、answerless、ttft cache、rocm、task-success、endurance。
get_claim
IDによる1つのクレームの完全な内容:完全な回答段落、測定値と単位、ワークロードスコープ、集計、制限、エビデンスレベル、GitHubリンク付きの生の実行ID、導出SQL、パーマリンク、引用文字列。
get_claim({ "id": "strix.qwen36.docsession.c1.ttft-q2-32k-cache" })未知のIDは、最も近い一致IDをリストしたエラーを返します。
list_measured
公開されたクレームを持つシステム×モデル×ランタイムの組み合わせの一覧で、クレーム数と例IDを含みます。これを最初に呼び出して、実際に何が測定されたかを確認します。
list_measured({})データ、ライセンス、帰属
サーバーコード:Apache-2.0。
クレームデータ:CC BY 4.0、帰属はAGmind Systems Lab (agmind.ai)。各ツール結果には、貼り付け可能なクレームごとの
cite文字列が含まれます。数値の再利用にはクレームのパーマリンクを保持してください。レジストリソース:https://agmind.ai/claims.json。生の実行記録と導出SQL:botAGI/agmind-lab。ベンチマークハーネスとコーパス:botAGI/agmind-bench。
方法論、エビデンスレベル、正誤表:agmind.ai/methodology、agmind.ai/errata。
動作上の注意
読み取り専用。サーバーはどこにも何も書き込みません。
テレメトリなし、分析なし、アカウントなし。唯一のネットワーク呼び出しはagmind.aiからのレジストリ取得です。
レジストリは起動時に取得され、1時間メモリにキャッシュされます。再取得に失敗した場合はキャッシュされたコピーにフォールバックします。必要に応じて
AGMIND_CLAIMS_URLを設定してレジストリのミラーを指定できます。
Available Tools
3 toolsget_claimGet one claim in fullARead-only
Fetch a single claim from the AGmind registry by id: full statement, measured value, unit, scope, limitations, evidence level, raw run ids and links, permalink, and the ready-made citation string. Ids look like "strix.qwen36.interactive2.c1.ttfa-nothink".
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Claim id, e.g. "strix.qwen36.longctx.c1.ttft-32k-en" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description does not need to emphasize safety. It adds value by disclosing the full set of returned fields, including raw run ids, links, permalink, and citation string, which goes beyond the schema. No contradiction is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core action ('Fetch a single claim'), followed by a list of contents and an id example. Every sentence adds value, and the format is ideal for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-id lookup tool with one parameter, complete schema coverage, and a readOnly annotation, the description fully covers what the agent needs to know: what it returns, what the id looks like, and the tool's non-mutating nature. No output schema exists, but the description lists the return fields explicitly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description reinforces the id format with an example ('strix.qwen36.interactive2.c1.ttfa-nothink') and mentions the id pattern. This adds practical guidance beyond the schema's basic type description, though the schema already documents the parameter adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a single claim by id and enumerates the exact fields returned, distinguishing it from siblings like list_measured and search_claims. The verb 'Fetch' and resource 'single claim from the AGmind registry' are specific, and the inclusion of an example id format adds clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a specific claim id is known and full detail is needed, but it does not explicitly state when not to use it or mention alternatives. Sibling tools exist (list_measured, search_claims), so more explicit guidance could help, but the context is clear enough for a targeted lookup tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_measuredList measured configurationsARead-only
List the distinct system × model × runtime combinations that have published, measured claims in the AGmind registry, with claim counts and example claim ids (with permalinks). Start here to see what hardware and models the lab has qualified.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so the description doesn't need to state it's a read operation. The description adds useful context about the output: includes claim counts and example claim ids with permalinks, which is beyond the annotations. However, it doesn't disclose any other behavioral traits like pagination or result ordering, which could be relevant for a list operation. The description adds some value but not extensive behavioral context beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. It states the core function in the first sentence, then adds a helpful guidance sentence about starting here. Every word earns its place; there is no fluff or repetition of schema/annotation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description does a good job explaining what the tool returns (combinations, counts, example IDs) and gives a clear use case. It could potentially mention the format of permalinks or how to interpret claim counts, but for a simple list tool with good guidance, it is complete enough. The sibling tools (get_claim, search_claims) provide further context for follow-up actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool has zero parameters, so the schema provides no parameter information. The description doesn't need to explain parameters, but it does clarify what the output represents (distinct combinations, claim counts, example IDs). With 0 parameters, baseline is 4 per the rubric, and the description adds sufficient context about the tool's purpose and output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists distinct system × model × runtime combinations with measured claims, including claim counts and example claim IDs with permalinks. It distinguishes from siblings by mentioning 'measured claims' and 'published' in the AGmind registry, which sets it apart from get_claim (retrieving a single claim) and search_claims (searching claims).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Start here to see what hardware and models the lab has qualified.' This explicitly guides the agent to use this tool as an entry point for exploring measured configurations. However, it does not explicitly state when not to use it or mention alternatives like search_claims, so it misses the when-not/alternatives aspect for a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_claimsSearch measured benchmark claimsARead-only
Keyword search over the AGmind claim registry of measured local-LLM benchmarks (hardware, runtime, model, metric, workload scope). Case-insensitive; every whitespace-separated term must match. Returns claim summaries with value, unit, evidence level, permalink, and a ready-made citation string. Use get_claim for the full record including limitations and raw run ids.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Keywords, e.g. "decode", "ttft 32k", "qwen vulkan", "answerless" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only and closed-world behavior. The description adds genuinely useful behavioral details beyond that: case-insensitive matching, the requirement that every whitespace-separated term matches, and the specific summary fields returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences cover the resource, matching behavior, return contents, and the main alternative. Every sentence earns its place with no wasted words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter search tool with no output schema, the description is thorough: it states the registry scope, search semantics, return fields, and the natural follow-up tool for deeper detail. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single query parameter with examples. The description adds extra meaning by explaining the matching rule ('every whitespace-separated term must match') and the scope of keywords across hardware, runtime, model, metric, and workload.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Keyword search over the AGmind claim registry of measured local-LLM benchmarks.' It also clearly distinguishes from get_claim by directing users to that tool for full records, and the keyword-search framing differentiates it from list_measured.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool and explicitly points to get_claim as the alternative for full records. However, it does not explicitly address when list_measured would be preferred over this keyword search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
get_claim - First observed
list_measured - First observed
search_claims
TDQS
Scored across 3 tools
Each tool serves a clearly distinct purpose: get_claim retrieves a single record by ID, list_measured provides an overview of measured combinations, and search_claims performs keyword queries. There is no overlap in functionality or potential for misselection.
All tool names follow a consistent snake_case verb_noun pattern (get_claim, list_measured, search_claims), which is predictable and easy to remember. No mixed conventions or ambiguous verbs.
With 3 tools, the set is compact yet sufficient for the server's purpose of querying a claim registry. Each tool earns its place, covering retrieval, overview, and search without unnecessary bloat.
For a read-only registry, the surface is complete: users can search, list overall combinations, and fetch full details by ID. There are no missing lifecycle operations (create/update/delete) that would apply to this domain, so no dead ends remain.
Maintenance
Related MCP Connectors
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Related MCP Servers
- AlicenseAqualityBmaintenanceAn MCP server that enables LLMs to pull-based search through Clawket's RAG repository for exploratory and conditional queries. It provides read-only access to search artifacts, tasks, and decisions via HTTP API.510 npmMIT

Local AI MCPofficial
AlicenseAqualityCmaintenanceUnified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.1619 npmCreative Commons Attribution Non Commercial No Derivatives 4.0 International- AlicenseCqualityAmaintenanceRead-only MCP server that exposes public TokenLab model catalog tools for agents to discover models, inspect request contracts, and compare pricing.32876 npmMIT
- AlicenseNot gradedqualityCmaintenanceA read-only MCP server for operator-grade release inspection and benchmark browsing.35 npm1MIT