Skip to main content
Glama
botAGI

agmind-mcp

by botAGI

agmind-mcp

측정된 로컬 LLM 벤치마크를 위한 MCP 서버입니다. 이 서버는 AGmind Systems Lab 클레임 레지스트리를 공개합니다. 현재 레지스트리에는 AMD Strix Halo 하드웨어(Ryzen AI Max+ 395, Radeon 8060S, 128 GB 통합 메모리)에서 Vulkan 및 ROCm 백엔드의 llama.cpp를 실행해 측정한 발표된 클레임 40개가 포함되어 있으며, 이를 세 가지 읽기 전용 Model Context Protocol 도구로 제공합니다. 같은 실험 벤치에 NVIDIA DGX Spark(GB10) 노드 두 대가 있으며, 실행이 발표되면 해당 노드의 클레임도 레지스트리에 등록됩니다. 연구소는 DGX Spark에서의 vLLM 작업을 별도로 발표했으며, 이에 대한 레지스트리 클레임도 동일한 파이프라인을 따릅니다.

모든 클레임은 구체적으로 측정된 하나의 수치입니다: time to first token, inter-token latency, task success rate, answerless-response rate, long-context needle success, endurance drift. 각 클레임에는 정확한 하드웨어, 런타임 빌드, 모델 버전 및 양자화, 고정된 워크로드 범위, 명시된 제한 사항, 증거 수준, 원시 실행 기록 링크, 그리고 바로 사용할 수 있는 인용 문자열이 포함되어 있습니다. 값은 레지스트리의 모든 CI 빌드에서 원시 실행 기록으로부터 다시 도출되므로, 모델이 이 서버를 통해 인용하는 수치는 게시된 증거와 일치합니다.

빠른 시작

Node.js 18 이상이 필요합니다. 설치 단계는 필요 없으며, npx가 GitHub에서 서버를 가져옵니다.

Claude Code

claude mcp add agmind -- npx -y github:botAGI/agmind-mcp

Claude Desktop (claude_desktop_config.json) 및 표준 구성 형식을 따르는 기타 MCP 클라이언트:

{
  "mcpServers": {
    "agmind": {
      "command": "npx",
      "args": ["-y", "github:botAGI/agmind-mcp"]
    }
  }
}

로컬 클론에서 사용하는 경우:

npm install
node server.mjs        # speaks MCP over stdio
npm test               # spawns the server and drives a real MCP session

Related MCP server: Local AI MCP

도구

세 도구 모두 읽기 전용입니다. 결과는 텍스트 콘텐츠 블록의 JSON이며, 각 결과의 모든 클레임에는 cite 문자열과 permalink가 포함되어 있어 에이전트가 인용한 내용의 출처를 명시할 수 있습니다.

search_claims

헤드라인, 지표, 시스템, 모델, 런타임, 범위 및 id에 대한 키워드 검색입니다. 대소문자를 구분하지 않으며, 공백으로 구분된 모든 용어가 일치해야 합니다.

search_claims({ "query": "ttft 32k" })

일치하는 항목마다 {id, headline, value, unit, evidence_level, permalink, cite}를 반환합니다. 유용한 검색어: decode, answerless, ttft cache, rocm, task-success, endurance.

get_claim

id로 클레임 하나를 전체 조회합니다: 완전한 답변 단락, 측정값과 단위, 워크로드 범위, 집계, 제한 사항, 증거 수준, GitHub 링크가 포함된 원시 실행 id, 파생 SQL, permalink, 그리고 인용 문자열입니다.

get_claim({ "id": "strix.qwen36.docsession.c1.ttft-q2-32k-cache" })

알 수 없는 id인 경우 오류를 반환하며, 가장 유사한 id 목록을 함께 보여 줍니다.

list_measured

발표된 클레임이 있는 서로 다른 시스템×모델×런타임 조합을 클레임 수와 예시 id와 함께 나열합니다. 실제로 무엇이 측정되었는지 보려면 이 도구를 먼저 호출하세요.

list_measured({})

데이터, 라이선스, 출처

동작 참고 사항

  • 읽기 전용입니다. 서버는 어디에도 쓰기 작업을 수행하지 않습니다.

  • 텔레메트리, 분석, 계정이 없습니다. 유일한 네트워크 호출은 agmind.ai에서 레지스트리를 가져오는 것입니다.

  • 레지스트리는 시작 시 가져와서 1시간 동안 메모리에 캐시됩니다. 재인출에실패하면 캐시된 사본으로 대체됩니다. 필요한 경우 AGMIND_CLAIMS_URL을 설정하여 레지스트리의 미러를 가리킬 수 있습니다.

Available Tools

3 tools
get_claimGet one claim in fullA
Read-only

Fetch a single claim from the AGmind registry by id: full statement, measured value, unit, scope, limitations, evidence level, raw run ids and links, permalink, and the ready-made citation string. Ids look like "strix.qwen36.interactive2.c1.ttfa-nothink".

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesClaim id, e.g. "strix.qwen36.longctx.c1.ttft-32k-en"

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description does not need to emphasize safety. It adds value by disclosing the full set of returned fields, including raw run ids, links, permalink, and citation string, which goes beyond the schema. No contradiction is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core action ('Fetch a single claim'), followed by a list of contents and an id example. Every sentence adds value, and the format is ideal for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-id lookup tool with one parameter, complete schema coverage, and a readOnly annotation, the description fully covers what the agent needs to know: what it returns, what the id looks like, and the tool's non-mutating nature. No output schema exists, but the description lists the return fields explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description reinforces the id format with an example ('strix.qwen36.interactive2.c1.ttfa-nothink') and mentions the id pattern. This adds practical guidance beyond the schema's basic type description, though the schema already documents the parameter adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches a single claim by id and enumerates the exact fields returned, distinguishing it from siblings like list_measured and search_claims. The verb 'Fetch' and resource 'single claim from the AGmind registry' are specific, and the inclusion of an example id format adds clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when a specific claim id is known and full detail is needed, but it does not explicitly state when not to use it or mention alternatives. Sibling tools exist (list_measured, search_claims), so more explicit guidance could help, but the context is clear enough for a targeted lookup tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_measuredList measured configurationsA
Read-only

List the distinct system × model × runtime combinations that have published, measured claims in the AGmind registry, with claim counts and example claim ids (with permalinks). Start here to see what hardware and models the lab has qualified.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, so the description doesn't need to state it's a read operation. The description adds useful context about the output: includes claim counts and example claim ids with permalinks, which is beyond the annotations. However, it doesn't disclose any other behavioral traits like pagination or result ordering, which could be relevant for a list operation. The description adds some value but not extensive behavioral context beyond the schema and annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded. It states the core function in the first sentence, then adds a helpful guidance sentence about starting here. Every word earns its place; there is no fluff or repetition of schema/annotation information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters and no output schema, the description does a good job explaining what the tool returns (combinations, counts, example IDs) and gives a clear use case. It could potentially mention the format of permalinks or how to interpret claim counts, but for a simple list tool with good guidance, it is complete enough. The sibling tools (get_claim, search_claims) provide further context for follow-up actions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

This tool has zero parameters, so the schema provides no parameter information. The description doesn't need to explain parameters, but it does clarify what the output represents (distinct combinations, claim counts, example IDs). With 0 parameters, baseline is 4 per the rubric, and the description adds sufficient context about the tool's purpose and output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists distinct system × model × runtime combinations with measured claims, including claim counts and example claim IDs with permalinks. It distinguishes from siblings by mentioning 'measured claims' and 'published' in the AGmind registry, which sets it apart from get_claim (retrieving a single claim) and search_claims (searching claims).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'Start here to see what hardware and models the lab has qualified.' This explicitly guides the agent to use this tool as an entry point for exploring measured configurations. However, it does not explicitly state when not to use it or mention alternatives like search_claims, so it misses the when-not/alternatives aspect for a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_claimsSearch measured benchmark claimsA
Read-only

Keyword search over the AGmind claim registry of measured local-LLM benchmarks (hardware, runtime, model, metric, workload scope). Case-insensitive; every whitespace-separated term must match. Returns claim summaries with value, unit, evidence level, permalink, and a ready-made citation string. Use get_claim for the full record including limitations and raw run ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesKeywords, e.g. "decode", "ttft 32k", "qwen vulkan", "answerless"

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only and closed-world behavior. The description adds genuinely useful behavioral details beyond that: case-insensitive matching, the requirement that every whitespace-separated term matches, and the specific summary fields returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences cover the resource, matching behavior, return contents, and the main alternative. Every sentence earns its place with no wasted words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter search tool with no output schema, the description is thorough: it states the registry scope, search semantics, return fields, and the natural follow-up tool for deeper detail. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the single query parameter with examples. The description adds extra meaning by explaining the matching rule ('every whitespace-separated term must match') and the scope of keywords across hardware, runtime, model, metric, and workload.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Keyword search over the AGmind claim registry of measured local-LLM benchmarks.' It also clearly distinguishes from get_claim by directing users to that tool for full records, and the keyword-search framing differentiates it from list_measured.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool and explicitly points to get_claim as the alternative for full records. However, it does not explicitly address when list_measured would be preferred over this keyword search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedget_claim
    • First observedlist_measured
    • First observedsearch_claims

TDQS

A4.5/5.0

Scored across 3 tools

Disambiguation5/5

Each tool serves a clearly distinct purpose: get_claim retrieves a single record by ID, list_measured provides an overview of measured combinations, and search_claims performs keyword queries. There is no overlap in functionality or potential for misselection.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (get_claim, list_measured, search_claims), which is predictable and easy to remember. No mixed conventions or ambiguous verbs.

Tool Count5/5

With 3 tools, the set is compact yet sufficient for the server's purpose of querying a claim registry. Each tool earns its place, covering retrieval, overview, and search without unnecessary bloat.

Completeness5/5

For a read-only registry, the surface is complete: users can search, list overall combinations, and fetch full details by ID. There are no missing lifecycle operations (create/update/delete) that would apply to this domain, so no dead ends remain.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that enables LLMs to pull-based search through Clawket's RAG repository for exploratory and conditional queries. It provides read-only access to search artifacts, tasks, and decisions via HTTP API.
    5
    10 npm
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Unified MCP server for managing local model runtimes (Ollama, LM Studio, etc.), enabling provider-agnostic discovery, lifecycle management, hardware-fit checks, and delegated inference.
    16
    19 npm
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International