Skip to main content
Glama

@altronis/tokenmark-mcp

An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.

Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.

Add to Claude Code

claude mcp add tokenmark -- npx -y @altronis/tokenmark-mcp

Related MCP server: LLM Benchmark MCP Server

Add to Cline

Cline CLI:

cline mcp add tokenmark --yes -- npx -y @altronis/tokenmark-mcp

Cline in VS Code: add the tokenmark entry from the JSON below to cline_mcp_settings.json. Step-by-step notes for agents are in llms-install.md.

Add to any MCP client

Run the server over stdio:

npx -y @altronis/tokenmark-mcp

Or in a client config:

{
  "mcpServers": {
    "tokenmark": {
      "command": "npx",
      "args": ["-y", "@altronis/tokenmark-mcp"]
    }
  }
}

Tools

  • tokenmark_recommend: { hardware, tasks?, prefer?, limit? } → ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.

  • tokenmark_configs: { model?, hardware?, limit? } → tracked benchmark configs, fastest first (25 by default, 50 max; matched gives the full count), with the run mode (speculative: true for MTP/DFlash/draft-model runs).

  • tokenmark_hardware: { platform? } → the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.

  • tokenmark_search: { term } → matching models/configs.

  • tokenmark_submit: { repo, source?, note? } → queues a GitHub/Hugging Face repo with benchmark numbers for human review.

  • tokenmark_submission_status: { id } → where a submission is.

A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.

Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.

The same data in your terminal

The CLI lives in cli/ of this repo:

npx @altronis/tokenmark-cli recommend --hardware "Strix Halo" --tasks coding

The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.

MIT licensed. Data aggregated from public community benchmarks with attribution.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Exposes queryable GPU inference benchmark data (quantization, throughput, VRAM, concurrent users) as tools for LLM clients.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.
    2
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.
    6 npm
    8
    Apache 2.0