TokenMark MCP
Allows queuing a GitHub repository with benchmark numbers for human review in TokenMark.
Allows queuing a Hugging Face repository with benchmark numbers for human review in TokenMark.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@TokenMark MCPwhat should I run on a Strix Halo for coding?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@altronis/tokenmark-mcp
An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.
Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.
Add to Claude Code
claude mcp add tokenmark -- npx -y @altronis/tokenmark-mcpRelated MCP server: LLM Benchmark MCP Server
Add to Cline
Cline CLI:
cline mcp add tokenmark --yes -- npx -y @altronis/tokenmark-mcpCline in VS Code: add the tokenmark entry from the JSON below to cline_mcp_settings.json. Step-by-step notes for agents are in llms-install.md.
Add to any MCP client
Run the server over stdio:
npx -y @altronis/tokenmark-mcpOr in a client config:
{
"mcpServers": {
"tokenmark": {
"command": "npx",
"args": ["-y", "@altronis/tokenmark-mcp"]
}
}
}Tools
tokenmark_recommend:{ hardware, tasks?, prefer?, limit? }→ ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.tokenmark_configs:{ model?, hardware?, limit? }→ tracked benchmark configs, fastest first (25 by default, 50 max;matchedgives the full count), with the run mode (speculative: truefor MTP/DFlash/draft-model runs).tokenmark_hardware:{ platform? }→ the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.tokenmark_search:{ term }→ matching models/configs.tokenmark_submit:{ repo, source?, note? }→ queues a GitHub/Hugging Face repo with benchmark numbers for human review.tokenmark_submission_status:{ id }→ where a submission is.
A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.
Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.
The same data in your terminal
The CLI lives in cli/ of this repo:
npx @altronis/tokenmark-cli recommend --hardware "Strix Halo" --tasks codingThe files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.
MIT licensed. Data aggregated from public community benchmarks with attribution.
This server cannot be deployed
Maintenance
Related MCP Connectors
Read-only: which local LLMs a GPU or Mac can run - VRAM fit, tokens/sec, model specs.
Compare LLM API prices, search models and providers, and access reviewed benchmark results.
GPU and LLM inference benchmarks, hardware evidence, deployment recommendations, and launch configs.
Source-backed AI model pricing, rankings, history, and benchmark data.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceExposes queryable GPU inference benchmark data (quantization, throughput, VRAM, concurrent users) as tools for LLM clients.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to compare LLM benchmarks, get pricing, and receive model recommendations for tasks and budgets.1MIT
- AlicenseNot gradedqualityAmaintenanceInferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.2MIT
- AlicenseNot gradedqualityCmaintenanceEnables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.6 npm8Apache 2.0