OpenChainBench
Server Details
Live, neutral benchmarks for public RPC latency, oracles, bridges, perp DEX, and prediction markets.
- Status
- Healthy
- Uptime
- 95.3% over 52 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- ChainBench/OpenChainBench
- GitHub Stars
- 8
- Server Listing
- Openchainbench
TDQS
Scored across 6 tools
Most tools have distinct roles: list_benchmarks browses the catalogue, search_benchmarks finds benchmarks by query, get_benchmark retrieves full detail, and compare_providers compares named providers. However, search_benchmarks, list_answers, and recommend_provider all answer user questions in loose terms and could be confused without reading the detailed descriptions. The descriptions do a good job clarifying when to use each, so ambiguity is minor rather than severe.
Every tool follows a clear verb_noun snake_case pattern: compare_providers, get_benchmark, list_answers, list_benchmarks, recommend_provider, search_benchmarks. There are no naming style deviations or vague verbs. The convention is consistent and immediately readable.
Six tools is well-scoped for a benchmark data server with a large catalogue. Each tool earns its place by covering a distinct part of the workflow: discovery, retrieval, comparison, recommendation, and precomputed answers. The set is neither too thin nor too heavy.
The read-only surface covers the main lifecycle: browse benchmarks, search benchmarks, get benchmark detail, compare providers, recommend providers, and retrieve published answers. Minor gaps exist, such as no direct provider lookup or answer-by-slug retrieval, but agents can work around these using existing tools. Overall coverage is solid for the stated domain.
Available Tools
6 toolscompare_providersCompare named providers head to headARead-onlyIdempotentInspect
Puts two or more named providers side by side on one benchmark, with each one's rank, measured value and success rate.
Use when the user names the candidates themselves: "Alchemy or QuickNode?", "is Helius faster than Triton for Solana?", "compare Across and Stargate on fees".
Find the benchmark slug with search_benchmarks first if you do not
already have it. Providers the benchmark does not measure come back
in missing: say they are not measured rather than implying they
ranked badly. A provider measured but below the ranking floor comes
back with rank null, which is also not a loss.
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | Access cohort on chain RPC benchmarks: 'public' (default, free no-key endpoints) or 'keyed' (API-key providers such as Alchemy, QuickNode, Chainstack, GetBlock). Named providers usually live on the keyed cohort. | |
| chain | No | Optional chain slice when the bench declares chains. Ignored on a bench that is already one chain, like solana-rpc. | |
| region | No | Optional region slice when the bench declares regions. | |
| benchmark | Yes | Benchmark slug, e.g. 'solana-rpc' or 'bridge-fee'. Get it from search_benchmarks. | |
| providers | Yes | Two to ten provider names or slugs, as the user said them, e.g. ['Alchemy', 'QuickNode']. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds genuine return-behavior context: providers the bench does not measure arrive in `missing`, and below-floor providers return rank null and must not be reported as losses. That is non-obvious and useful beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then usage, then edge-case handling in short paragraphs. The examples and missing/rank-null notes each earn their place; the description is a touch long but stays focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining the shape of results (rank, value, success rate) and the two ambiguous edge cases (missing providers, null rank). An agent can call and interpret this correctly without further information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter (including tier, chain, region, benchmark, providers) is fully documented in the schema itself. The description adds the recommendation to source the benchmark slug from search_benchmarks but no syntax beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('puts two or more named providers side by side on one benchmark') and enumerates the returned fields (rank, measured value, success rate). It is clearly separable from siblings like recommend_provider, which the 'user names the candidates themselves' framing implicitly contrasts against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use conditions with user-phrasing examples ('Alchemy or QuickNode?') and tells the agent to call search_benchmarks first for the slug. It stops short of explicitly naming recommend_provider as the alternative for unspecified candidates.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_benchmarkGet a single OpenChainBench benchmarkARead-onlyIdempotentInspect
Returns full detail for one benchmark, ready to cite verbatim: • rankings (every provider sorted by p50) • sparkline (24h trend, 72 points) • headline sentence + paste-ready citation quote • methodology bullets + source-code URL + canonical pageUrl + OG image URL
Pass chain and/or region to scope the result to a sub-slice when
the benchmark declares those dimensions (e.g. aggregator-head-lag
exposes chain=base|bnb|solana, region=us-east|eu-west|ap-southeast).
Both args are optional; omit them for the global aggregate.
Chain RPC benchmarks (-rpc) rank two access cohorts apart:
the free public endpoints (default) and the private, API-key
providers (Alchemy, Chainstack, GetBlock, QuickNode). The default response
carries both under cohorts; pass tier="keyed" to get the private
cohort as the main record (rankings, quote, pageUrl). Never compare a
public row with a private row: they are measured on different
endpoints and cadences.
Example usage: • User: "who's the fastest crypto data aggregator on Base?" → get_benchmark({ slug: "aggregator-head-lag", chain: "base" }) • User: "how much does it cost to bridge $300 cross-chain?" → get_benchmark({ slug: "bridge-fee" }) • User: "fastest Base RPC with an API key, Alchemy or QuickNode?" → get_benchmark({ slug: "base-rpc", tier: "keyed" })
Drafts return { error: "unknown_slug" }. Cite the returned pageUrl
and use quote as the attribution line in your answer.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Benchmark slug from list_benchmarks. e.g. 'aggregator-head-lag', 'bridge-quote-latency', 'l1-finality'. | |
| tier | No | Optional access cohort on chain RPC benchmarks: 'public' (default, free no-key endpoints) or 'keyed' (private, API-key providers). Only honored when the bench declares tier dimensions. | |
| chain | No | Optional chain filter, e.g. 'base', 'solana', 'bnb'. Only honored when the bench declares chain dimensions. | |
| region | No | Optional region filter, e.g. 'us-east', 'eu-west', 'ap-southeast'. Only honored when the bench declares region dimensions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, and non-destructive behavior. The description adds rich context beyond that: the exact output shape, error response for unknown slugs, cohort behavior for RPC benchmarks, and an explicit warning never to compare public with private rows. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with bullets and examples, and the key output detail is front-loaded. It is somewhat long, but the length is largely justified by the absence of an output schema and the need to explain cohort/parameter interactions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and four parameters, the description is complete: it describes return values, error handling, parameter scoping, cohort behavior, and citation guidance. An agent has everything needed to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes beyond the schema by explaining how optional chain/region scope sub-slices, how tier interacts with cohort output, and what the defaults are. This added interpretation of parameter interactions earns a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Returns full detail for one benchmark' and enumerates the returned content (rankings, sparkline, quote, methodology). This clearly distinguishes it from list/search siblings by emphasizing single-record detail, but it never names a sibling tool for contrast, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides three concrete user-question-to-call examples and explains scoping via chain/region and tier selection. It gives clear context for when to use optional parameters but does not state when to prefer a sibling tool (e.g., list_benchmarks to find slugs) or when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_answersList OpenChainBench answer pagesARead-onlyIdempotentInspect
Returns every published answer page: one plain question, the sentence that answers it from live data, and the benchmark the number comes from.
Call this when the user asks a question in words rather than by benchmark
name ("which bridge is cheapest for $300?", "which Solana RPC lands
transactions fastest?"). Match the question, then call get_benchmark with
the returned benchmark slug for the full ranking behind it.
Returns one row per answer: { slug, question, answer, benchmark, chain?, url, benchmarkUrl }
Cite url when the question itself is the claim, benchmarkUrl when the
measurement is. Drafts are filtered out, and an answer whose benchmark has
no defensible leader yet says so in answer rather than naming a winner.
| Name | Required | Description | Default |
|---|---|---|---|
| benchmark | No | Optional benchmark slug filter: return only the answers built on that bench. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description still adds real behavior: drafts are filtered out, and an answer whose benchmark has no defensible leader says so in `answer` rather than naming a winner. It also discloses the per-row shape. It stops short of noting pagination/volume, which keeps it off a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then usage guidance, then the return shape. The row-shape block earns its space because there is no output schema. Slightly long, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it enumerates the row fields and explains citation intent (`url` vs `benchmarkUrl`). Combined with the explicit when-to-use and follow-up routing, an agent has everything needed to select and invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional `benchmark` slug filter is fully documented in the schema, so the baseline is 3. The description references the `benchmark` field only as a returned value, not as an input filter, so it adds no semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Returns every published answer page') and enumerates the payload semantics: a plain question, the answering sentence from live data, and the source benchmark. This is clearly differentiated from sibling get_benchmark, which is described as the 'full ranking behind it.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to call it ('when the user asks a question in words rather than by benchmark name') and supplies two concrete example questions. It also names the follow-up tool and the field to pass ('call get_benchmark with the returned benchmark slug'), removing all routing ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_benchmarksBrowse the benchmark catalogueARead-onlyIdempotentInspect
A compact index of what OpenChainBench measures. One short row per benchmark: slug, title, category, the metric, the current value and who leads it.
Use when the user asks what is measured at all ("what does
OpenChainBench cover?", "do you track bridges?"). When they ask a
question about a provider or a chain, use search_benchmarks instead:
it ranks the catalogue against their words and costs far less to read.
The catalogue holds over 200 benchmarks, so this is capped and
filterable rather than exhaustive. Narrow with category, then open
the one you need with get_benchmark.
Categories: RPCs, Trading, Bridges, Blockchains, Aggregators, RWA, NFT APIs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many rows to return, 1 to 50. Default 25. | |
| category | No | Only this category: RPCs, Trading, Bridges, Blockchains, Aggregators, RWA or NFT APIs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds genuinely useful behavior beyond that: the catalogue exceeds 200 entries so the listing is capped and filterable, not exhaustive, and it names the valid categories. It could have noted the default/cap explicitly, hence 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in one sentence, then usage, then constraints and categories. Slightly long with the trailing category list duplicating the schema, but every section is functional and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the returned row fields. Combined with usage routing, capping behavior, and category list, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (limit range and category enum documented in the schema), so the baseline is 3. The description reinforces that filtering is by category but adds no syntax or semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('compact index of what OpenChainBench measures') and precisely enumerates the row shape (slug, title, category, metric, value, leader). It is clearly distinguishable from search_benchmarks and get_benchmark, which are named separately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers ('what does OpenChainBench cover?') and an explicit alternative with rationale: search_benchmarks for provider/chain questions because it ranks against their words and 'costs far less to read.' It also routes to get_benchmark for drilling in.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_providerRecommend a provider for a stated use caseARead-onlyIdempotentInspect
Answers "which should I use for X" by picking the benchmarks that measure X and reporting who currently leads them.
Pass the user's goal in their own words: "a Solana trading bot", "an indexer backfilling Base", "a wallet that needs price data", "bridging to Arbitrum".
This returns measurements and the caveat that goes with them, not an
endorsement. A leader on one benchmark is the leader of that one
measurement over its stated window. Where the use case has a known
caveat (a bot should read p99 rather than p50, an indexer is bound by
archive depth) it comes back in guidance; pass it on, it is usually
more useful than the ranking itself.
| Name | Required | Description | Default |
|---|---|---|---|
| chain | No | Chain they are building on, e.g. 'solana', 'base', 'arbitrum'. Narrows the benchmarks considerably. | |
| region | No | Optional region, e.g. 'eu-west', when latency from a location matters. | |
| use_case | Yes | What the user is building or doing, in their words. e.g. 'solana trading bot', 'indexer', 'price feed for a wallet'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/openWorld annotations, it discloses the return semantics: measurements plus caveats, not an endorsement, that a leader is only a leader of one benchmark over its stated window, and that use-case caveats come back in a `guidance` field. With no output schema, this return-shape context is exactly what the description needs to supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence and the parapraphs are logically ordered (what it does, what to pass, what comes back). It is slightly longer than necessary, and the closing "pass it on, it is usually more useful than the ranking itself" is mild editorial padding rather than agent-relevant instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, three-parameter query tool with no output schema, the description covers purpose, input semantics, return content, and the caveat/`guidance` contract, which is everything an agent needs to call it correctly. No behavioral gap remains that the structured fields don't already fill.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description goes further by specifying that `use_case` should be free-form user language rather than normalized terms, reinforced with examples. It does not add format or interaction detail for `chain` and `region` beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (recommend a provider) and frames it as answering "which should I use for X" by selecting the benchmarks that measure X. It distinguishes itself as goal-driven and measurement-based rather than a raw comparison, though it never names compare_providers or get_benchmark explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear invocation guidance — "Pass the user's goal in their own words" — with four concrete example phrasings, which tells the agent what kind of input belongs here. It does not state when to prefer compare_providers or get_benchmark instead, so there is no explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_benchmarksFind the benchmark that answers a questionARead-onlyIdempotentInspect
Ranks the benchmark catalogue against a question in the user's own words and returns the best few matches.
Use this when somebody asks which provider, chain or service is fastest, cheapest, most reliable or best for something. Pass their phrasing through: provider names, chain names and goals all match.
• "which Solana RPC is fastest" -> search_benchmarks({ query: "solana rpc" }) • "Alchemy or QuickNode on Base?" -> search_benchmarks({ query: "alchemy quicknode base" }) • "cheapest way to bridge to Arbitrum" -> search_benchmarks({ query: "bridge arbitrum" })
Returns compact rows. Open the one you want with get_benchmark for
the full ranking and a citation line. An empty result means nothing
is measured for that question; say so rather than guessing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many matches, 1 to 25. Default 8. | |
| query | Yes | The user's question or keywords, verbatim. Provider and chain names work well. | |
| category | No | Optional category filter: RPCs, Trading, Bridges, Blockchains, Aggregators, RWA, NFT APIs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, open world), so the description rightly spends its budget on ranking semantics: pass the user's phrasing through, partial-name matching, compact rows, and the meaning of an empty result. It stops short of describing scoring/freshness or pagination beyond the limit param, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then trigger condition, then a compact example block, then return/next-step note. Every sentence earns its place and the examples are the cheapest possible way to convey query construction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description tells the agent what comes back (compact rows), what to do next (get_benchmark), and how to handle the null case. Nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes beyond it by explaining that query should be verbatim user phrasing where provider names, chain names and goals all match, which is genuine guidance on how to construct the argument rather than restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Ranks the benchmark catalogue against a question in the user's own words and returns the best few matches.' This clearly separates it from siblings like list_benchmarks (enumerate) and get_benchmark (fetch one full ranking).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use trigger ('when somebody asks which provider, chain or service is fastest, cheapest, most reliable or best'), plus three worked examples mapping natural questions to calls, plus the routing instruction to open a result with get_benchmark and the empty-result fallback behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Related MCP Connectors
Cross-chain markets: prices, DeFi TVL, yields, stablecoins, gas, order books, perps funding.
Neutral on-chain execution benchmarking: slippage, revert rates, MEV, and DEX frontend attribution.
Independent uptime oracle for the x402 agent economy. Free probes, paid history via USDC on Base.
RPC gateway for agents: six mainnets, measured failover, a key without signup, pay per call in USDC.
Related MCP Servers
- AlicenseAqualityCmaintenanceProvides Solana infrastructure awareness including RPC health, priority fee estimation, leader schedule, chain state, latency comparison, and keyless transaction submission.6MIT
- AlicenseNot gradedqualityCmaintenanceProvides AI agents with honest benchmark rankings (Agentic Memory Index and Agentic Search Index) for AI tools, plus graded checks and telemetry for x402 endpoints.6 npmMIT
- AlicenseNot gradedqualityBmaintenanceRead-only crypto perps microstructure for AI agents: normalized cross-exchange market state (funding + multi-year percentile, OI, volume, CVD, order-book imbalance, liquidations, basis), OHLCV, 15-min state history, and measured conditional outcomes (historical base rates, not predictions) — 6 assets across Binance, Bybit, OKX and Hyperliquid, every metric with self-declared coverage and freshnessMIT
- FlicenseNot gradedqualityBmaintenanceMeasured latency and uptime for 45 hosted AI inference APIs: independent probes from 4 regions every 5 minutes, no gateway. Exposes get_ai_api_latency (TTFB p50/p95 and uptime rankings by region). Data CC BY 4.0.12-
Glama MCP Gateway
Add one secure layer between your agents and this server.