BenchGecko
Server Details
AI model prices per provider, benchmarks, own AI behavior tests and AI economy data, with sources.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- BenchGecko/datasets
- GitHub Stars
- 0
TDQS
Scored across 9 tools
Both search and search_models can be used to find models, and fetch overlaps with get_model/get_gecko_test/get_scorecard by retrieving similar records. Descriptions clarify intended use (ids vs slugs, full record vs core facts), but an agent could still misselect.
Names are consistently snake_case and mostly follow verb_noun, but fetch and search are bare verbs, and cheapest_provider/latest_findings are noun phrases. Still readable and predictable overall.
9 tools is well-scoped for a model benchmark and pricing catalog. Each tool has a plausible role, and the set is neither thin nor bloated.
Covers search, model details, provider pricing, comparisons, Gecko test results, scorecard, and latest findings. Minor gap: no explicit list-all-models or filter-by-price/provider query, but core read-only workflows are covered.
Available Tools
9 toolscheapest_providerCheapest provider for a modelARead-onlyIdempotentInspect
Every provider serving a model, sorted by input price per 1M tokens (refreshed daily from provider endpoints), with output price, quantization and 1 day uptime.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model slug or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so safety is covered. The description adds genuinely useful behavioral context beyond that: the data is refreshed daily from provider endpoints, and the response includes output price, quantization and 1-day uptime. Freshness cadence and return composition are exactly the kind of detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with zero filler, front-loading the resource and sort order before the secondary return fields. Every clause earns its place and nothing is repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description does the heavy lifting by enumerating the returned fields (input/output price, quantization, uptime) and the refresh cadence, which compensates for the absent output schema. It is essentially complete, with only the when-to-use gap remaining.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage ('Model slug or name'), so the schema carries the semantics. The description mentions 'a model' but adds no format, matching, or validation guidance beyond the schema, which is the expected baseline here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (all providers serving a model) and the ordering criterion (input price per 1M tokens), which is specific enough to act on. It reads as a full listing sorted by price rather than a single 'cheapest' pick, a mild mismatch with the name/title, but the scope is unambiguous. It does not differentiate itself from siblings like compare_models or get_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the alternatives (compare_models, get_model, search_models), nor any prerequisite or exclusion. The agent must infer that this is the price-comparison lookup from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_modelsCompare two modelsARead-onlyIdempotentInspect
Side by side: BenchGecko score, list price, cheapest provider, context window, shared benchmark scores and Gecko Tests grades for two models.
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | First model slug or name | |
| b | Yes | Second model slug or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds value beyond that by enumerating the compared dimensions, telling the agent what content comes back from a tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the compared fields are listed inline and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return content, and it does so by naming the compared fields. It stops short of explaining structure or whether missing benchmark overlap is handled, which is a minor gap for a simple read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters documented as 'First/Second model slug or name'. The description adds no format, slug-vs-name resolution behavior, or ambiguity handling beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation (side-by-side comparison) on a specific resource (two models) and enumerates exactly what is compared: BenchGecko score, list price, cheapest provider, context window, shared benchmark scores and Gecko Tests grades. This distinguishes it clearly from get_model (single model) and cheapest_provider (single dimension).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied by the shape of the tool ('for two models') but there is no explicit when-to-use guidance and no mention of alternatives such as calling get_model twice or using get_scorecard. An agent can infer the intent, but nothing routes it deliberately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchFetch a BenchGecko recordARead-onlyIdempotentInspect
Fetch the full record for an id returned by search (model:, gecko-test: or scorecard).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds the useful fact that the id prefix determines which record type is returned, but says nothing about what happens on an unknown or malformed id, or about the size/shape of the returned record.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and resource and packs the id-format detail into a parenthetical. No filler, nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool whose annotations already establish safety, the definition covers the input format and the general return ('full record'). With no output schema the agent still cannot know the record's fields, but nothing essential for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the bare 'id' string parameter carries no meaning on its own; the description compensates by documenting the accepted id forms (model:<slug>, gecko-test:<slug>, scorecard). That is real semantic value beyond the schema, though it omits format constraints for each prefix.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('the full record') and clarifies it operates on an id produced by search, covering model, gecko-test and scorecard records. It implicitly spans the typed siblings get_model/get_gecko_test/get_scorecard but never names them, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description ties usage to ids 'returned by search', which implies the workflow (search first, then fetch), but it never says when to prefer this generic fetch over the typed siblings get_model, get_gecko_test or get_scorecard. Usage is implied, not prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gecko_testGecko Test resultsARead-onlyIdempotentInspect
Results of one of BenchGecko's own tests (own measurements, CC BY 4.0): who-are-you (Does the model know which lab made it?) world-map (How well does the model draw the world map from memory?) censorship-index (How often does the model refuse legitimate questions?) knowledge-horizon (Where does the model's knowledge of world events actually stop?) tokenizer-tax (How many more tokens does the same text cost outside English?) same-model-different-host (Do providers serving the same open model give the same quality?) model-drift-index (Do models quietly change behind the same name?)
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | ||
| limit | No | Rows in the text summary (structured result has all rows) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds provenance ('own measurements, CC BY 4.0'), which annotations do not carry, but says nothing about result shapes, freshness or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first clause, and every parenthetical gloss earns its place by disambiguating an enum value. It is a long run-on sentence crammed with parentheticals, which hurts scannability slightly, but no content is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should carry more of the return-value burden; it never says what a result record contains beyond the hint that there is a text summary and a structured result. For a read-only data-retrieval tool this leaves a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the 'limit' parameter is documented in the schema, but the required 'test' enum has no per-value descriptions there. The description fills that gap by explaining what each of the seven test names actually measures, which is meaningful semantics the schema alone lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (results/retrieval) and resource (one of BenchGecko's own tests), then enumerates every valid test value with a short gloss of what each measures. That routing detail lets an agent distinguish this from get_scorecard or compare_models without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'own tests' and the enumerated test names, but there is no explicit when-to-use vs. when-not, and no mention of the sibling tools (get_scorecard, compare_models, get_model) that might overlap for benchmark-style questions. An agent can infer the context but is not routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_modelGet model factsARead-onlyIdempotentInspect
Core facts for one AI model: BenchGecko score and rank, list price, price at every provider, benchmark scores with their original sources, Gecko Tests grades and as-of dates.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model slug (from search_models) or name, e.g. "claude-opus-5-5" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds real value beyond that by disclosing the contents of the response payload (per-provider pricing, benchmark sources, grades with as-of dates), which is important because no output schema exists. It omits any note on coverage gaps or stale data beyond the 'as-of dates' hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence whose clause is a tight enumeration of returned fields; nothing is padded or repeated. Slightly dense as one long clause chain, but every element earns its place by describing the payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read with no output schema, the field enumeration adequately tells the agent what it will receive and the 'as-of dates' phrase hints at data freshness. Missing only an explicit pointer to search_models for obtaining the slug (present in the schema) and any statement about what happens when the model is unknown.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter at 100% schema description coverage, the schema already documents the 'model' slug and gives a concrete example. The description adds no further parameter semantics (no casing rules, no aliasing behavior for names), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Core facts for one AI model') and then enumerates exactly what those facts are: score/rank, price, per-provider pricing, benchmark scores with sources, and Gecko Test grades. The 'one AI model' scoping makes it distinguishable from bulk siblings like search_models and compare_models, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the singular 'one AI model' suggests single-entity lookup versus compare_models, and the schema notes the slug comes from search_models. No explicit when-to-use, when-not-to-use, or named alternative appears in the description itself, leaving the agent to infer routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scorecardGecko ScorecardBRead-onlyIdempotentInspect
BenchGecko Gecko Scorecard: every model across BenchGecko's own profile tests with a letter grade per test and an overall Gecko Score (behavior, not intelligence).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description adds a genuinely useful semantic note that the score measures behavior rather than intelligence, but says nothing about result volume or how the 'every model' claim interacts with a limiting parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that identifies the tool and its return contents with no filler. The parenthetical '(behavior, not intelligence)' earns its place as disambiguation, though the sentence is dense enough that a second short clause about scope would have helped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does describe the shape at a high level (per-test letter grades plus an overall Gecko Score). It stops short of covering the pagination/limit behavior implied by the schema's default of 15 against the claim of 'every model'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single `limit` parameter, and the description never mentions it, so an agent gets no indication of whether it caps models, tests, or rows, nor of its default of 15 versus the maximum of 100. One parameter with no textual compensation falls below the baseline for documented params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (Gecko Scorecard), its coverage (every model across BenchGecko's profile tests), and its contents (letter grade per test plus an overall Gecko Score). This implicitly distinguishes it from get_gecko_test (single test) and get_model, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the aggregate-scorecard view, in contrast to get_gecko_test or compare_models. There is no explicit 'use this when' statement, no exclusion, and no mention of when a per-model or per-test tool would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
latest_findingsLatest findingsCRead-onlyIdempotentInspect
Latest notable results detected in the Gecko Tests and price data (fixed rules, numbers and quotes from stored results).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so safety is covered. The description adds genuinely new context by clarifying the output is derived from stored results via fixed rules rather than live computation, but it says nothing about ordering, freshness window, or result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler, and the core subject ('latest notable results') is front-loaded. The parenthetical about fixed rules is slightly cryptic but does carry meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameter documentation, the description is the only source of truth, yet it never explains what a returned 'finding' contains or what 'notable' means. For a tool whose output is opaque, that is a substantial gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'limit' parameter, and the description never mentions it, its default of 5, or its 1-20 range. With one undocumented parameter the description had room to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a resource (latest notable results in Gecko Tests and price data) but supplies no clear verb and never defines what makes a result 'notable'. It gives domain context yet does not distinguish this tool from siblings such as get_scorecard or search, so an agent cannot confidently route to it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this tool, when not to, or which sibling to prefer (e.g. search vs get_scorecard). The agent is left to infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchSearch BenchGeckoBRead-onlyIdempotentInspect
Search BenchGecko models and Gecko Tests. Returns ids for fetch.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, idempotent=true, destructive=false and a closed world, so the safety profile is covered. The description's one addition is that results are ids intended for a follow-up fetch call, which is useful since no output schema exists, but nothing is said about pagination, result caps, or match semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the scope front-loaded ahead of the return-value hint; nothing is redundant. It is arguably too terse given the unspecified query parameter, but deducting for that would double-count the parameter gap.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read search, the description covers what it searches and what it returns (ids for fetch), which is the minimum viable set. It omits result-set behavior (limits, ordering, pagination) that an agent would want when deciding how to consume or repeat the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole required parameter 'query' is not described at all in the description. For a search tool, the query syntax and matching behavior (name vs id vs free text) are the most important semantics, and the description compensates for none of it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and the two resources it covers (BenchGecko models and Gecko Tests), plus the return type (ids). That distinguishes it implicitly from narrower siblings like search_models and get_model, but no sibling is named, so the boundary is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Returns ids for fetch" implies the search-then-fetch workflow, which is genuine usage guidance. However there is no when-to-use/when-not guidance and no explicit mention of alternatives such as search_models or compare_models for the agent to choose between.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_modelsSearch AI modelsARead-onlyIdempotentInspect
Find AI models in the BenchGecko catalog by name, slug or lab. Returns slugs to use with get_model, cheapest_provider and compare_models, with BenchGecko score and list price.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | Model name, slug or lab, e.g. "claude opus", "gpt-5", "deepseek" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds genuine behavioral value by disclosing the return payload shape (slugs, BenchGecko score, list price), which the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The purpose is front-loaded and the return-value information follows immediately, so every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the burden of describing the return contents (slugs, score, price). It is nearly complete for a search tool, but the unexplained 'limit' parameter and absent guidance on result volume leave a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% - 'query' is documented in the schema, but 'limit' has no description anywhere. The description's 'by name, slug or lab' reinforces the query semantics but adds no syntax or default/range info for limit, so it does not fully compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) and resource (AI models) with scope (BenchGecko catalog) and the searchable fields (name, slug or lab). It also distinguishes itself from the generic 'search' sibling by being model-specific and names the downstream tools it feeds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly signals the workflow context - results return slugs to use with get_model, cheapest_provider and compare_models - which tells the agent when this tool is the right entry point. It stops short of stating when not to use it or how it differs from the sibling 'search' tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
- First observed
cheapest_provider - First observed
compare_models - First observed
fetch - First observed
get_gecko_test - First observed
get_model - First observed
get_scorecard - First observed
latest_findings - First observed
search - First observed
search_models
Related MCP Connectors
Source-backed AI model pricing, rankings, history, and benchmark data.
Compare LLM API prices, search models and providers, and access reviewed benchmark results.
Sourced AI-model pricing and capability data — compare and route to the cheapest capable model.
LLM API prices across 70+ providers: cheapest offer, comparisons, history and cost estimates.
Related MCP Servers
AlicenseNot gradedqualityAmaintenanceLive, sourced pricing and benchmark data across image, language, video, and audio AI models, seven tools for comparison, competitor lookups, and cost-aware routing. Also powers the free Modelglass VS Code extension.1MIT- AlicenseAqualityBmaintenanceGlobal price benchmarking for AI inference across 2,600+ SKUs from 47 vendors. Query live pricing, market indexes, and model specs via 8 tools. Free tier available.855 npmMIT
- AlicenseAqualityAmaintenanceLive LLM API pricing: current token prices, model comparisons, cheapest-model lookups, and The LLM Price Index for 150+ models across 20+ providers, re-verified daily. No API key required.41MIT
- AlicenseAqualityDmaintenanceGive your AI assistant real-time LLM/VLM knowledge. Pricing, benchmarks, and recommendations — updated every hour, not every training cycle.496 npm2MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.