Skip to main content
Glama

BenchGecko

Gecko Test results

get_gecko_test
Read-onlyIdempotent

Results of one of BenchGecko's own tests (own measurements, CC BY 4.0): who-are-you (Does the model know which lab made it?) world-map (How well does the model draw the world map from memory?) censorship-index (How often does the model refuse legitimate questions?) knowledge-horizon (Where does the model's knowledge of world events actually stop?) tokenizer-tax (How many more tokens does the same text cost outside English?) same-model-different-host (Do providers serving the same open model give the same quality?) model-drift-index (Do models quietly change behind the same name?)

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
testYes
limitNoRows in the text summary (structured result has all rows)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds provenance ('own measurements, CC BY 4.0'), which annotations do not carry, but says nothing about result shapes, freshness or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose in the first clause, and every parenthetical gloss earns its place by disambiguating an enum value. It is a long run-on sentence crammed with parentheticals, which hurts scannability slightly, but no content is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should carry more of the return-value burden; it never says what a result record contains beyond the hint that there is a text summary and a structured result. For a read-only data-retrieval tool this leaves a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: the 'limit' parameter is documented in the schema, but the required 'test' enum has no per-value descriptions there. The description fills that gap by explaining what each of the seven test names actually measures, which is meaningful semantics the schema alone lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (results/retrieval) and resource (one of BenchGecko's own tests), then enumerates every valid test value with a short gloss of what each measures. That routing detail lets an agent distinguish this from get_scorecard or compare_models without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'own tests' and the enumerated test names, but there is no explicit when-to-use vs. when-not, and no mention of the sibling tools (get_scorecard, compare_models, get_model) that might overlap for benchmark-style questions. An agent can infer the context but is not routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.