Skip to main content
Glama

Server Details

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
mauricekleine/nonobench
GitHub Stars
6

TDQS

A3.5/5.0

Scored across 11 tools

Disambiguation4/5

Each tool targets a distinct resource+action: puzzles, models, runs, families, providers, and cross-cutting result views. There is some mild overlap between get_model_results, get_leaderboard, and compare_models (all surface accuracy metrics) and between get_model_puzzles, get_puzzle_results, and list_runs, but descriptions make the boundaries clear enough to select correctly.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (get_*, list_*, compare_*, check_*). No mixed conventions or vague verbs; the prefix also telegraphs intent (get=single item, list=collection, compare=multi-entity).

Tool Count5/5

11 tools is well-scoped for a benchmark-results server and each earns its place across the puzzle, model, run, and provider dimensions. Nothing feels redundant or padded.

Completeness4/5

Good coverage of the benchmark lifecycle: enumerate puzzles/models/families/providers/runs, inspect per-model and per-puzzle results, compare, rank, and validate a solution. Minor gaps (e.g. no direct run-detail fetch or aggregate stats tool) are easily worked around via list_runs and the existing get_* tools.

Available Tools

11 tools
check_solutionCheck solutionA
Read-only
Inspect

Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesPuzzle id from list_puzzles
gridYesRow-major string of 0 (empty) and 1 (filled), width × height characters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
correctYes
rowViolationsYes
columnViolationsYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds meaningful context beyond that: it uses 'the same rule as the benchmark grader' and reports specific mismatched rows and columns, which helps set expectations for the check behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the tool's purpose and followed by the reporting behavior. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only check tool, the description covers purpose and key behavior, while the schema documents parameters and an output schema exists for return values. It is nearly complete, though usage context relative to sibling tools is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the two parameters (id and grid) are fully documented in the schema, including the grid's row-major 0/1 format. The description adds no parameter-level detail beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Check a grid against a puzzle's clues' and clarifies it uses the benchmark grader's rule. It does not explicitly distinguish itself from siblings such as get_puzzle_results, so it is clear but lacks sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the tool does but gives no guidance on when to use it versus alternatives like get_puzzle_results or get_puzzle. There are no prerequisites, exclusions, or alternative tools named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_modelsCompare modelsA
Read-only
Inspect

Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelsYes2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelsYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds which metric families are compared (accuracy, cost, latency, tokens; overall and per-size), which is useful, but since an output schema exists this is largely a preview of return contents rather than disclosure of behavior an agent couldn't infer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the comparison framing and then the dimensions compared. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, one-parameter comparison tool with a full output schema, the description covers what is compared closely enough for correct invocation. The remaining gap is routing guidance versus sibling tools such as get_model_results or get_leaderboard.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single 'models' array is fully documented, including the 2–20 bound and id examples. The description's 'model or family names' merely restates the schema, adding no new format or syntax guidance; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action (side-by-side comparison) plus the resource (models/families) and enumerates the compared dimensions: overall and per-size accuracy, cost, latency, and token results. An agent can distinguish it from single-model tools like get_model_results, though the description never explicitly names that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'side-by-side' and the plural 'names' imply the multi-model use case, and the schema's 2–20 min/max reinforces it, but there is no explicit 'use this when…' statement and no reference to get_model_results, get_leaderboard, or other alternatives for single-model or ranked lookups.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_leaderboardGet leaderboardB
Read-only
Inspect

Models ranked by accuracy; defaults to all effort levels for compatibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoGrid size to filter on
effortNobest, all (default), or one effort level; empty means all
familyNoComma-separated family ids; empty means no filter
versionNoComma-separated benchmark versions: 1.0, 1.1, 1.2; empty means all
providerNoComma-separated provider ids; empty means no filter
reasoningNoOnly reasoning (true) or non-reasoning (false) variants
min_correctNoMinimum puzzles solved in the selected tier; default 0 includes unsolved variants
open_weightsNoOnly open-weight (true) or closed (false) models

Output Schema

ParametersJSON Schema
NameRequiredDescription
sizeYes
modelsYes
updatedAtYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral fact — that omitted effort defaults to all levels for compatibility — but says nothing about ranking metric mechanics or tie-breaking. Adequate but thin against an annotation-covered read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight clause with the core ranking statement front-loaded and the default-behavior caveat trailing. No padding, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema and 100% parameter coverage, the description is not required to explain returns or params. However, for an eight-filter leaderboard tool it omits any hint of sort order, scope semantics, or how 'accuracy' is measured, leaving an agent to infer ranking details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all eight parameters, including the effort default and empty-means-all conventions, so the schema carries the load. The description's effort-default remark marginally reinforces the schema rather than adding new meaning. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and operation — models ranked by accuracy — so an agent knows it returns an ordered list of models. It does not distinguish itself from lookalike siblings such as get_model_results or compare_models, which also surface model performance data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use or when-not-to-use guidance and names no alternative among the ten siblings. The phrase about effort-level defaults is a behavioral note, not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_puzzlesGet model puzzlesB
Read-only
Inspect

Which puzzles one model solved, missed, timed out on, or has not run.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel variant id as listed on the leaderboard, e.g. claude-opus-5.5-high

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYes
solvedYes
puzzlesYes
attemptedYes
displayNameYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe read operation (readOnlyHint=true, openWorldHint=false), so the safety profile is covered. The description adds the useful behavioral detail that results are partitioned into solved/missed/timed-out/not-run categories, but says nothing about auth, scope, or cost; with an output schema present, disclosing return values is not required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence fragment with zero filler and the key content front-loaded. It is arguably too terse — a verbless clause rather than a call-to-action sentence — which slightly limits how self-explanatory it reads in a tool list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and the only parameter fully documented in the schema, the description carries little remaining burden. The one gap is routing: it never clarifies how this differs from get_model_results, which matters given how many retrieval siblings exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'model' parameter is clearly documented there with an example id format. The description adds no syntax or format detail beyond the schema, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (puzzles for one model) and enumerates the outcome buckets returned (solved, missed, timed out, not run), which is more specific than restating the title. It does not, however, distinguish this from the close sibling get_model_results, so an agent cannot tell the two apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_model_results or get_puzzle_results. The agent must infer from the name alone which of the sibling retrieval tools to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_resultsGet model resultsB
Read-only
Inspect

Accuracy, cost, latency and token use for one model, broken down by grid size.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel name as listed on the leaderboard, e.g. gpt-5.4-xhigh

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYesModel variant id, e.g. claude-opus-5.5-high
totalYes
bySizeYes
effortYes
familyYes
correctYes
accuracyYesPercentage of puzzles solved
providerYes
reasoningYes
failedRunsYes
displayNameYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the useful detail that results are segmented by grid size, but says nothing about aggregation, time window, or what happens for models without runs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word (metrics list, single-model scope, grid-size breakdown) carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated, and the parameter list is trivially small. The remaining gap is routing context: nothing tells the agent why to pick this over get_model_puzzles or compare_models.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'model' parameter, including a naming example ('gpt-5.4-xhigh'), so the schema does the heavy lifting. The description's phrase 'for one model' adds no format or constraint detail beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific resource (results for one model) and enumerates the metrics returned (accuracy, cost, latency, token use) with a scope qualifier ('broken down by grid size'). It does not explicitly distinguish itself from close siblings like get_model_puzzles or get_puzzle_results, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and names no alternatives. With siblings such as compare_models, get_leaderboard, and get_model_puzzles in the same family, the agent gets no signal on when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_puzzleGet puzzleA
Read-only
Inspect

One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesPuzzle id from list_puzzles
include_solutionNoInclude a reference solution

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
urlYes
sizeYes
indexYes
widthYes
heightYes
promptYesClue text as models received it
rowCluesYes
columnCluesYes
referenceSolutionNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine behavioral context beyond that: the clue text is included, the reference solution is opt-in, and some puzzles have several valid solutions — a useful caveat that affects downstream comparison/checking. It stops short of auth or rate-limit details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler. The core result (the puzzle and its clue text) is front-loaded, followed by the solution-conditional caveat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary, and the description covers the non-obvious piece an agent needs — the optional solution and the existence of multiple valid solutions. It could add a note about the id source or sibling routing, but is largely complete for a two-param read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents both parameters. The description nevertheless adds meaning: it explains that the solution is 'only included on request,' clarifying the conditional semantics of include_solution beyond its terse schema label.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('One puzzle') and clarifies it returns the clue text models were prompted with, which distinguishes it from the plural list_puzzles sibling. It does not, however, explicitly name any sibling or state the singular-by-id retrieval role to contrast with get_puzzle_results or get_model_puzzles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (fetch a single puzzle, pass an id from list_puzzles), and it notes the solution is only included on request, which guides use of include_solution. But there is no explicit when-to-use/when-not guidance and no named alternative for adjacent needs like results or model comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_puzzle_resultsGet puzzle resultsB
Read-only
Inspect

Per-model outcomes for one puzzle. Answers are omitted unless requested.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesPuzzle id from list_puzzles
effortNobest, all (default), or one effort level
familyNoComma-separated family ids; empty means no filter
providerNoComma-separated provider ids; empty means no filter
reasoningNoOnly reasoning (true) or non-reasoning (false) variants
open_weightsNoOnly open-weight (true) or closed (false) models
include_answersNoInclude each model's answer grid

Output Schema

ParametersJSON Schema
NameRequiredDescription
runsYes
sizeYes
indexYes
solvedYes
attemptsYes
puzzleIdYes
updatedAtYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description usefully discloses a non-obvious default (answer grids are omitted unless requested), but says nothing about pagination, result volume, or behavior when filters match nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no waste, with the resource scope front-loaded and the output-default caveat trailing. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the resource scope plus the answers-omission behavior. For a 7-param filtered query tool the definition is adequate, though the absence of any usage routing leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there are 7 well-documented parameters, so the schema carries the parameter burden. The description adds nothing about parameter syntax, defaults, or filter interaction beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Per-model outcomes for one puzzle" gives a specific verb+resource with a clear scope (one puzzle, all models), which implicitly separates it from get_puzzle (single puzzle metadata) and get_model_results (one model across puzzles). It never names a sibling, so differentiation is inferable rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no routing to alternatives such as get_puzzle or get_model_results. The scoping phrase hints at the context, but the agent must infer when this is the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_familiesList familiesC
Read-only
Inspect

Model families, available efforts and best variants.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
familiesYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, covering the safety profile. The description adds no behavioral details such as whether the list is complete, static, paginated, or otherwise constrained, and its content is largely restated by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The seven-word fragment is very concise and front-loads the returned content categories. It is not a complete sentence, but it wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter list tool with an output schema and read-only annotations, the description supplies the core resource and return categories. It still omits usage context relative to sibling list/get tools, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics for the description to clarify. The baseline for zero parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (model families) and return content (available efforts and best variants), which helps distinguish it from siblings like list_providers. However, it is a noun phrase with no explicit verb or scope statement, so the tool's action is only implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool versus list_providers, list_puzzles, or get_leaderboard. Usage is only inferable from the resource name, with no alternatives or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_providersList providersC
Read-only
Inspect

Provider ids, names, families and variant counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
providersYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds no behavioral context beyond the annotations: no statement about ordering, cost, or whether results are cached or exhaustive. What it does say (field names) duplicates the existing output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no filler, but it is telegraphic rather than front-loaded prose — a nominal phrase with no verb, which forces the reader to reconstruct intent. Brief but under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema present, the description needn't explain return values or arguments, so the scope of what's missing is narrow. The remaining gap is any routing guidance among the many sibling list_* and get_* tools, which goes unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric this is baseline 4. There is nothing parameter-related for the description to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Provider ids, names, families and variant counts' tells the agent the resource is providers and enumerates the fields returned, so the purpose is inferable but never stated as an action. It also does not distinguish itself from the sibling list_families, which is a plausible alternate target. Vague-but-decodable rather than a clear verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no condition selecting this tool over list_families or check_solution, and no prerequisites. The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_puzzlesList puzzlesC
Read-only
Inspect

The benchmark puzzles with their ids and row/column clues.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoGrid size to filter on

Output Schema

ParametersJSON Schema
NameRequiredDescription
puzzlesYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds only that ids and row/column clues are included, which overlaps with the existing output schema, and says nothing about volume, pagination, or result size. Little behavioral value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, front-loaded with the resource and its contents. It is efficient, though arguably too terse for a listing tool, bordering on under-specification rather than maximal information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need not be spelled out, and annotations cover safety. Still, for a listing tool the description omits anything about result volume, ordering, or how it relates to get_puzzle. Adequate but with clear gaps given the crowded sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one optional parameter (size) with 100% schema description coverage and an enum of grid sizes, so the schema fully documents it. The description does not mention filtering by size at all, adding no meaning beyond the schema. Per the baseline rule for high schema coverage, 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (benchmark puzzles) and the fields returned (ids and row/column clues), so an agent knows this is a listing of puzzle definitions. However, it is a noun phrase with no verb and no differentiation from siblings like get_puzzle or get_model_puzzles, which also concern puzzles. Purpose is decipherable but not sharply scoped.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus get_puzzle (single puzzle), get_model_puzzles, or get_puzzle_results. No prerequisites, no exclusions, no alternatives are mentioned. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsList runsB
Read-only
Inspect

Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoGrid size to filter on
limitNoDefault 100
modelNoOnly runs of this model variant id
offsetNoRuns to skip, for paging; default 0
puzzle_idNoOnly runs on this puzzle id from list_puzzles
include_outputNoInclude raw prompt and model output (large)

Output Schema

ParametersJSON Schema
NameRequiredDescription
runsYes
limitYes
totalYes
offsetYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds only that output can optionally include raw prompt/model output, which largely restates the schema's include_output note rather than disclosing new behavior like pagination defaults or payload size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words or filler. It is efficient, though it is arguably too thin to carry the tool's full scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering the safety profile, return values and safety need not be explained. However, a six-parameter listing tool whose description omits any filtering or sibling-routing guidance is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (size, limit, model, offset, puzzle_id, include_output) is documented in the schema itself. The description adds no syntax, format, or interaction detail beyond what the schema provides, so it meets the baseline only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource ('individual benchmark runs') and defines it precisely as 'one model on one puzzle', which is genuinely clarifying. It does not, however, distinguish this list tool from siblings like get_model_results or get_puzzle_results that may return overlapping data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no condition selecting this over the sibling result tools, and no mention of how filtering should be applied. The agent must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 11 tool updates
    • First observedcheck_solution
    • First observedcompare_models
    • First observedget_leaderboard
    • First observedget_model_puzzles
    • First observedget_model_results
    • First observedget_puzzle
    • First observedget_puzzle_results
    • First observedlist_families
    • First observedlist_providers
    • First observedlist_puzzles
    • First observedlist_runs

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server for puzzle generation: word search generator, crossword generator, and sudoku generator + solver, with printable PDF worksheets, themed word banks, and verifiable LLM evals. Works with Claude Desktop, Cursor, Windsurf, and any Model Context Protocol client. From the makers of puzzletide.com.
    51 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A large-scale benchmark that evaluates AI agents' tool-use competency across 36 real MCP servers using a reproducible Docker sandbox and LLM-as-judge scoring.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.