Skip to main content
Glama

gambit


Testing License: MIT

gambit is a local-first chess analysis and tutoring server for the Model Context Protocol (MCP). It connects an AI tutor to Stockfish, opening knowledge, and persistent learning history.

It features:

  • Bounded Stockfish analysis with warm engine workers, cached results, candidate comparisons, and progressive feedback.

  • Whole-game analysis with persistent checkpoints, cancellation, and restart recovery.

  • Position-specific teaching sessions, graduated hints, game-derived lessons, and spaced review.

  • A bundled corpus of 3,864 named opening lines, study imports, transposition-aware repertoires, and source-attributed plans.

  • Tactical puzzles, pawn-structure analysis, endgame evaluation, and local Syzygy tablebase probing.

  • Structured speech plans, optional local speech synthesis, and optional screenshot recognition.

  • Local stdio and HTTP transports, persistent Linux services, and an optional outbound tunnel.

  • SQLite storage with explicit backup, restore, and retention commands.

It is authored by Mayaz Rakib.

Installation and Setup

Install Python 3.14 or newer, uv, and Stockfish 18 or a newer compatible UCI engine. From the package directory, install dependencies and initialize the bundled opening knowledge:

uv sync --extra vision
uv run gambit initialize
uv run gambit

Omit --extra vision if screenshot recognition is not needed. The distribution and Python import package are both named gambit; the source package is src/gambit.

The default transport is MCP stdio. Configure your MCP host to run uv --directory /absolute/path/to/gambit run gambit. Protocol messages use stdout. For the background service and tunnel, keep credentials in the ignored package-local .env, never in committed configuration or a parent dotenv.

Set STOCKFISH_PATH to override the executable. Set GAMBIT_DATABASE to select the local SQLite database. Set GAMBIT_CONFIG to a TOML configuration file. Configuration keys and defaults are defined in configuration.py.

Related MCP server: Chess Engine MCP Server

Background Services and Connections

Install the official OpenAI tunnel client. Copy .env.example to .env inside this package, and replace its placeholders with your runtime API key and tunnel ID. Existing environment variables take precedence when launching Gambit directly. The loader reads only this package's dotenv, not parent directories.

On Linux with systemd, run from this package after uv sync:

chmod 600 .env
uv run python scripts/install_services.py
loginctl enable-linger "$USER"
systemctl --user status gambit.service gambit_tunnel.service

The installer creates and enables two user services from the service templates. Gambit serves streamable HTTP at http://127.0.0.1:3025/mcp; the tunnel connects outbound to OpenAI. Both restart after failures. Linger keeps them running after logout and starts them when the systemd host boots. Closing a shell does not stop them. Shutting down Windows or WSL, suspending the host, or stopping the Linux machine still interrupts availability.

In ChatGPT developer mode, create an app from the Plugins page, select Tunnel as its connection, and select the tunnel matching CONTROL_PLANE_TUNNEL_ID. The runtime key must have Tunnels Read and Use permissions, and the tunnel must belong to the relevant workspace. The key authenticates the tunnel, not an LLM: the model selected in ChatGPT supplies the conversational tutoring. See the official connection instructions.

Check tunnel readiness, restart the services, or stop them with these commands:

curl --noproxy '*' --max-time 5 http://127.0.0.1:3026/readyz
systemctl --user restart gambit.service gambit_tunnel.service
systemctl --user stop gambit_tunnel.service gambit.service

Use disable --now instead of stop to also disable automatic startup. After changing .env, restart both services. The tunnel's local status interface is http://127.0.0.1:3026/ui. Keep both listeners on loopback; do not expose Gambit's unauthenticated HTTP server directly to the network. The tunnel can reconnect without restarting Gambit or interrupting its analysis jobs. Restarting Gambit itself interrupts unfinished jobs.

Architecture

The server exposes factual chess tools and structured teaching turns to an MCP host. The host's language model provides conversational explanation. No model key or network model call is required by Gambit. Engine evidence, corpus provenance, learning state, board directives, and speech segments remain separate:

  • engine.py owns warm Stockfish subprocesses and bounded searches.

  • analysis.py provides configurable worker pooling, duplicate-search coalescing, persistent analysis, and background game jobs.

  • games.py validates PGN and produces adaptive game analysis and review classifications.

  • knowledge.py imports PGN studies and opening TSV, matches transpositions, and probes local Syzygy files.

  • training.py manages learner preferences, source-attributed puzzles, spaced repetition, and lessons.

  • features.py computes geometric and pawn-structure facts.

  • tutoring.py assembles grounded teaching turns.

  • speech.py pronounces legal moves and optionally synthesizes local WAV audio with espeak-ng.

  • ocr.py recognizes axis-aligned board screenshots with the optional vision dependencies.

The sibling chessboard package can consume the FEN, legal UCI variations, and arrow directives. A client owns playback, animation timing, audio interruption, and stale-position rejection.

Configuration

Copy the example configuration to a local TOML file and set GAMBIT_CONFIG to its absolute path. Environment overrides include STOCKFISH_PATH for the executable and GAMBIT_DATABASE for storage.

Setting

Default

Purpose

worker_count

2

Number of warm engine processes.

threads_per_worker

1

Engine threads per process.

hash_mb_per_worker

64

Hash memory per process, in MiB.

cache_capacity

512

Maximum in-memory analysis cache entries.

instant_nodes

12000

Preliminary search budget.

quick_nodes

100000

Routine search budget.

deep_nodes

1000000

Deeper verification budget.

candidate_count

3

Candidate variations for quick and deep searches.

search_time_ms

5000

Per-search time ceiling, in milliseconds.

search_depth

64

Per-search depth ceiling, in plies.

max_jobs

16

Background job capacity.

is_reproducible

false

Clear engine hash before uncached work.

syzygy_path

Empty

Local tablebase directories.

speech_executable

espeak-ng

Optional local speech executable.

speech_directory

~/.cache/gambit/speech

Generated audio directory.

Analysis Budgets and Evidence

Profiles use node budgets, not latency guarantees: instant defaults to 12,000 nodes and one variation, quick to 100,000 nodes and three variations, and deep to 1,000,000 nodes and three variations. Customize instant_nodes, quick_nodes, deep_nodes, and candidate_count in the example configuration. Searches additionally stop at search_time_ms (5,000 by default) or search_depth (64 by default), whichever search limit is reached first. End-to-end latency includes queueing and depends on CPU, position, and concurrent work. Each engine operation has a 30-second response deadline.

The default pool has two workers, one thread per worker, and 64 MiB of hash per worker. Adjust worker_count, threads_per_worker, and hash_mb_per_worker together to avoid CPU oversubscription. is_reproducible resets search hash before uncached work. Exact results can still differ between engine versions and hardware. New analysis responses record the engine version and executable SHA-256.

New evaluation tools use White-relative scores. Legacy analyze_position and tutor_move preserve side-to-move scoring. Mate values are separate from centipawns. Review labels are Gambit policy, not native Stockfish labels or a rating estimate.

Supply initial FEN plus move history to evaluate_position when repetition history matters. A standalone FEN cannot establish previous repetitions. Opening transposition keys intentionally omit move counters; engine cache keys retain counters and supplied history.

Background game searches use at most one fewer worker than the pool size, reserving interactive capacity when there are at least two workers. A one-worker configuration cannot reserve separate interactive capacity. get_analysis_metrics reports a bounded window of queue and search latency samples, cache hits, and pending requests.

Game jobs checkpoint completed searches to SQLite. Checkpoint identity includes the root position, move history, search settings, engine binary digest, and configuration. Queued, running, and interrupted jobs with saved inputs resume at startup within the job limit. resume_analysis_job also resumes explicitly cancelled or failed jobs. Old jobs without saved inputs remain interrupted and require resubmission. Cancellation sends UCI stop to active searches and drains the engine response before reuse. Completed reports remain readable. Job status distinguishes completed and reused search counts; neither is an estimated completion percentage.

start_progressive_analysis returns preliminary evidence and a request ID for deeper background analysis. Reusing a context ID cancels the previous request. Poll get_progressive_analysis, then retrieve its analysis ID through get_analysis; use cancel_progressive_analysis when a board is closed. Clients must still reject mismatched context IDs and FENs. These short-lived requests do not resume after a process restart; whole-game jobs do.

assess_candidate_move compares an unrestricted root search with a candidate-restricted search at quick and deep budgets. It preserves supplied move history, reports mate transitions separately from centipawn loss, and includes legal lines and observed captures or checks. A serious-mistake label requires agreeing classifications and root best moves, a larger configured budget, and increased measured search work. Disagreement returns uncertain. Search agreement is not calibrated confidence or proof of optimality. evaluate_move preserves its previous fields and adds this assessment. compare_candidate_moves and create_tutor_turn accept history; the tutor's compare intent now returns candidate assessments.

Public Tools

The MCP host discovers input schemas from the server. Tools return engine evidence and stored records; the host model supplies conversational explanations.

Area

Tools

Positions

inspect_position, list_legal_moves, analyze_position, evaluate_position, analyze_positions, explain_position, analyze_pawn_structure

Candidate moves

tutor_move, evaluate_move, compare_candidate_moves, assess_candidate_move

Games and jobs

parse_game, analyze_game, review_game, start_analysis_job, get_analysis_job, cancel_analysis_job, resume_analysis_job, get_analysis

Progressive analysis

start_progressive_analysis, get_progressive_analysis, cancel_progressive_analysis, get_analysis_metrics

Opening knowledge

initialize_openings, import_study, import_opening_tsv, find_opening, get_opening_position, import_opening_plan, get_opening_plans, import_opening_statistics, get_opening_coverage

Repertoires

get_repertoire_move, create_repertoire_drill, get_due_repertoire_drill, submit_repertoire_move

Endgames

get_tablebase_result, evaluate_endgame, get_tablebase_diagnostics

Learners and puzzles

update_learner_profile, get_learner_progress, import_puzzles, get_tactic, submit_tactic_move, record_training_attempt

Lessons and teaching

start_lesson, advance_lesson, create_tutor_turn, start_teaching_session, get_teaching_hint, submit_teaching_move, get_concept_progress, create_game_lessons

Speech and vision

get_speech_plan, synthesize_speech, recognize_chessboard, analyze_chessboard

Discovery

get_capabilities

The server also publishes a teaching prompt and an analysis resource. Tool contracts are implemented in main.py, tools.py, and extended_tools.py.

Knowledge and Training

Call initialize_openings once to load the bundled 3,864 named Lichess opening lines. The corpus is pinned to revision 5a13018164f6bd88f48b3dc31a8e2a39f31a060a and includes CC0 attribution. Import one PGN study with all its variations through import_study. Use the repertoire flag for personal preparation. import_opening_tsv accepts the Lichess opening format with eco, name, and pgn columns. Sources are required and retained. Imports are local and do not silently download datasets.

import_puzzles accepts bounded batches of the Lichess puzzle CSV schema. Its first move is the opponent's setup move. Subsequent moves form the source solution. Legality is checked, but importing does not independently prove optimality. An alternative move is described as not matching the source, not automatically as a chess mistake.

No corpus can promise every known or studied line. Import versioned sources and personal studies, then use engine analysis beyond their coverage. Opening matches are evidence of corpus membership, not proof that a move is best.

Configure syzygy_path for local tablebases. Multiple directories use the platform path separator, a colon on Linux. Paths are expanded consistently for Gambit and Stockfish. Gambit's tablebase handles stay open within a bounded descriptor budget and close with the runtime. Restart after installing or replacing table files. Missing files produce an explicit unavailable result. Raw WDL and DTZ values are returned with their side-to-move perspective and draw-rule caveats; distance to zeroing is not distance to mate.

get_tablebase_diagnostics reports missing directories and unpaired WDL/DTZ files. Pairing alone does not prove complete material coverage. For offline integrity verification, supply a trusted SHA-256 manifest containing plain filenames:

uv run python scripts/check_tablebases.py
uv run python scripts/check_tablebases.py --manifest /absolute/path/to/checksums.sha256

Tablebase datasets are not downloaded automatically.

Adaptive Teaching and Opening Plans

Use start_teaching_session, get_teaching_hint, and submit_teaching_move for a persisted diagnose, ask, assess, explain, and revisit cycle. Hint levels reveal the concept, candidate piece, candidate move, and legal line in that order. get_concept_progress reports attempts, independent successes, and spaced-review dates. Assisted answers do not earn independent-success credit, and inconclusive engine assessments do not count as failed attempts. These counts are learning history, not a calibrated mastery probability. Ask the host model to discuss the learner's reasoning; Gambit assesses the move, not the semantic quality of an explanation.

create_game_lessons generates up to twenty position-specific lessons from a saved game report for the learner's chosen side. It retains the source report and move history. Repeated topics identify practice candidates; geometric observations do not establish why an engine evaluation changed.

Puzzle answers matching the source continue its line. Alternatives receive two-budget root verification. A sound alternative can be accepted separately, but does not automatically satisfy the source's teaching objective or earn source-solution mastery credit. Unstable alternatives remain inconclusive. Repertoire drills use position identity for spaced review, so transpositions share recall history. get_due_repertoire_drill selects a due position without revealing its moves.

import_opening_plan accepts plans, pawn breaks, piece placements, and common mistakes with a source, source version, and source date. get_opening_plans matches these records by position, including transpositions. Imported advice is explicitly not engine-verified; check a proposed tactical refutation with assess_candidate_move in its actual position. import_opening_statistics records aggregate outcomes with a source date and population description, separately from theoretical advice. get_opening_coverage reports exact local records and positions by source. No external annotated-plan or game-statistics corpus is silently fetched or claimed complete.

The new workflows use versioned output contracts in contracts.py and typed assessment and teaching records. Existing legacy tools retain their response shapes unless an additive field is documented here.

Opening initialization also imports four original starter teaching plans for open games, the Sicilian, the Queen's Gambit Declined, and the French. They are explicitly labeled instructional advice, not engine-verified theory. Tutor turns include matching plans except when hiding a solution. This is a small starter set, not comprehensive annotated opening coverage.

Storage Maintenance and Credentials

Backups use SQLite's online backup API and an integrity check, not a copy of a potentially active database file. Backup and restore destinations must be new files. Restore into a new database, stop Gambit, set GAMBIT_DATABASE to the restored path, and restart. The restore command never overwrites the active database.

Use the local maintenance commands to back up, restore, preview retention, or apply retention with a backup:

uv run python scripts/maintain_storage.py backup --path /absolute/path/to/new_backup.sqlite3
uv run python scripts/maintain_storage.py restore --path /absolute/path/to/new_backup.sqlite3 --destination /absolute/path/to/restored.sqlite3
uv run python scripts/maintain_storage.py prune --age-days 90
uv run python scripts/maintain_storage.py prune --age-days 90 --apply --path /absolute/path/to/new_retention_backup.sqlite3

Retention defaults to a preview. Applying it first creates a backup and removes at most 1,000 old, unreferenced analysis records and 1,000 old checkpoints belonging to completed jobs per invocation. It preserves learner records, lessons, opening knowledge, reports referenced by other collections, and checkpoints for resumable jobs. Nothing is pruned automatically.

The tunnel continues reading credentials from the ignored package-local .env. Gambit's service explicitly removes tunnel credential variables, its dotenv loader imports only GAMBIT_* and STOCKFISH_PATH, and Stockfish subprocesses receive an environment without OpenAI or control-plane variables. Administrative backup, restore, and checksum operations are local CLI commands rather than remote MCP tools.

Speech and AI Clients

create_tutor_turn returns concise teaching text, evidence, and a speech plan. get_speech_plan expands algebraic and UCI move notation into unambiguous spoken moves. synthesize_speech uses an installed espeak-ng executable and returns a local WAV URI. No voice model is bundled. A remote MCP client cannot automatically access local file URIs.

For higher-quality streaming speech, send individual plan segments to the client's preferred TTS provider. The server does not stream audio over MCP or manage microphone input. A client should interrupt playback on a new move and discard directives whose FEN no longer matches the board.

The model should use the tutor prompt and supplied evidence, ask short questions, and call evaluate_move to check the learner's answer. It should not invent tactical motifs from an evaluation number. Learner preferences customize rating band, Socratic/direct style, detail, speech rate, and voice.

Development and Verification

Run these commands from the package directory. The project formatter implements the required four-space indentation, call layout, and logical blank-line rules.

Command

Purpose

uv run pytest

Run the test suite, including installed-engine checks.

uv run ruff check src test scripts

Check imports and Python errors.

uv run python scripts/format_code.py

Apply the project formatting rules.

uv run python scripts/format_code.py --check

Verify formatting without changes.

uv run python scripts/benchmark_analysis.py --iterations 3

Measure local analysis latency and check authored positions.

uv build

Build the source distribution and wheel.

The test suite contains 87 tests. Install Stockfish and the vision extra to exercise the corresponding integrations. The testing badge records the verified suite count, not code coverage; coverage has not been measured.

Tests include strict PGN parsing, engine protocol fakes, storage persistence, training state, knowledge imports, HTTP reconnects, schema publication, interruption, engine recovery, checkpoint reuse, and MCP round trips. Tests named test_live_* exercise installed Stockfish. The benchmark runs fixed authored positions, validates legal variations, tactical results, terminal states, positional facts, and hint leakage, and reports platform and engine provenance alongside cold and warm latency samples. These measurements exclude tunnel latency, model generation, and human teaching-quality assessment. They do not establish an improvement over a previous version without a comparable baseline run. Speech playback remains client-owned; no dedicated voice client is added.

Licensing

Gambit is authored by Mayaz Rakib and distributed under the MIT license.

Stockfish is a separately installed GPLv3 executable; python-chess has its own GPL license. The optional screenshot-recognition model and detector attribution are preserved in third-party notices. The bundled Lichess opening corpus retains its CC0 terms. These third-party terms are separate from Gambit's license.

Available Tools

58 tools
advance_lessonC

Mark a lesson step complete without equating completion to demonstrated mastery.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepYes
lesson_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and mostly does not: it does not say whether the advance is persisted, whether it is idempotent, what happens if the step is out of order, or what the response contains. The only behavioral content is the semantic caveat that completion is not mastery, which is useful framing but not operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is well-sized, though the second clause is a semantic caveat rather than the operational content an agent most needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a state-mutating tool with zero annotations, zero parameter documentation, and no usage context, the description leaves major gaps about persistence, ordering, and prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters, so the description must compensate and does not. It only loosely implies a 'lesson' and a 'step'; there is no explanation of lesson_id format, whether step is 1-indexed or 0-indexed, or valid ranges.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Mark ... complete') and resource ('a lesson step'), so the agent can distinguish it from siblings like start_lesson or get_learner_progress. It stops short of explicitly naming the alternative tools, but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The clause about mastery hints at intent but does not tell an agent in what situation to call this rather than record_training_attempt or start_lesson.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_chessboardB

Recognize a chessboard image and analyze it with Stockfish in one fast call. Intended for axis-aligned screenshots and diagrams. Unreliable OCR is rejected by default so analysis is not based on a silently wrong board.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYes
nodesNo
multipvNo
orientationNoauto
side_to_moveNowhite
allow_unreliableNo
infer_castling_rightsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that unreliable OCR is rejected by default, which explains a key precondition. It still does not describe authentication needs, rate limits, what happens to the rejected image, or the exact output shape and analysis behavior beyond Stockfish being used.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short, front-loaded sentences with no obvious filler. The first sentence states the core action and the last two add important constraints efficiently. It could be slightly more structured around when to use it versus alternatives, but it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is moderately complex: image input, Stockfish analysis, seven parameters, and an output schema. The description covers the high-level pipeline and the reliable-OCR default, which are the most important safety behaviors. Given the output schema exists, return values need not be explained, but the zero parameter documentation and lack of alternative guidance leave clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions none of the seven parameters. Important semantics for nodes, multipv, orientation, side_to_move, allow_unreliable, and infer_castling_rights are left entirely undocumented. The only implied parameter relationship is that allow_unreliable toggles the rejection behavior, which is too thin for a tool with this many knobs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear two-part verb+resource: recognize a chessboard image and analyze it with Stockfish. This distinguishes it from recognize_chessboard (recognition only) and analyze_position (analysis of a known position), because it combines both steps. However, it does not explicitly name or contrast those closest siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an implied usage context: axis-aligned screenshots and diagrams. It does not point to alternatives or say when the tool should be avoided, such as for non-axis-aligned photos or when only recognition/analysis is already available. This is enough to infer the intended input class but not a full usage policy.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_gameC

Analyze a complete standard-chess PGN with adaptive Stockfish searches. Scores use White's perspective; nodes is the configured per-position budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
pgnYes
sideNoboth
multipvNo
initial_nodesNo
critical_nodesNo
critical_loss_cpNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
movesYes
engineYes
resultYes
headersYes
initial_fenYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that scores use White's perspective and that nodes is a per-position budget, adding useful behavioral context. However, it omits key traits like expected runtime (likely long-running due to Stockfish), resource consumption, whether it blocks, or how to retrieve results—critical for an analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no waste, and the core purpose is front-loaded. The second sentence efficiently clarifies two important semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 6 parameters with no schema descriptions, no annotations, and takes a full PGN—a complex operation. The description is far too brief, omitting parameter details, behavioral expectations (runtime, blocking), and distinctions from similar analysis tools. An output schema exists, so return values needn't be explained, but significant gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains 'nodes' (initial_nodes/critical_nodes) and scoring perspective, but leaves multipv, side, and critical_loss_cp unexplained. With 6 parameters and 0% coverage, the description falls well short of providing needed semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Analyze a complete standard-chess PGN'), which is clear. However, it provides no differentiation from siblings like analyze_position, analyze_positions, or review_game, which likely overlap in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as analyze_position or evaluate_position. The phrase 'adaptive Stockfish searches' hints at a specific mode but doesn't state prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_pawn_structureB

Return factual isolated, doubled, and passed-pawn geometry for both sides.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden of behavioral disclosure. It says what is returned but not how the analysis is performed, whether it's purely factual (as 'factual' implies), what permissions or game state it requires, or whether it works on any FEN. The single 'factual' adjective is the only qualifier, and it is vague.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the return scope precisely: pawn structure types (isolated, doubled, passed) and both sides. No filler, no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, so it is reasonably complete for that dimension. However, given no annotations and no usage guidance, an agent lacks context on when this tool is preferable to other position-analysis tools available in the sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the schema provides only a parameter name ('fen') with no description. The description does not explain what a FEN is or how the analysis depends on it, and with only one parameter the baseline is 4, but since the description adds no parameter-level meaning beyond what's in the schema, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('isolated, doubled, and passed-pawn geometry for both sides'), clearly distinguishing it from generic position analyzers like analyze_position or evaluate_position. It's a focused, well-scoped purpose. However, it doesn't explicitly name a sibling alternative or contrast with overlapping tools like inspect_position, keeping it out of 5 territory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus other analysis tools such as analyze_position, inspect_position, or evaluate_endgame. The description only states what it returns, leaving the agent to infer the use case (pawn-structure-specific analysis). No preconditions or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_positionB

Analyze a FEN position with Stockfish. Evaluations are from the side-to-move perspective.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
nodesNo
multipvNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. It usefully discloses the evaluation convention (from the side-to-move perspective), which prevents sign errors, but says nothing about compute cost, latency, blocking behavior, or determinism for an engine call with a configurable node budget.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences with no waste; the core operation is front-loaded and the evaluation convention follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Still, given three undocumented parameters and numerous sibling analysis tools, the definition is only minimally sufficient to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all three parameters. It addresses only 'fen' implicitly and says nothing about 'nodes' (search budget) or 'multipv' (number of lines returned), which are non-obvious tuning knobs an agent needs explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (analyze) plus resource (FEN position) and names the engine (Stockfish), so the operation is unambiguous. However, it does not distinguish itself from close siblings such as evaluate_position, analyze_positions, or evaluate_move, leaving the agent to guess which entry point applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no mention of the many alternative analysis tools (evaluate_position, analyze_positions, evaluate_move). The agent must infer the choice from names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_positionsC

Analyze a batch of FEN requests through one warm Stockfish process. All scores use White's perspective.

ParametersJSON Schema
NameRequiredDescriptionDefault
positionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. 'Through one warm Stockfish process' implies batch efficiency and shared process state but the description does not disclose cost/rate implications, whether results are deterministic given fixed nodes, ordering guarantees, or failure behavior per position. The White's-perspective score convention is a useful addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no waste, front-loading the batch-over-warm-process behavior and the score convention.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the 0% param coverage and absence of annotations leave the agent without needed context on array shape, nodes/multipv semantics, or how this batch tool compares to the many sibling analysis endpoints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It mentions 'FEN requests' generically but does not explain the 'positions' array structure or the per-item nodes/multipv fields (and defaults) that actually control analysis depth and multi-line output. The White-perspective note maps to output meaning rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (analyze FEN requests) and notes the batch/warm-process implementation. However, with many sibling analysis tools (analyze_position, evaluate_position, start_analysis_job, analyze_game, start_progressive_analysis), the description gives no differentiation of when this batch tool is preferred over those.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance. The 'batch through one warm Stockfish process' hints at a multi-position efficiency use case, but the agent is not told to prefer this over single-position siblings or job-based tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_candidate_moveC

Verify a candidate against unrestricted and forced-root searches at two budgets, retaining full supplied history.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
moveYes
historyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
fenYes
caveatYes
historyYes
root_fenYes
verifiedYes
preliminaryYes
best_move_uciYes
classificationYes
schema_versionYes
evidence_statusYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. It mentions 'two budgets' and 'retaining full supplied history', which hints at cost and state behavior, but it doesn't name the budgets, state whether the call is synchronous or long-running, or clarify what is returned. For a comparatively expensive multi-search tool, this is a material gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though the compression is partly why the behavioral and parameter details are absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with zero annotation coverage, zero parameter documentation, and no sibling differentiation, the description is incomplete for a 3-parameter analysis tool with several near-identical siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and all three parameters (fen, move, history) are undocumented in both schema and description. The phrase 'full supplied history' is the only hint about history, but it doesn't explain format, purpose, or when to include it. The description does essentially nothing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Verify' and its object 'a candidate' are stated, along with the search method and budgets. However, the description does not differentiate this tool from close siblings like evaluate_move, compare_candidate_moves, or evaluate_position, which an agent would otherwise confuse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what the tool does but gives no when-to-use context, no conditions under which this is preferred over evaluate_move or compare_candidate_moves, and no exclusions. The agent is left to infer the divot from the bare description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_analysis_jobB

Cancel a job at its next search boundary. An active search has a bounded deadline.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and does usefully disclose the non-obvious deferred semantics: cancellation occurs at the next search boundary, not immediately, and the active search is bounded. But it omits idempotency, whether partial results are retained, and whether the job_id must reference a running job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, and the deferred-cancellation behavior is front-loaded after the core action. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the deferred-cancel nuance is captured. However, given a dense family of siblings (start/get/resume/cancel_analysis_job, cancel_progressive_analysis), the description leaves routing and the job_id parameter under-resolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is exactly one parameter (job_id) with 0% schema description coverage, so the description must compensate. It says nothing about where job_id comes from or what a valid identifier looks like, leaving the single parameter semantically thin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: cancels an analysis job, with a meaningful scope detail ('at its next search boundary'). It does not, however, distinguish itself from the sibling cancel_progressive_analysis, so an agent must infer which cancellable entity each tool targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — you cancel a job you started via start_analysis_job. There is no explicit when-to-use, when-not-to-use, or reference to alternatives such as cancel_progressive_analysis or resume_analysis_job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_progressive_analysisB

Cancel queued or running progressive analysis and mark its result stale.

ParametersJSON Schema
NameRequiredDescriptionDefault
request_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorYes
statusYes
context_idYes
request_idYes
analysis_idYes
schema_versionYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. It does disclose the core effect: queued or running analyses are cancelled and the result becomes stale. However, it omits permissions/ownership requirements, reversibility, idempotency, and error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action and effect with zero wasted words. It is appropriately sized for a one-parameter cancellation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers the core mutation and its visible effect, but for a tool with no annotations it is still missing prerequisites such as how to obtain request_id, permission requirements, and interaction with sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the single required parameter request_id is not explained in the description at all. The description does not say where the request_id comes from or how it relates to start_progressive_analysis or get_progressive_analysis, so it fails to compensate for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Cancel'), a precise resource ('progressive analysis'), and scope ('queued or running'), plus the side effect of marking the result stale. It distinguishes itself from the generic cancel_analysis_job via the resource name, but it does not explicitly say how it differs from that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The eligible states ('queued or running') imply when cancellation is applicable, but the description never states when to choose this over cancel_analysis_job, nor what to do if the request has already completed. No alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_candidate_movesC

Analyze and compare up to eight candidate moves under the same search profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
movesYes
historyNo
profileNoquick

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It hints that all moves share one 'search profile' for fair comparison, but does not disclose cost/compute intensity, whether the call is synchronous or job-based, permission needs, or what 'profile' means behaviorally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the scope constraint ('up to eight') arrives early. It is efficient, though the compression leaves real ambiguity unresolved.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no elaboration. However, for a four-parameter tool with 0% schema coverage, no annotations, and a non-trivial history/profile model, the one-line description leaves essential invocation details missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It touches two of four parameters indirectly ('up to eight ... moves', 'search profile'), but leaves the required 'fen' string and the 'history' array completely unexplained and gives no hint of valid profile values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'analyze and compare ... candidate moves' — with the notable scope detail 'up to eight' and 'under the same search profile'. It is clear what the tool does, but it never distinguishes itself from siblings like assess_candidate_move or evaluate_move, which sound like near-overlapping evaluations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance and no named alternative, despite several sibling tools that evaluate individual moves. The phrase 'compare ... under the same search profile' implies batch comparison, but the agent must infer the selection condition itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_game_lessonsC

Create up to twenty position-specific practice lessons from the learner's mistakes in a saved game analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
learner_idYes
analysis_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
lessonsYes
groupingYes
schema_versionYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a useful cap ('up to twenty') and the origin of the content, but says nothing about side effects on learner progress, permissions required, idempotency, or what happens if the analysis yields fewer mistakes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every clause (quantity cap, position-specific, source of the lessons) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be documented, and the core purpose is stated. However, with no annotations and 0% schema coverage, the missing 'side' semantics and unstated write behavior leave it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It indirectly implies learner_id ('the learner's mistakes') and analysis_id ('saved game analysis'), but the 'side' parameter is entirely unexplained and its expected values are unspecified, leaving a required field ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (create) and resource (game lessons) and scopes it precisely: up to twenty position-specific lessons sourced from mistakes in a saved game analysis. This source framing distinguishes it from generic lesson tools like start_lesson or create_repertoire_drill, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the tool is used after a saved game analysis exists, but gives no explicit when/when-not guidance and does not mention any sibling alternative such as start_lesson or create_repertoire_drill. The agent must infer the workflow position from the phrase 'saved game analysis'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_repertoire_drillC

Create a repertoire recall drill without exposing its expected moves.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it only discloses one trait: expected moves are hidden from the caller. It says nothing about what gets created, whether it mutates learner state, what happens if the FEN is not in the repertoire, or any auth/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with no filler, front-loading the verb and resource. However it is arguably under-specified rather than concise for a state-creating tool with two undocumented parameters, so it sits at minimum viable rather than exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but the description still leaves critical semantics unexplained: what a recall drill is, how fen and learner_id interact, and what committing it does. For a creation tool with no annotations and 0% schema coverage, this is not enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters. The description does not mention 'fen' or 'learner_id' at all, so an agent gets no help understanding that fen is the drill's starting position or how learner_id scopes the hidden expected moves. This is the main gap, since the schema is bare titles only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a repertoire recall drill') that clearly marks it as the creation counterpart to get_due_repertoire_drill and submit_repertoire_move. The clause 'without exposing its expected moves' adds a distinctive scope note, though it doesn't explain what a drill actually is or contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to create a drill versus retrieving one (get_due_repertoire_drill) or submitting a move (submit_repertoire_move), and no prerequisites such as whether the FEN must already belong to the learner's imported repertoire.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_tutor_turnC

Create a grounded tutor turn with engine evidence, teaching guidance, board actions, and speech segments.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
intentNoexplain
historyNo
profileNoinstant
candidatesNo
learner_idNodefault

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says it 'creates' something ('tutor turn') but does not disclose whether this is persisted, whether it is expensive, whether it requires a prior session, or what state it mutates. For a mutation-style tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded, no filler. It is efficient, though perhaps under-specified for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 6 parameters, one required, a 0% schema coverage, no annotations, and an output schema. The description is a single one-line sentence that does not cover parameters, usage context, mutation behavior, or relationship to siblings. It is inadequate for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 6 parameters. The description does not explain any of fen, intent, history, profile, candidates, or learner_id. The only parameter with an obvious meaning is fen, and even that is left implicit. The description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Create a grounded tutor turn'. Lists the constituent content (engine evidence, teaching guidance, board actions, speech segments). However it doesn't differentiate from siblings like tutor_move, start_teaching_session, or get_teaching_hint, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives, no prerequisites. The description does not mention when this should be invoked versus tutor_move or get_teaching_hint, leaving the agent to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_endgameC

Combine endgame features with tablebase evidence or engine analysis when unavailable.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It hints at fallback behavior (using engine analysis when tablebase evidence is unavailable), but says nothing about permissions, safety, rate limits, or output characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, but it is under-specified rather than concise. The most critical information (what the tool actually does) is not front-loaded or clearly stated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an endgame evaluation tool with no annotations and an undocumented parameter, the description is far too sparse. It omits prerequisites, when to use it, and how it relates to sibling tools, leaving major gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention the sole 'fen' parameter at all. It provides no semantic meaning beyond the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says to combine endgame features with tablebase evidence or engine analysis, but it never states that the tool evaluates an endgame position or what it returns. It does not distinguish itself from siblings like evaluate_position, analyze_position, or get_tablebase_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the many other evaluation tools. The mention of fallback to engine analysis is internal behavior, not a usage condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_moveB

Compare a legal move against engine analysis before and after it. Scores are White-relative.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
moveYes
historyNo
profileNoquick

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It usefully states that scores are White-relative, an important interpretation detail, but says nothing about permissions, cost/rate limits, engine requirements, or side effects for what is presumably a read-only analysis call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, no filler. Every clause conveys something useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter analysis tool with no annotations and zero schema description coverage, two sentences are insufficient. Missing: parameter formats, when to prefer this over sibling evaluators, and any behavioral context (auth, engine cost). The output schema does relieve it of explaining return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It hints at the 'move' parameter ('a legal move') and obliquely at 'history' ('before and after'), but leaves 'fen', the move notation format, and the 'profile' parameter (default 'quick') completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compare) and resource (a legal move against engine analysis before and after it), which distinguishes it from position-level siblings like evaluate_position. It stops short of naming the alternative tools (compare_candidate_moves, assess_candidate_move) explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance or conditions for choosing this tool over evaluate_position, compare_candidate_moves, or assess_candidate_move. The pre/post framing implies a use case but is not spelled out as routing advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_positionC

Analyze a position with White-relative scores, immutable evidence, and optional move history.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
movesNo
profileNoquick

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
fenYes
engineYes
profileYes
elapsed_msYes
evaluationYes
history_scopeYes
schema_versionYes
evidence_statusYes
configured_nodesYes
score_perspectiveYes
has_complete_historyYes
principal_variationsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints that scores are White-relative and that evidence is immutable, but says nothing about cost, latency, permissions, whether the profile trades speed for depth, or how the optional move history affects the result. The output-facing hints overlap with the existing output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is appropriately tight. It is arguably too terse for a 3-parameter tool with zero schema coverage, but there is no wasted wording to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, but the definition still omits parameter meaning, sibling routing, and any behavioral profile for a compute-style evaluation tool. Given the dense cluster of similar chess-analysis siblings, one sentence is not enough for correct selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all three parameters and largely does not. It alludes to the optional move history (the 'moves' param), but leaves 'fen' unexplained and gives no indication of what the 'profile' values are or what 'quick' (the default) changes about the evaluation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (analyze/evaluate a position) and adds scope hints about White-relative scores and immutable evidence. However, it does not distinguish itself from near-synonymous siblings such as analyze_position, inspect_position, explain_position, and evaluate_move, leaving the agent to guess which position-scoring tool to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of any alternative tool. In a namespace crowded with analyze_position, analyze_chessboard, inspect_position, and evaluate_move, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_positionB

Inspect legal state, material, pins, attacks, and pawn structures without an engine.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses one behavioral trait (no engine used), but says nothing about whether this is read-only, what computational cost to expect, or what form the inspection output takes. For an analysis tool with zero annotation coverage this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and enumerating the covered aspects. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and one obvious parameter keeps the input side simple. But with no annotations and a crowded sibling cluster, the description should say more about when this engine-free inspection is preferable and how it differs from the many neighboring position-analysis tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (fen) exists and schema description coverage is 0%, so the description must compensate — but the parameter name 'fen' is self-explanatory to a chess-domain agent and the description correctly frames the input as a position to inspect. Baseline for a single obvious parameter with an output schema present is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (inspect) and enumerates specific aspects (legal state, material, pins, attacks, pawn structures), which is more specific than a generic 'analyze position'. However, it doesn't differentiate from close siblings like inspect_position, analyze_position, or analyze_pawn_structure, which likely overlap significantly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without an engine' implies a when-to-use condition (engine-free inspection vs engine-based evaluation), but this is oblique. There are many siblings in the analysis cluster (evaluate_position, analyze_position, inspect_position) and the description doesn't explicitly state when to choose this one over them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_openingB

Search locally imported opening names with pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the data scope ('locally imported') and pagination behavior, but omits matching rules, permissions, rate limits, and return behavior; partial behavioral context only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is appropriately sized for a simple search operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read/search operation and has an output schema, so return details are not required in the description. However, with 0% parameter coverage and many sibling opening tools, the description leaves important query semantics and when-to-use guidance unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It weakly maps 'opening names' to the query and 'pagination' to offset, but does not explain query matching semantics, case sensitivity, or how offset/pagination behaves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Search'), resource ('opening names'), and scope ('locally imported'), plus pagination. It does not explicitly name sibling alternatives, but the local-import qualifier helps distinguish it from analysis and plan-oriented opening tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when searching locally imported opening names, but gives no explicit when-not conditions or alternatives such as get_opening_plans or get_opening_position. Usage is only implied by the tool's narrow scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysisC

Retrieve an immutable position analysis or completed game report.

ParametersJSON Schema
NameRequiredDescriptionDefault
analysis_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Immutable' and 'completed' hint that the artifact is final (vs. an in-progress job), but there is nothing about permissions, what happens if the id refers to an unfinished job, or error behavior for unknown ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, though it is arguably under-specified rather than tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the tool is a simple one-param getter. Still, with no annotations and an undocumented id parameter, the definition falls short of what an agent needs to distinguish it from the many other analysis-retrieval siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, so the description must compensate and does not: it never explains what analysis_id is, its format, or where to obtain it. The only semantic hint is that it retrieves one analysis/report.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (position analysis or completed game report), and the 'immutable'/'completed' qualifiers distinguish it from the live-job siblings. It does not name get_progressive_analysis or get_analysis_job, so an agent must still infer which of the several analysis getters applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no exclusion of alternatives, despite many sibling getters (get_progressive_analysis, get_analysis_job, get_analysis_metrics). The agent is left to infer that this is the completed/immutable report fetch rather than a running-job poll.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysis_jobC

Read analysis job status and completed-search progress.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Read' implies a non-mutating operation, but nothing is said about polling cadence, behavior for invalid/expired job IDs, or whether completed-search progress is a partial vs terminal state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence that front-loads the resource and is free of filler. It is efficient, though the brevity reflects under-specification rather than disciplined trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a job-status reader tied to a start/cancel/resume workflow the description should say how the job is obtained and when polling is appropriate. Those workflow links are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter job_id has 0% schema description coverage and the description adds nothing about it - not its origin, format, or validity constraints. With low coverage the description should compensate, and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('analysis job status and completed-search progress'), which is more than a restatement of the name. However, it does not distinguish this from close siblings such as get_analysis, get_progressive_analysis, or start/cancel/resume_analysis_job, leaving the agent to guess which job-reader to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus the sibling analysis tools, and no indication that job_id is expected to come from start_analysis_job or that this is the polling endpoint. Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysis_metricsB

Report bounded engine latency samples, queue delays, cache hits, and reserved interactive capacity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
cache_hitsYes
queue_p95_msYes
sample_countYes
search_countYes
pending_countYes
latency_p50_msYes
latency_p95_msYes
schema_versionYes
has_reserved_interactive_capacityYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states what metrics are reported but does not say whether this is a read-only operation, what permissions are required, whether sampling is continuous or point-in-time, or how the bounded samples are scoped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that names the operation and enumerates the returned metric categories without filler. It is appropriately sized for a zero-argument reporting tool, though the phrase "bounded engine latency samples" is somewhat jargon-heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values, and the zero-parameter schema requires no parameter documentation. However, with no annotations and no usage guidance, the description is only minimally complete for an agent deciding when and why to call this metrics tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no argument semantics to explain and the schema is trivially complete. The description appropriately does not introduce parameter details, matching the baseline for a zero-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ("Report") and lists the exact metric categories it returns: engine latency samples, queue delays, cache hits, and reserved interactive capacity. It does not, however, distinguish this metrics tool from sibling analysis tools such as get_analysis_job or get_progressive_analysis, so an agent must infer the intended monitoring context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, no alternatives, and no conditions for selecting this tool over the many sibling analysis tools. An agent is left to infer that this is for observing analysis engine metrics rather than running or inspecting analysis jobs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_capabilitiesA

Report engine settings and exact local corpus counts without launching an engine.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Report' implies a read-only operation and 'without launching an engine' is a useful behavioral trait, but it does not explicitly state side effects, permissions, or performance characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It immediately states what is reported and the key constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter capability-reporting tool with an output schema, the description is largely complete: it says what is reported and that no engine is launched. It could be slightly richer on when to choose this over sibling inspection tools, but no critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema fully documents this. With no parameters to explain, the baseline score of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and names the exact resources ('engine settings' and 'exact local corpus counts'), plus a distinguishing constraint ('without launching an engine'). This clearly separates it from sibling analysis, evaluation, and job-control tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage context: call this when you need engine settings or local corpus counts without starting an engine. It does not name alternative tools or state explicit exclusions, but the context is specific enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_concept_progressC

Return concept-specific attempts, independent successes, and due review dates.

ParametersJSON Schema
NameRequiredDescriptionDefault
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
conceptsYes
schema_versionYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it only restates output fields (attempts, successes, due dates) that the output schema already documents. It never discloses that this is a read-only operation, whether learner authorization is required, or any scoping/rate constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single terse sentence with no filler, front-loaded with the verb and the returned content. It earns its place but is arguably too sparse for the ambiguity it faces.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Return values are covered by the existing output schema and the single param is trivial, so the basics are present. However, with no annotations and several overlapping siblings, the description omits the usage routing an agent needs to invoke this tool correctly rather than get_learner_progress.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter (learner_id) is never mentioned in the description. The param name is largely self-explanatory, which keeps this from being worse, but the description does nothing to compensate for the missing coverage (e.g., format, source of the ID).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and a clear resource set (concept-specific attempts, independent successes, due review dates). It is distinct from a generic progress tool by the word 'concept-specific', but it never names or differentiates itself from the similarly-scoped sibling get_learner_progress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use, when-not-to-use, or alternative guidance. With siblings like get_learner_progress and get_due_repertoire_drill clearly overlapping in the learning-progress/review space, the agent has no basis for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_due_repertoire_drillB

Select a due repertoire position using transposition-aware spaced review and hide expected moves.

ParametersJSON Schema
NameRequiredDescriptionDefault
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses two non-obvious traits: review is transposition-aware and expected moves are hidden (so the agent must not assume the answer is visible). However, it omits whether selecting a drill mutates review state or advances a schedule, and gives no auth/permission context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core action front-loaded and no filler. It is efficient, though the trailing 'hide expected moves' clause could be clearer about whether hiding is a guarantee or a side effect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. Still, for a stateful training tool with no annotations, the definition leaves key call-time questions open: does invoking it consume or advance a review item, and what prerequisites accompany learner_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, learner_id, with 0% schema description coverage and no mention in the description. The description adds no meaning about whose repertoire is being drawn from or what a valid learner_id looks like, so the agent must infer it entirely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (select) and resource (a due repertoire position) plus the mechanism (transposition-aware spaced review), which distinguishes it from siblings like get_repertoire_move and create_repertoire_drill. It stops short of explicitly naming an alternative, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus get_repertoire_move, create_repertoire_drill, or submit_repertoire_move, and no prerequisites or sequencing notes. The phrase 'due' implies a scheduling condition but never states what makes a position due or when an agent should invoke this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_learner_progressC

Read persisted learner preferences and a paginated training-progress view.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNo
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It mentions 'persisted' and 'paginated' but does not disclose permissions, whether preferences are cached, pagination defaults, or rate limits. An agent lacks key operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core purpose. Minimal waste, though it tries to combine two concepts without clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, an output schema present (so returns need not be described fully), and 0% schema description coverage, the description is insufficient. It should clarify the two data sources, pagination behavior, and parameter usage to be callable correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only implies a pagination mechanism via 'paginated' but does not explain the 'offset' parameter's role, default, or the meaning/format of 'learner_id'. Significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States two distinct actions (reading learner preferences and a paginated training-progress view) but is vague about the resource relationship. No sibling differentiation despite many related learner/training tools in the list (e.g., get_concept_progress, update_learner_profile, record_training_attempt).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as get_concept_progress or update_learner_profile. No prerequisites, authentication context, or exclusions mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_opening_coverageB

Report exact local opening coverage by source, keeping theoretical plans and practical statistics distinct.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourcesYes
source_limitYes
schema_versionYes
is_exhaustive_chess_coverageYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It implies a read-only report and adds useful context that theoretical plans and practical statistics are kept distinct in the output, but it never explicitly states read-only safety, permissions, or scope limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no redundant phrasing. Every clause (exact, local, by source, theoretical vs practical) carries meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with zero parameters the schema/invocation side is trivial. However, the description omits when-to-use routing against the many sibling opening tools, leaving a meaningful gap for selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric baseline is 4. There is nothing for the description to clarify beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('local opening coverage by source'), which is reasonably distinct from siblings like get_opening_plans or find_opening. It does not name a sibling or explicitly contrast itself, and 'coverage' is left abstract, but an agent can infer this is a coverage-summary read tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or named alternatives are given. The clause about keeping theoretical plans and practical statistics distinct hints at output content but does not tell the agent when to pick this tool over get_opening_plans, import_opening_statistics, or get_opening_position.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_opening_plansC

Retrieve transposition-matched opening plans separately from practical game statistics.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
plansYes
caveatYes
schema_versionYes
practical_statisticsYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses only that matching is transposition-based and results are kept apart from statistics; it says nothing about prerequisites (e.g., whether an opening plan must be imported first), error behavior, or result scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the key discriminator (transposition-matched plans vs. game statistics) is stated immediately. It is arguably under-specified rather than verbose, which does not hurt conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with no annotations, a 0%-covered required parameter, and no indication of prerequisites or how this differs from the many neighboring opening/analysis tools, the description is not complete enough to call the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'fen' has 0% schema description coverage, so the description must compensate — yet it says nothing about the FEN format, whether a full or partial position is accepted, or how transposition matching uses it. The parameter's meaning is left entirely implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (opening plans) and adds the qualifier 'transposition-matched,' which distinguishes the result set from generic opening data. It also implicitly separates itself from statistics-oriented siblings, though it does not name any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'separately from practical game statistics' hints that this tool is the alternative to statistics tools, implying when to prefer it. However, no explicit when-to-use condition, prerequisites, or named alternative (e.g., get_opening_coverage, import_opening_statistics) is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_opening_positionC

Get source-attributed corpus moves at a position, including transpositions.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions 'source-attributed' and 'transpositions', which hints at behavior, but does not state whether the operation is read-only, what the output contains beyond a generic 'moves' return, whether pagination applies (offset parameter exists), or any rate or permission constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no filler. It front-loads the action and the key modifiers. It could be slightly more structured to separate conditions or return behavior, but it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, 0% schema description coverage, and many sibling tools, the description is not complete enough. It does not explain required FEN syntax, the offset/pagination behavior, or when this tool is preferable to alternatives like find_opening or inspect_position. An agent would need to infer too much.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate. It implies the position is specified (presumably via 'fen'), but does not explain the required FEN format or the purpose of the 'offset' parameter. Baseline is 3 when schema descriptions are high, but here the low coverage means the description falls short; however, it does at least convey what the queried entity is.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('corpus moves at a position'), and adds distinguishing details like 'source-attributed' and 'including transpositions'. However, it is not sharply differentiated from siblings such as find_opening, evaluate_position, or inspect_position, which also operate on positions. The purpose is clear but lacks explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. With many sibling tools that analyze, inspect, or find openings, an agent has no indication of which context should select get_opening_position. No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_progressive_analysisC

Read progressive analysis status and its completed evidence ID, withholding stale results.

ParametersJSON Schema
NameRequiredDescriptionDefault
request_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorYes
statusYes
context_idYes
request_idYes
analysis_idYes
schema_versionYes

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose one important behavioral trait: stale results are withheld, which is valuable beyond schema. But it omits whether the call is read-only, whether it blocks or polls, what a pending vs completed status looks like, or any rate/refresh guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single compact sentence front-loaded with the verb and resource, then the behavioral caveat. No wasted words. Slightly terse given how much it leaves unsaid, but structurally sound.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is not required, but the description should still cover when to poll and how request_id is obtained given 0% schema coverage. As a status-polling tool in a dense sibling family with no annotations, it leaves too much unspecified for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter 'request_id', so the description should compensate but does not mention it at all. With only one parameter and no enum, the baseline is moderate; the description leaves the identity of the request_id (which start call returns it) entirely to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and resource ('progressive analysis status') plus a notable behavior ('withholding stale results'). However, it does not distinguish itself from siblings like get_analysis, get_analysis_job, or get_progressive_analysis's natural counterpart cancel_progressive_analysis; it sits in a crowded family of polling/status tools without clarifying its niche.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. An agent cannot infer whether to call this instead of get_analysis_metrics or get_analysis_job, nor when to escalate to start_progressive_analysis. No prerequisites or timing guidance (e.g., poll after start) are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_repertoire_moveC

Get personal repertoire moves at a position.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it says almost nothing beyond 'get' implying a read. It never explains what 'personal' means, whose repertoire is consulted, whether an empty position returns an error, or what the offset is for (pagination?).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient rather than padded, though the brevity edges into under-specification rather than crisp conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema does relieve it of explaining return values, but with no annotations, 0% parameter coverage, and an ambiguous offset argument, the definition is not complete enough for an agent to invoke it confidently in a crowded repertoire-related toolset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. 'At a position' loosely implies the fen input, but the offset parameter is left completely unexplained — is it pagination over matching moves, and what are its units? Two undocumented parameters at 0% coverage is a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get personal repertoire moves') with a scoping qualifier ('at a position'). It clearly contrasts with the mutation sibling submit_repertoire_move, though it does not distinguish itself from get_due_repertoire_drill or the opening-lookup tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named despite many related siblings (submit_repertoire_move, get_due_repertoire_drill, find_opening, get_opening_position). Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_speech_planC

Create interruptible speech segments and legal SAN/UCI pronunciations.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
movesNo
segmentsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that segments are 'interruptible' and that pronunciations are legal, but it does not describe side effects, auth requirements, idempotency, or what the plan actually contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. However, it is under-specified rather than appropriately concise for a tool with three parameters and no annotation or schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with no annotations, 0% input schema description coverage, and no mention of required parameters like 'fen' or 'segments', the description is not complete enough for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention any of the three parameters ('fen', 'moves', 'segments'). The required chess position FEN and segment list are completely undocumented, so the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific output ('interruptible speech segments and legal SAN/UCI pronunciations'), which helps identify the tool's role. However, the tool name is 'get_speech_plan' while the description says 'Create', and it never explains what a speech plan is or how it relates to chess positions, leaving the purpose only partially clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as 'synthesize_speech' or other speech/teaching tools. The description implies a use case but does not state conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tablebase_diagnosticsA

Inspect configured tablebase directories and WDL/DTZ file pairing without claiming complete coverage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
caveatYes
dtz_filesYes
wdl_filesYes
directoriesYes
schema_versionYes
unpaired_filesYes
has_complete_pairsYes
missing_directoriesYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the caveat that the report may be incomplete ('without claiming complete coverage'), which is genuine behavioral context, but it says nothing about read-only status, cost, or what a partial result implies for the caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the operation front-loaded and the completeness caveat attached economically. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema already covering the return shape, the description needs only to convey intent and scope, which it does. The only shortfall is the absence of any guidance on when to reach for this versus sibling tablebase tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description adds no parameter detail, but none is needed given the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('inspect') and resource ('configured tablebase directories and WDL/DTZ file pairing'), which is concrete and distinguishable from the sibling get_tablebase_result (which returns a result rather than checks configuration). The scope is clear even without naming the sibling it contrasts with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as get_tablebase_result or evaluate_endgame. The diagnostic intent ('diagnostics') is implied, but an agent cannot tell from the text when this should be preferred over the other tablebase tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tablebase_resultC

Probe configured local Syzygy WDL and DTZ with explicit availability and draw semantics.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose two behavioral traits beyond the name — explicit availability reporting (tablebases may not be configured) and draw semantics — which is real value. It stops short of describing failure modes when Syzygy files are absent or the position is out of range.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler, and the resource scope leads. It borders on under-specification rather than verbosity, but the structure itself is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail need not appear in prose, and the operation is low-complexity (one parameter). Still, the absence of any parameter guidance and of usage routing leaves the definition thin for a tool the agent must call with a correctly formatted FEN.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter (fen) has 0% schema description coverage, and the description adds nothing about its format or expectations. Since FEN is a standard but easily malformed input, this leaves the agent without guidance on valid values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a resource (Syzygy WDL/DTZ tablebases) and a verb (probe), so the general purpose is recoverable. However, "probe configured local Syzygy WDL and DTZ" is jargon-heavy and does not clearly state that it returns the tablebase evaluation for a given position, nor does it distinguish itself from the sibling get_tablebase_diagnostics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer (e.g., get_tablebase_diagnostics or evaluate_endgame). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tacticB

Select a due puzzle near the learner's rating and begin a solution-hidden session.

ParametersJSON Schema
NameRequiredDescriptionDefault
themeNo
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the session is 'solution-hidden' and that puzzle selection is rating-matched, but it does not say whether this starts/creates stateful session data, what prerequisites or permissions apply, or what happens if no due puzzle exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler: verb, resource, and scope qualifiers in order of importance. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a stateful session-starting tool with 0% parameter coverage and no annotations, the description omits prerequisites and the meaning of its parameters, leaving real gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for two undocumented parameters. It only gestures at learner context ('the learner's rating') and says nothing about the required learner_id or the optional theme filter, leaving the theme parameter entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb (select) and resource (a due puzzle) plus scope qualifiers (near the learner's rating, solution-hidden session). It is clear what the tool does, though it never names the sibling it differs from (e.g., get_due_repertoire_drill), and the name 'get_tactic' vs 'puzzle' wording is slightly loose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the usage context (training flow, learner is due for a puzzle) but gives no explicit when-to-use/when-not guidance and does not point to alternatives like get_due_repertoire_drill or the follow-up submit_tactic_move. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_teaching_hintB

Reveal one of four progressive hint levels, recording assistance without exposing the full solution early.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelYes
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
fenYes
statusYes
conceptYes
questionYes
hint_levelYes
session_idYes
schema_versionYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a side effect ('recording assistance') and a progressive constraint ('without exposing the full solution early'), which is meaningful. However it says nothing about session prerequisites, whether hints are consumed/monotonic, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the reveal action front-loaded and the behavioral caveat trailing. No wasted words, though it is arguably too terse for a zero-annotation, zero-coverage tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations and 0% parameter coverage, the description should do more to cover session requirements and the semantics of the level range to be fully actionable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning for 'level' by indicating a range of four progressive hints, but says nothing about 'session_id' (its format, source, or required lifecycle). Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reveal) and resource (hint), with the scope clarified as 'one of four progressive hint levels'. This clearly distinguishes it from siblings like submit_teaching_move or start_teaching_session, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'recording assistance without exposing the full solution early' implies pedagogical usage during a teaching session, but there is no explicit when-to-use, when-not, or named alternative. The usage context is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_opening_planB

Import a dated, versioned opening teaching plan with source attribution. Advice is not engine-verified automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
fenYes
nameYes
plansYes
sourceYes
pawn_breaksYes
source_dateYes
position_keyYes
evidence_kindYes
schema_versionYes
source_versionYes
common_mistakesYes
piece_placementsYes
is_engine_verifiedYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one genuinely useful trait: the imported advice is not engine-verified. But it says nothing about mutation semantics – whether an import overwrites an existing plan, how source_version conflicts are handled, or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose and followed by the one behavioral caveat. No filler, no restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a mutation tool with no annotations and undocumented nested fields, the description omits conflict/overwrite behavior and prerequisites. It is roughly minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It loosely maps to the source/source_date/source_version fields by saying 'dated, versioned ... with source attribution', but the remaining required fields (fen, name, plans, pawn_breaks, piece_placements, common_mistakes) are given no meaning anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (import) and resource (dated, versioned opening teaching plan with source attribution), which is clearer than a tautology. However, it does not differentiate from siblings like import_opening_statistics, import_opening_tsv, or import_study, which are also 'import' tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of why to choose this importer over import_opening_tsv, import_opening_statistics, or import_study. The only contextual hint is the verification caveat, which is not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_opening_statisticsC

Import source-attributed aggregate opening outcomes for an explicitly described game population.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
moveYes
drawsYes
sourceYes
black_winsYes
populationYes
white_winsYes
source_dateYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
drawsYes
sourceYes
move_uciYes
black_winsYes
populationYes
white_winsYes
source_dateYes
position_keyYes
schema_versionYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a write operation and hints at provenance ('source-attributed') and population scoping, but never states whether existing statistics are merged or overwritten, whether duplicates are detected, what permissions are needed, or what happens on partial input. For an 8-required-parameter write tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is appropriately sized, though its brevity is partly under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but the tool has 8 required parameters with zero schema documentation, no annotations, and no output-schema relief for input semantics. The description does not come close to filling that gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 required parameters (fen, move, source, source_date, population, and three win/draw counts), and the description explains none of them. It only gestures at 'source-attributed' and 'population', leaving formats like the FEN string, move notation, and source_date encoding undocumented everywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Import') and resource ('source-attributed aggregate opening outcomes') with a scope qualifier ('explicitly described game population'). This distinguishes it from sibling writers like import_study and import_opening_plan, though it never explains how it differs from import_opening_tsv, which appears to be the closest analogue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool, when not to, or which sibling import path to prefer for a given data format. With four other import_* siblings in the list, the absence of routing guidance is costly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_opening_tsvC

Import a bounded Lichess-format opening TSV batch with explicit source attribution.

ParametersJSON Schema
NameRequiredDescriptionDefault
tsvYes
sourceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden for what is clearly a write operation. It does not disclose permissions needed, idempotency/duplicate handling, batch size limits implied by 'bounded', or what happens to existing opening data on import.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the core purpose is front-loaded. It is appropriately sized, though the compression is partly responsible for the missing detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for an unannotated mutation tool with 0% parameter coverage the description is thin. It never defines the batch bound, the source format, or the failure/conflict behavior an agent would need before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for two undocumented parameters. It does add real meaning to both: the tsv parameter is a Lichess-format TSV batch and the source parameter is an explicit attribution field, but no column format, delimiter, or expected source values are given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource (import an opening TSV batch) plus the source format (Lichess-format), which helps separate it from siblings like import_opening_plan and import_opening_statistics. It stops short of explicitly stating what makes it different from those other import tools, so an agent must infer the distinction from the format word alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement about when another import tool is preferable, and no prerequisites. The word 'bounded' hints at a size constraint but the description never states what the bound is or how to comply with it, leaving the agent to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_puzzlesC

Import a bounded puzzle CSV batch, validating the setup move and complete source line.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
csv_textYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose that each batch validates the setup move and the completeness of the source line, which is useful, but it omits what happens on validation failure, whether imports are all-or-nothing, size limits implied by 'bounded', and the fact that this is a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with the verb and resource first, and no padding. It is appropriately sized even if under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, but for a mutation/import tool with no annotations and 0% parameter coverage the description leaves too much unstated: format expectations, failure behavior, batch limits, and how it relates to the other import_* siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is described in the schema. The description implies csv_text holds CSV content and references a 'source line', giving marginal meaning, but it never defines the expected CSV columns or what values 'source' accepts, so the coverage gap is not compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (import) and resource (puzzle CSV batch), and the name itself separates it from siblings like import_opening_tsv or import_study. However, the description never explicitly contrasts with those alternative importers, so differentiation relies on the tool name and domain nouns rather than stated guidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'bounded' hints that batch size is limited, but there is no when-to-use guidance, no prerequisites, and no mention of the alternative import tools (import_opening_tsv, import_opening_plan, import_study) an agent could pick instead. An agent is left to infer routing entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_studyC

Import a source-attributed PGN study, including variations and transpositions.

ParametersJSON Schema
NameRequiredDescriptionDefault
pgnYes
nameYes
sourceYes
is_repertoireNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says what is imported, including variations and transpositions, but does not disclose side effects, permissions, duplicate handling, or whether the import creates or overwrites data. For an import/mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is efficient and easy to scan, though its brevity leaves much unsaid for a 4-parameter import tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, with no annotations and 0% schema description coverage, the description is far too sparse for a 4-parameter import operation. It omits parameter meanings, side effects, and usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all 4 parameters. It only implicitly touches on 'pgn' and 'source' via 'PGN study' and 'source-attributed', while 'name' and 'is_repertoire' are not addressed. It adds little beyond the parameter titles already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: importing a PGN study. It adds scope with 'source-attributed' and 'including variations and transpositions', distinguishing it from generic parsing tools. It does not explicitly name sibling alternatives such as import_opening_plan or parse_game, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no alternatives. It does not say when to choose import_study over parse_game, import_opening_plan, or other import tools. Only the implicit context of importing a PGN study is available.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

initialize_openingsB

Load the bundled, version-pinned Lichess named opening corpus into local storage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that data is bundled (implying no network fetch) and version-pinned, and that it writes into local storage — a mutation of local state — but it says nothing about idempotency, whether it overwrites existing local data, expected runtime, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the key qualifiers ('bundled', 'version-pinned', 'local storage') are packed efficiently into one clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and no parameters, the definition has little left to explain about return values, so it is structurally adequate for a simple zero-arg seed operation. However, for an initialization tool it omits the when-to-run and repeat-safety information an agent needs before triggering it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document beyond what the empty schema already conveys; the description adds the only meaningful context (what corpus is loaded and where). This meets the zero-parameter baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Load') and a precise resource ('bundled, version-pinned Lichess named opening corpus'), making the operation unambiguous. It implicitly contrasts with user-supplied import siblings (import_opening_plan, import_opening_tsv) by stressing 'bundled' and 'version-pinned', though it never names an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus the many other opening-corpus tools, nor any prerequisites such as 'run once before find_opening/get_opening_position'. The agent must infer the initialization role from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_positionA

Inspect a FEN position without engine analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the tool performs no engine analysis, which is a meaningful behavioral trait, but it does not explicitly state read-only safety, side effects, or what 'inspect' entails beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the core action and the key constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple one-parameter schema and the presence of an output schema, the description is nearly complete: it states the action, the input type, and the no-engine-analysis scope. It could be improved by clarifying what 'inspect' returns, but the output schema likely covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single 'fen' parameter. It identifies the parameter as a 'FEN position,' which adds some meaning beyond the bare schema name, but gives no format details or examples. This partially compensates but leaves a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Inspect') and resource ('FEN position'), and distinguishes itself from sibling analysis tools by saying 'without engine analysis.' However, 'inspect' remains somewhat vague about what aspects of the position are examined, so it falls short of the most precise 5-level clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without engine analysis' implies when this tool is appropriate compared to engine-based siblings like evaluate_position or analyze_position, but it does not explicitly name alternatives or state when-not-to-use conditions. Usage is only implied, not fully guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parse_gameA

Parse one standard-chess PGN without Stockfish. Plies are one-based from the supplied starting position.

ParametersJSON Schema
NameRequiredDescriptionDefault
pgnYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
movesYes
resultYes
headersYes
initial_fenYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that no Stockfish is used and that plies are one-based from the supplied starting position, but it does not state whether the operation is read-only, how invalid PGN is handled, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no wasted wording. The core purpose is front-loaded and the output indexing detail follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return values. For a single-parameter parsing tool, it covers the essential purpose, engine exclusion, and ply indexing; only error and format constraints are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented 'pgn' parameter. It adds meaningful constraints by specifying 'one standard-chess PGN' and clarifying ply indexing, but it does not explain accepted PGN variants, headers, comments, or invalid-input behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Parse'), a clear resource ('standard-chess PGN'), and a distinguishing constraint ('without Stockfish'). An agent can tell this apart from Stockfish-based evaluation or analysis siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without Stockfish' implies lightweight parsing rather than engine-backed analysis, which gives some usage context. However, it does not explicitly say when to choose this over siblings like analyze_game, review_game, or import_study, nor does it state exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recognize_chessboardA

Convert an axis-aligned 2D chessboard screenshot or book diagram to FEN locally. Pass image as a local path, file URL, raw base64, or base64 data URL. Returns confidence and reliability; do not silently trust a result whose reliable or plausible field is false.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYes
orientationNoauto
side_to_moveNowhite
infer_castling_rightsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses that processing happens locally (no upload), that confidence/reliable/plausible style fields are returned, and warns about unreliable outputs. It omits whether the operation is read-only/idempotent or what happens on failure, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core conversion and its scope, followed by input formats and the reliability caveat. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description helpfully warns about the confidence/reliability semantics anyway. It is complete for the primary flow but leaves three of four parameters undescribed, which is the main gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 4 parameters, so the description must compensate. It richly documents the required 'image' param (local path, file URL, raw base64, base64 data URL) but says nothing about orientation, side_to_move, or infer_castling_rights, leaving three params to enum/default inference alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Convert an axis-aligned 2D chessboard screenshot or book diagram to FEN') and adds a distinguishing scope qualifier ('locally'), letting an agent separate it from siblings like analyze_chessboard or inspect_position without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear input-context guidance (the kinds of images accepted: screenshot or book diagram) and an explicit caution on when NOT to trust the result ('do not silently trust... reliable or plausible field is false'). It stops short of naming an alternative sibling or the condition that routes to analysis tools, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_training_attemptC

Record an explicitly reported training outcome with idempotent attempt identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
puzzle_idYes
attempt_idYes
learner_idYes
was_successfulYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses idempotency tied to attempt identity, which is real behavioral value, but says nothing about permission requirements, whether re-submission mutates existing records, or error behavior on conflicting attempt_ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and no filler. Nothing is wasted, though it is arguably too terse given the documentation gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, but with 0% parameter documentation, no annotations, and four required inputs, the definition leaves the agent guessing about argument formats and write semantics for a state-mutating tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and all four parameters are undocumented. The description maps loosely to was_successful (outcome) and attempt_id (attempt identity), but learner_id and puzzle_id are given no meaning, nor are formats or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (record) and resource (training outcome/attempt), and the idempotency qualifier hints at the distinct semantics versus a generic write. It is reasonably distinguishable, though it never names a sibling it should be preferred over.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'explicitly reported training outcome' implies the caller is passing a known result rather than triggering evaluation, but there is no explicit when-to-use, when-not-to-use, or alternative among the many sibling analysis/lesson tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_analysis_jobA

Resume an interrupted, failed, or cancelled game job using compatible durable search checkpoints.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
pgnYes
errorNo
statusYes
analysis_idNo
should_reviewYes
schema_versionYes
reused_searchesNo
completed_searchesYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the dependence on durable search checkpoints and checkpoint compatibility, which is useful non-obvious behavior, but says nothing about what happens when checkpoints are incompatible, whether resume is idempotent, or whether prior results are preserved or discarded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding. The precondition (interrupted/failed/cancelled) and the mechanism (compatible durable search checkpoints) are both stated compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation here, and for a one-parameter tool the description covers the essential state precondition and mechanism. The remaining gap is behavioral detail around checkpoint incompatibility, which would need annotations or a longer description to close.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter, job_id, has 0% schema description coverage and the description adds no format or sourcing guidance for it. The only added meaning is an implicit constraint that the referenced job be in a resumable state, which is marginal over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resume) and resource (game/analysis job) with the exact precondition states (interrupted, failed, cancelled). This inherently separates it from start_analysis_job and cancel_analysis_job, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when to call it: only for an existing job in an interrupted, failed, or cancelled state, and only with compatible checkpoints. No alternative is named for the incompatible-checkpoint case, but the conditional framing is genuinely actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_gameB

Analyze and classify a complete PGN with deterministic Gambit review labels and per-player summaries. No prose coaching is generated.

ParametersJSON Schema
NameRequiredDescriptionDefault
pgnYes
sideNoboth
multipvNo
initial_nodesNo
critical_nodesNo
critical_loss_cpNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
movesYes
engineYes
resultYes
headersYes
summariesYes
initial_fenYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden; it does disclose determinism and the absence of prose output, which is genuinely useful. However, it says nothing about the tool being compute/engine-heavy (implied by multipv and node budgets) or about expected latency, so key behavioral traits remain hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and with the no-prose caveat placed at the end. No wasted words, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But for a 6-parameter analysis tool with zero schema coverage and no annotations, the description leaves the entire configuration surface and any cost/latency expectation unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, and the description explains none of them. Opaque tuning knobs like multipv, initial_nodes, critical_nodes, and critical_loss_cp are left entirely undefined in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Analyze and classify a complete PGN') and characterizes the output as deterministic review labels with per-player summaries, which separates it from prose-generating lesson tools. It does not explicitly name a sibling, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'No prose coaching is generated' implicitly routes the agent away from coaching/tutor tools toward this one for structured review, but there is no explicit when-to-use statement or named alternative. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_analysis_jobC

Start a bounded, persistent background game analysis or review job.

ParametersJSON Schema
NameRequiredDescriptionDefault
pgnYes
should_reviewNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose useful async semantics ('persistent, background job'), implying the caller must poll for results, but it never explains what 'bounded' means, whether the job is cancellable, or how the returned job handle is used for retrieval.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the object of the action appears immediately. It is perhaps too terse for the amount of unstated behavior, but it wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema covers return values, but the definition omits the job lifecycle (retrieve via get_analysis_job, cancel via cancel_analysis_job) and any usage context, which is a substantial gap for a start-job mutation with zero annotation support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for two undocumented parameters. 'Analysis or review job' faintly hints at the should_review flag, but the required 'pgn' parameter is never explained (format, size, source), leaving the most important input ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (start) and resource (analysis/review job), and the qualifiers 'bounded, persistent, background' begin to distinguish it from the sibling start_progressive_analysis. However, it never names an alternative, so an agent must infer the split between the two job starters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g. a parsed game or valid PGN), and no pointer to the sibling tools get_analysis_job / cancel_analysis_job that the caller will obviously need next. The description gives no selection criteria among the many analysis siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_lessonC

Start a persisted lesson in tactics, openings, middlegames, positioning, endgames, or calculation.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYes
learner_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it only hints at state creation via "persisted". It does not say whether the lesson requires a prior learner profile, whether it mutates existing state, what happens if a lesson is already active, or whether the call is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the verb and resource front-loaded and the topic enumeration following. Nothing is wasted, though the enumeration is long enough that the core action could be stated more sharply.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the topic domain is covered. However, for a stateful start tool with no annotations and 0% schema coverage, the description omits session lifecycle, prerequisites, and relationship to the many sibling analysis/lesson tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both parameters are bare strings, so the description must compensate. It usefully enumerates valid topic values (tactics, openings, middlegames, positioning, endgames, calculation), which is real added meaning, but learner_id remains completely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Start") and resource ("persisted lesson") and enumerates the topic domain, which lets an agent distinguish it from sibling utilities like advance_lesson or create_game_lessons. It is clear but does not explicitly differentiate itself from the closest sibling, start_teaching_session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and never names an alternative such as start_teaching_session, advance_lesson, or create_game_lessons. An agent must infer the selection criteria from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_progressive_analysisA

Return immediate preliminary evidence and queue deeper analysis; supersede older requests with the same context ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
historyNo
context_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
context_idYes
request_idYes
preliminaryYes
position_fenYes
schema_versionYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it does real work: it discloses that the call is non-blocking (immediate preliminary evidence plus queued deeper work) and that concurrent requests sharing a context_id supersede older ones. That supersede semantics is a behavioral trait not derivable from the schema. It stops short of stating auth, rate limits, or output shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the semicolon cleanly splits the immediate-return clause from the queuing/supersede clause. 'Preliminary evidence' is slightly vague about what is actually returned, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But with no annotations and 0% parameter coverage, the description leaves fen and history (including their relationship to one another) undocumented, which is a real gap for an analysis-start tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It gives context_id real meaning via the supersede rule, but fen and history are left entirely unexplained in both places. Partial compensation for a three-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific two-phase behavior: return immediate preliminary evidence and queue deeper analysis. This clearly separates it from the retrieval sibling get_progressive_analysis and from the job siblings (start_analysis_job). However it never names an alternative, so differentiation is by inference rather than explicit contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The supersede clause implies usage context (re-issuing with the same context_id invalidates the prior request), which tells the agent how to think about repeated calls. But there is no explicit when-to-use-versus-alternatives guidance, and the crowded analysis-start sibling family (start_analysis_job, analyze_position, evaluate_position) makes that omission meaningful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_teaching_sessionC

Start a solution-hidden teaching cycle with a concept-specific learning objective.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
historyNo
learner_idNodefault

Output Schema

ParametersJSON Schema
NameRequiredDescription
fenYes
statusYes
conceptYes
questionYes
hint_levelYes
session_idYes
schema_versionYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Solution-hidden' does disclose one meaningful trait (the answer is withheld from the learner), but nothing is said about persistence, learner state side effects, permissions, or what starting a cycle implies for subsequent calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is structurally sound. It is arguably too terse for a three-parameter session-start tool, but there is no wasted language.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but with zero annotations, 0% schema description coverage, and an extremely terse sentence, the definition is not complete enough for an agent to invoke this tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for three parameters (fen, history, learner_id), and the description names none of them. 'Concept-specific learning objective' hints at session semantics but does not clarify that fen is a chess position, or what history/learner_id control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a ... teaching cycle') and adds the concept-specific objective framing, so an agent understands the operation. However, it offers no differentiation from close siblings like start_lesson or advance_lesson, and the term 'solution-hidden' is unexplained jargon.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many adjacent teaching tools (start_lesson, create_tutor_turn, advance_lesson, get_teaching_hint). No preconditions or exclusions are stated, leaving the agent to infer intent entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_repertoire_moveC

Submit a legal move to a repertoire recall drill.

ParametersJSON Schema
NameRequiredDescriptionDefault
moveYes
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the move must be 'legal' but does not describe what happens on success or failure, whether the session must be active, what response is returned, or any side effects. While 'legal move' implies validation, critical behavioral context is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence that is front-loaded and contains no wasted words. It is concise but perhaps too terse given the complexity of the tool, though the structure itself is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has an output schema, the description need not explain return values. However, it still lacks essential context for using the tool correctly: no annotation coverage, no parameter documentation, and no guidance on interaction with sibling tools (e.g., create_repertoire_drill, get_repertoire_move). The description is incomplete for a mutation tool in a complex domain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning neither parameter has any description in the schema. The description does not explain the format of 'session_id' (e.g., how to obtain it) or the expected notation for 'move' (e.g., SAN, UCI). With two required parameters and no supporting documentation, the description fails to compensate for the lack of schema detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (submit) and resource (repertoire move), which gives a clear sense of the action. However, it is almost identical to sibling tools like submit_tactic_move and submit_teaching_move, and the description does not differentiate this tool from those or from get_repertoire_move. The purpose is understandable but lacks the differentiation needed to confidently choose among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as submit_tactic_move, submit_teaching_move, or get_repertoire_move. There is no mention of prerequisites, context (e.g., after starting a repertoire drill), or exclusions. The agent must infer usage from the name and context alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_tactic_moveC

Check a puzzle move, play the source reply, and schedule completed puzzles for review.

ParametersJSON Schema
NameRequiredDescriptionDefault
moveYes
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden — but it does disclose non-obvious side effects: it plays the source reply (mutates game position) and schedules completed puzzles for review (persistent learner state change). It omits permissions, idempotency, and error behavior, so it is useful but incomplete for a state-mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the main action front-loaded; no filler or repetition. It is efficient, though the three comma-separated clauses compress distinct behaviors with little structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a state-mutating tool with zero annotations and 0% parameter coverage, the description should at minimum define the move format and session requirements. Those gaps leave the definition under-specified for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so neither 'move' nor 'session_id' is documented anywhere. The description does not state the move notation (SAN/UCI), whether the session must be pre-created, or how an invalid move is handled, leaving an agent to guess the primary input format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource (check a puzzle move) and two follow-on effects, which is far more specific than a tautology. However, it never distinguishes itself from near-identical siblings like submit_repertoire_move or submit_teaching_move, so an agent must infer the domain split from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, prerequisites, or alternative selection guidance. The description does not explain why an agent would pick this over submit_repertoire_move or submit_teaching_move, which share the same submit-move shape.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_teaching_moveC

Assess an answer, explain verified candidate lines, update concept progress, and schedule review.

ParametersJSON Schema
NameRequiredDescriptionDefault
moveYes
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
due_atYes
assessmentYes
session_idYes
explanationYes
is_acceptedYes
next_questionYes
schema_versionYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses important side effects—updating concept progress and scheduling review—which indicate mutation, but omits details like permissions, error handling, reversibility, or what 'verified candidate lines' means operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded with the primary action and lists the rest efficiently, but the list of four comma-separated actions makes it slightly dense. It avoids filler, so it is concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool performs multiple state-changing operations and has two required parameters with no descriptive coverage, yet the description offers no detail on inputs or when to use it. The output schema covers return values, but invocation context and parameter meaning are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides no information about the two parameters (session_id, move), and the schema has 0% description coverage. An agent must guess the format and meaning of both required arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates four specific actions (assess, explain, update, schedule), giving a clear picture of what happens when the tool is called. It is distinct from sibling tools like submit_tactic_move by the 'teaching' context and mention of concept progress, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as submit_repertoire_move, submit_tactic_move, or tutor_move. The phrase 'Assess an answer' implies a teaching-session context, but no prerequisites or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

synthesize_speechB

Synthesize a short local WAV using espeak-ng. Returns a local file URI, not streamed audio.

ParametersJSON Schema
NameRequiredDescriptionDefault
rateNo
textYes
voiceNoen

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses two useful behavioral facts: output is a local file URI rather than streamed audio, and the audio is short. It omits file location/lifetime, overwrite behavior, voice availability, and whether the URI is fetchable by other tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the most important fact for an agent (a file URI comes back, not a stream) is front-loaded in the second sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details need not be repeated, and the description does note the URI form. However, for a 3-parameter tool with no annotations and 0% schema coverage, the definition is thin on parameter meaning and operational constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for three parameters (text, rate, voice), so the description must compensate and largely does not. Only the implicit 'short' qualifier hints at the text parameter; rate and voice are entirely undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (synthesize) and resource (speech, as a short local WAV via espeak-ng). No sibling tool does TTS, so differentiation is less critical, but the description usefully distinguishes itself from a streaming-audio tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'short local WAV' implies a use case and a size constraint, but there is no explicit when-to-use guidance, no prerequisites, and no mention of related siblings like get_speech_plan or create_tutor_turn that might precede or consume this output.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tutor_moveB

Assess a proposed legal SAN or UCI move. Returns the resulting position and Stockfish's best replies, for move-by-move tutoring.

ParametersJSON Schema
NameRequiredDescriptionDefault
fenYes
moveYes
nodesNo
multipvNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose that the tool returns the resulting position plus Stockfish's best replies (engine-backed analysis). It says nothing about cost, latency, or how nodes/multipv affect the depth of analysis, which matters for an engine call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the action and followed by the return behavior. No wasted words and the scope is immediately clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return-value sentence is somewhat redundant but harmless. The remaining gap is the undocumented engine parameters (nodes, multipv) and the absence of sibling disambiguation, which leaves the definition only marginally sufficient for a 4-parameter analysis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters. The description mentions SAN/UCI move format, which lightly informs the 'move' parameter, but 'nodes' and 'multipv' are engine-tuning parameters left completely unexplained in both schema and description, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource: 'Assess a proposed legal SAN or UCI move,' and adds the tutoring framing and engine-backed output. It does not, however, differentiate itself from close siblings like evaluate_move or assess_candidate_move, so an agent cannot confidently route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'for move-by-move tutoring' implies a teaching context, but there is no explicit when-to-use or when-not-to-use guidance and no naming of alternatives such as evaluate_move or assess_candidate_move. Usage is only weakly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_learner_profileC

Persist a learner's rating band, teaching preferences, and voice settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNosocratic
voiceNoen
detailNobrief
ratingNo
learner_idYes
speech_rateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it only says 'Persist'. It does not clarify whether this creates or updates, whether omitted optional fields reset to their schema defaults or remain unchanged (critical given every non-id param has a default), or what permissions are needed. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and scope come first. It is efficient, though arguably too terse to be genuinely useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a six-parameter mutation tool with no annotations and no parameter documentation the description is far too thin. The default-vs-unchanged question for omitted fields is the most important thing an agent needs and it is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only loosely groups the fields ('rating band', 'teaching preferences', 'voice settings') and never explains style, detail, voice, speech_rate, or the required learner_id, leaving five of six parameters without semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Persist') and resource (a learner's profile) and enumerates the field groups being written: rating band, teaching preferences, voice settings. An agent can tell it is a write tool for learner configuration, though nothing distinguishes it from nearby siblings like get_learner_progress or record_training_attempt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool, no preconditions, and no named alternative. With siblings such as get_learner_progress and record_training_attempt in the same family, the agent gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 58 tool updatesv0.1.0
    • First observedadvance_lesson
    • First observedanalyze_chessboard
    • First observedanalyze_game
    • First observedanalyze_pawn_structure
    • First observedanalyze_position
    • First observedanalyze_positions
    • First observedassess_candidate_move
    • First observedcancel_analysis_job
    • First observedcancel_progressive_analysis
    • First observedcompare_candidate_moves
    • First observedcreate_game_lessons
    • First observedcreate_repertoire_drill
    • First observedcreate_tutor_turn
    • First observedevaluate_endgame
    • First observedevaluate_move
    • First observedevaluate_position
    • First observedexplain_position
    • First observedfind_opening
    • First observedget_analysis
    • First observedget_analysis_job
    • First observedget_analysis_metrics
    • First observedget_capabilities
    • First observedget_concept_progress
    • First observedget_due_repertoire_drill
    • First observedget_learner_progress
    • First observedget_opening_coverage
    • First observedget_opening_plans
    • First observedget_opening_position
    • First observedget_progressive_analysis
    • First observedget_repertoire_move
    • First observedget_speech_plan
    • First observedget_tablebase_diagnostics
    • First observedget_tablebase_result
    • First observedget_tactic
    • First observedget_teaching_hint
    • First observedimport_opening_plan
    • First observedimport_opening_statistics
    • First observedimport_opening_tsv
    • First observedimport_puzzles
    • First observedimport_study
    • First observedinitialize_openings
    • First observedinspect_position
    • First observedlist_legal_moves
    • First observedparse_game
    • First observedrecognize_chessboard
    • First observedrecord_training_attempt
    • First observedresume_analysis_job
    • First observedreview_game
    • First observedstart_analysis_job
    • First observedstart_lesson
    • First observedstart_progressive_analysis
    • First observedstart_teaching_session
    • First observedsubmit_repertoire_move
    • First observedsubmit_tactic_move
    • First observedsubmit_teaching_move
    • First observedsynthesize_speech
    • First observedtutor_move
    • First observedupdate_learner_profile

TDQS

C2.7/5.0

Scored across 58 tools

Disambiguation2/5

Many tools overlap heavily: analyze_position, evaluate_position, analyze_positions, inspect_position, and explain_position all inspect or evaluate a FEN, while analyze_game, review_game, start_analysis_job, and get_analysis all analyze entire games. Move assessment is similarly fragmented across evaluate_move, compare_candidate_moves, tutor_move, assess_candidate_move, and submit_teaching_move. Descriptions attempt to differentiate, but the agent must read carefully to avoid picking the wrong tool.

Naming Consistency4/5

All tool names use snake_case and mostly follow a verb_noun pattern (e.g., create_game_lessons, get_opening_plans, start_progressive_analysis). Minor deviations exist, such as 'inspect_position' vs 'analyze_position' and 'tutor_move' vs 'evaluate_move', but the overall convention is predictable and readable.

Tool Count1/5

With 58 tools, the set is far beyond a reasonable scope, and many tools duplicate functionality (e.g., multiple analysis and evaluation endpoints). The domain is broad (engine analysis, openings, repertoire, tutoring, speech, tablebases), but the sheer number creates a heavy, confusing surface that is not well-scoped.

Completeness4/5

The surface covers a wide range of chess analysis and teaching workflows: engine evaluation, game review, openings, repertoire drills, tactics, tablebases, speech synthesis, and learner progress. Minor gaps include no deletion or update for some entities (e.g., no delete_repertoire, no update_game_lesson), but the core lifecycle operations are present.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables comprehensive chess analysis through Stockfish engine integration, positional evaluation, puzzle training, game review, and access to extensive chess databases. Provides visual board rendering, interactive game viewers, and tactical puzzle training with 3+ million problems from Lichess.
    31
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to perform professional-grade chess analysis using Stockfish and optionally Leela Chess Zero, including position analysis, full game review, opening lookup, and puzzle generation.
    15 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A powerful chess engine and game server built with the Model Context Protocol (MCP). Play chess against AI, analyze positions, and integrate chess functionality into your AI applications.
    13 npm
    1
    ISC