Skip to main content
Glama

ScoreCompute

Server Details

Independent Rust compute tools, typed capability planning and checked native composition.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

B3.4/5.0

Scored across 23 tools

Disambiguation3/5

Most tools target clearly different domains or inputs, but the scheduling/composition cluster (build_robust_plan, plan_mission, execute_composition, compose_capabilities, list_capabilities, list_composition_capabilities, simulate_plan) has overlapping boundaries that could lead to misselection. Shadow-related tools (check_shadow_track, verify_shadows) also overlap in purpose, though descriptions provide distinctions. The verbose descriptions help, but the set still requires careful reading.

Naming Consistency4/5

All names use snake_case, and most follow a clear verb_noun pattern (analyze_chess_game, audit_statistics, compute_orbit, simulate_pi). Minor deviations like network_status and route_optimizer are noun phrases rather than verb_noun, but the convention remains readable and largely predictable.

Tool Count3/5

23 tools is on the heavy side for a single MCP server, especially because several tools serve the same planning/composition family. The count is not extreme, but the set would benefit from splitting or consolidating related tools. It lands in the borderline-heavy range.

Completeness4/5

The server covers many specialized computational and verification workflows, including network job submission/status/retrieval, planning/execution for compositions, statistical audits, and multiple astronomy/geometry checks. Minor gaps exist, such as no cancel or list-own-jobs operation for network jobs, but agents can generally work around them.

Available Tools

23 tools
analyze_chess_gameA
Read-onlyIdempotent
Inspect

Compare supplied chess positions and played moves against Stockfish UCI: first-choice agreement and centipawn loss per position. Accept SAN or UCI moves. Returns a statistical profile, never a cheating verdict. The complete analysis has a 25-second deadline.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNo
positionsYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower; the description adds genuinely non-redundant traits: a 25-second execution deadline, the disclaimer that it produces a statistical profile rather than a verdict, and accepted move encodings. It omits failure behavior (e.g., what happens on timeout) but adds real value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, front-loaded with the core action and metrics before the qualifiers. No filler, though the first sentence is dense and packs metric definitions alongside the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly names the return shape (statistical profile with per-position metrics) and flags the time budget. The main remaining gap is the undocumented depth control, but for the essential call path the description is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 2 parameters, so the description must compensate. It partially does for the move format ('Accept SAN or UCI moves') and implies the positions input, but the depth parameter (range 4–20, default 12) and the FEN field format are entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (compare) and resource (supplied chess positions and played moves), names the engine (Stockfish UCI), and states the exact metrics computed: first-choice agreement and centipawn loss per position. An agent knows precisely what this does, and no sibling tool overlaps with the chess domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies clear context — accepted move formats (SAN or UCI), the output nature (statistical profile, never a cheating verdict), and an operational bound (25-second deadline). It stops short of explicit when-not-to-use guidance or naming an alternative, but no sibling is a plausible substitute, so little is left ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_kinematicsA
Read-onlyIdempotent
Inspect

Convert observed angular speed into minimum transverse speed at candidate distances using v = angular speed times distance. Compare illustrative balloon, drone, aircraft and satellite limits and check a claimed speed-distance pair. Makes no claim about object origin.

ParametersJSON Schema
NameRequiredDescriptionDefault
distance_max_kmNo
distance_min_kmNo
claimed_speed_kmhNo
claimed_distance_kmNo
angular_speed_deg_per_sYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the non-destructive, idempotent read profile, so the bar is lower. The description adds real behavioral context: the exact formula used, that the balloon/drone/aircraft/satellite limits are illustrative comparisons, and the important interpretive caveat that it makes no claim about object origin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core computation, no filler. Slightly dense packing of formula, comparison set, and caveat but all sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, no-output-schema computation tool, the description covers the operation, the comparison outputs, and the scope disclaimer. It omits what the distance sweep actually produces (tables? per-band speeds?) and unit expectations beyond the schema's implicit names, leaving the agent to infer the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does conceptually: it accounts for angular speed, candidate distance range, and claimed speed/distance. However it never explains that distance_min_km/distance_max_km define a scan sweep, their bounds, or how the claimed pair interacts with the range, leaving real gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific transformation (angular speed -> minimum transverse speed via v = ω·d) plus reference comparisons and a claimed-pair check. It clearly distinguishes itself from the sibling set of orbital/statistical/simulation tools, though it doesn't name a near-neighbor explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context of use is implied: an agent with an observed angular rate wanting a distance-constrained speed bound, or a claimed speed-distance pair to test. There is no explicit when-not-to-use or routing to an alternative, so guidance is present but only inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_statisticsB
Read-onlyIdempotent
Inspect

Apply necessary GRIM/GRIMMER arithmetic consistency checks to reported statistics for bounded integer observations. Examine rounded means and sample standard deviations. Passing the filters is not proof that a dataset exists; an inconsistency is not evidence of fraud.

ParametersJSON Schema
NameRequiredDescriptionDefault
nYes
sdNo
meanYes
decimalsNo
max_valueNo
min_valueNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare that this is a read-only, non-destructive, idempotent operation, so the safety profile is covered. The description adds meaningful behavioral context beyond the annotations by clarifying what the checks examine and by warning that passing is not proof of a dataset and inconsistency is not evidence of fraud.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loads the core operation, and avoids repetition. It is appropriately sized for the tool, though the abstract phrasing could be slightly more structured for an agent parsing required inputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters, no output schema, and zero schema description coverage, the description is incomplete on operational details. It provides an important interpretive warning but does not explain input requirements, parameter usage, or what a result consists of.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the six undocumented parameters, but it does not. It hints at means and sample standard deviations and bounded integer observations, but it never explains parameters such as n, decimals, min_value, or max_value, leaving most parameter meaning unspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Apply') and a specific resource ('GRIM/GRIMMER arithmetic consistency checks to reported statistics'), which is clear enough to distinguish this tool from unrelated siblings. It does not explicitly name or differentiate itself from any sibling tool, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the tool is used for reported statistics with bounded integer observations and rounded means/sample standard deviations, but it does not state when to prefer this tool over alternatives or when not to use it. The caveat about passing filters not proving a dataset exists adds interpretive context but not operational routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_robust_planB
Read-onlyIdempotent
Inspect

Compare up to three task selections in native Rust using optimize_tasks and, when needed, simulate_plan. Select the highest-value examined candidate whose simultaneous 95% Monte Carlo lower bound meets the requested success target under independent uniform duration uncertainty. Use an exact bound instead of simulation when possible and reuse identical selections within this mission. Return inconclusive if none passes, or blocked for unsupported contracts. Streams actual native events via MCP progress; one composite MCP call, no arbitrary tool generation.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
itemsYes
budgetYes
samplesNo
uncertaintyYes
allowed_toolsNo
duration_modelNoindependent_uniform
max_candidatesNo
min_success_rateNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, non-destructive, closed-world behavior. The description adds meaningful traits beyond that: it streams native events via MCP progress, is a single composite call, performs no arbitrary tool generation, and returns distinct terminal statuses (inconclusive, blocked). That is genuine extra behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core action and avoids padding, but the pivotal sentence is a dense stack of statistical qualifiers ('simultaneous 95% Monte Carlo lower bound meets the requested success target under independent uniform duration uncertainty') that is hard to parse at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-param, no-annotated-output, composite orchestration tool, the description explains the algorithm and the result statuses but leaves the parameter surface almost entirely uncovered and gives no sense of the return shape or event payloads. Adequate for orientation, not for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are 9 parameters, so the description carries the full burden. It gestures at the statistical model (95% Monte Carlo lower bound, independent uniform duration uncertainty, success target) and the 'up to three' candidate cap, but never explains seed, budget, samples, allowed_tools, or duration_model by name, leaving most parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete action (compare up to three task selections, select the highest-value candidate meeting a success target) and names the specific sibling tools it orchestrates (optimize_tasks, simulate_plan). It clearly conveys what the tool produces (a robust plan or an inconclusive/blocked status), though it never explicitly contrasts itself with plan_mission, another planning sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives useful internal guidance — 'use an exact bound instead of simulation when possible', reuse identical selections, and the fallback conditions for inconclusive/blocked. However, it does not say when an agent should choose build_robust_plan over optimize_tasks, simulate_plan, or plan_mission directly, which is the core routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_ephemerisB
Read-onlyIdempotent
Inspect

Check a claimed Sun or Moon position for a UTC date and location. Uses NOAA/Meeus approximations, topocentric lunar parallax, angular separation and lunar phase. The object parameter uses soleil (Sun) or lune (Moon). Returns compatible, contredit or intestable.

ParametersJSON Schema
NameRequiredDescriptionDefault
objectNolune
date_utcYes
latitudeYes
longitudeYes
tolerance_degNo
claimed_azimuth_degYes
claimed_elevation_degYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuine context beyond them: the approximation model (NOAA/Meeus), topocentric lunar parallax, angular separation, lunar phase, and the three verdict outcomes (compatible, contredit, intestable).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core action, then method, then the object enum and return values. No filler, though the enum explanation is partly redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does list the three verdict strings, plus the computational approach, which helps. However, for a 7-parameter tool with 0% schema coverage it omits the date_utc format and tolerance behavior, leaving key invocation details to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters. The description only clarifies the object enum values (already an enum in the schema) and never explains the required date_utc format (fixed 20-char string), the role of tolerance_deg, or the units/meaning of claimed_azimuth_deg and claimed_elevation_deg.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: checking a claimed Sun/Moon position for a UTC date and location, with the method (NOAA/Meeus) and verdict outputs. This distinguishes it from shadow-oriented siblings like verify_shadows or check_shadow_track, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'check a claimed ... position' – i.e., validate a human-provided azimuth/elevation against computed ephemeris. There is no explicit when-to-use/when-not guidance or reference to the sibling tools that might overlap (verify_shadows, check_shadow_track).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_shadow_trackB
Read-onlyIdempotent
Inspect

Compare measured shadow azimuths over time with the expected solar ephemeris. Check slope, time-lapse speed, rotation direction and individual residuals. Inputs are measured minute/azimuth pairs, not video files. Returns compatible, contredit or intestable.

ParametersJSON Schema
NameRequiredDescriptionDefault
samplesYes
date_utcYes
latitudeYes
longitudeYes
tolerance_degNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the full safety profile (readOnly, idempotent, non-destructive, closed-world), so the description is free to add substance, and it does: the specific phenomena it evaluates and the three possible verdicts it returns. It stops short of explaining what each verdict means or how tolerance affects the outcome.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core comparison and ending with the return vocabulary; no filler. Minor friction from the non-standard verdict spellings ('contredit', 'intestable') that force the reader to interpret rather than recognize.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% schema coverage, the description is the only source of meaning, and it delivers the return-value vocabulary but not their semantics. Combined with the undocumented tolerance parameter, an agent has enough to call the tool but not enough to interpret or tune it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden. 'Measured minute/azimuth pairs' partially documents the samples array, but date_utc, latitude, longitude, and especially the tolerance_deg threshold are left completely unexplained, leaving four of five parameters to be reverse-engineered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('compare measured shadow azimuths over time with the expected solar ephemeris') and enumerates the checks performed (slope, time-lapse speed, rotation direction, residuals). It is concrete enough to act on, but it never names or differentiates itself from the plausible siblings check_ephemeris and verify_shadows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'Inputs are measured minute/azimuth pairs, not video files' is a useful precondition on the expected data source. However, it gives no explicit when-to-use guidance relative to check_ephemeris or verify_shadows, so selection among the shadow/ephemeris siblings remains inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_capabilitiesA
Read-onlyIdempotent
Inspect

Let native Rust compose a dependency plan from goal contracts and declared available contracts. Enforce allowed tools, locality, step and search limits. Return the selected graph, examined alternatives, blockers, search completeness and any existing-tool execution binding. Structural cost counts registered graph steps; it is not measured runtime, financial cost or scientific quality. Available facts are declarations; actual arguments are validated at execution. No engines run during planning.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalsYes
inputsNo
adaptersNo
availableYes
constraintsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior: what the result contains (selected graph, examined alternatives, blockers, search completeness, execution binding), the precise meaning of 'structural cost' (registered graph steps, not runtime/financial/scientific cost), and the caveat that declared facts are validated only at execution. It does not discuss determinism, limits behavior, or failure modes in detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Six sentences, front-loaded with the core purpose, then constraints, then return shape, then two clarifying caveats. Every sentence carries information, though the density is high and could be tightened slightly. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex planner with nested objects, 0% schema coverage and no output schema, the description does valuable work by enumerating return fields and clarifying cost and validation semantics. The main gap is parameter-level detail (inputs, adapters, constraint fields) that neither schema nor description explains, which matters for a tool with 5 parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It maps loosely onto the constraints object ('Enforce allowed tools, locality, step and search limits') and mentions goals and available contracts by concept, but never explains inputs, adapters (route_durations.v1), or the max_structural_cost/max_search_states semantics. It partially compensates but leaves several nested parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: composing a dependency plan from goal contracts and declared available contracts, with enforcement of tool/locality/step/search limits. It clearly distinguishes planning from execution ('No engines run during planning'). However, it never names or contrasts with the closest siblings such as build_robust_plan, execute_composition, or plan_mission, so the agent must infer differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the tool is a planning-only step, and the note that no engines run and that arguments are validated at execution hints that execution is handled elsewhere. There is no explicit 'use this when / do not use this when' guidance, nor any reference to alternative planning or execution siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_orbitA
Read-onlyIdempotent
Inspect

Compute an ideal elliptical orbit around one solar mass. Time is measured from periapsis and distances are in astronomical units. This two-body model is not a real celestial ephemeris.

ParametersJSON Schema
NameRequiredDescriptionDefault
eccentricityYes
elapsed_daysYes
semi_major_auYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a safe, read-only, idempotent, closed-world computation. The description adds genuine behavioral context beyond that: the two-body idealization, that time is measured from periapsis, that distances are in AU, and an explicit scope limitation. It does not, however, describe return values or accuracy bounds.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, with the core action front-loaded and the disambiguating caveat placed last. Every sentence carries distinct information (purpose, units/frame, scope limit).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameter descriptions, the description should carry more burden. It covers the physical setup and units but never says what the computation returns (position, velocity, orbital elements) or how results should be interpreted, which is a real gap for an agent invoking it blind.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It supplies units and reference frame (AU for distances, time from periapsis) that the bare schema lacks, covering elapsed_days and semi_major_au implicitly. It says nothing about the eccentricity parameter or the meaning of the numeric bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ("Compute an ideal elliptical orbit") plus the physical model (one solar mass, two-body). It implicitly separates itself from check_ephemeris by declaring it is not a real ephemeris, but it does not name or contrast with any sibling tool directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing caveat ("not a real celestial ephemeris") implies when not to rely on this tool for real-world ephemeris work, which routes the agent away from it. However, there is no explicit when-to-use statement or named alternative, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_compositionA
Read-onlyIdempotent
Inspect

Recompile and execute an approved optimize_tasks → selected_durations.v1 → simulate_plan or route_optimizer → route_durations.v1 → simulate_plan path in Rust. Supply the engine arguments plus simulate_plan configuration (budget, uncertainty, samples, seed; omit durations). The route path also requires adapters["route_durations.v1"].speed_kmh from 1 to 150, explicit independent-uniform assumptions and 1–32 nonzero route legs. Whole-minute durations use ceil(distance_km / speed_kmh * 60). Shared 1–1440 minute budget, each duration 1–1440 minutes, 100–20000 fixed samples. Geometric distance and a declared speed are not road or traffic predictions. Return the engine result, a model-conditional risk assessment, a 95% Hoeffding sampling interval and real native events. Completed means the assessment ran, not that the deadline is safe. Other graph paths must use their advertised existing tool or remain planning-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalsYes
inputsNo
adaptersNo
availableYes
constraintsNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/no-open-world, so the safety profile is covered; the description goes further by disclosing inputs to omit (durations), numeric operating envelopes (1–150 km/h, 1–1440 min budget, 100–20000 samples, 1–32 legs), the recompute-and-execute behavior in Rust, and the important caveat that 'Completed means the assessment ran, not that the deadline is safe'. It does not address failures, compile cost, or runtime, but adds substantive context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence and each subsequent sentence carries non-redundant information (usage constraint, route-path requirements, ceil formula, numeric bounds, disclaimers, return shape). It is dense and jargon-heavy for a single paragraph, but there is little filler to remove given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex graph-execution tool with no output schema and 0% schema coverage, the description supplies the return shape (engine result, risk assessment, 95% Hoeffding interval, native events), the two supported graph paths, and their input constraints. It is largely self-sufficient, with the main remaining gap being the meaning of the goals/available/constraints inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden. It does explain adapters['route_durations.v1'].speed_kmh (1–150), the simulate_plan configuration knobs (budget, uncertainty, samples, seed) and the requirement to omit durations, but it never explains what goes in goals, available, or the constraints block (max_steps, local_only, allowed_tools, max_search_states, max_structural_cost), leaving several top-level parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action (recompile and execute) on a specific resource (the approved optimize_tasks → selected_durations.v1 → simulate_plan / route_optimizer → route_durations.v1 → simulate_plan graph path) in Rust. It also explicitly distinguishes itself from sibling tools by naming simulate_plan, optimize_tasks and route_optimizer and stating that other graph paths must use their advertised tool or stay planning-only.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear usage condition (only 'approved' graph paths) and a real exclusion (other graph paths must use their existing tool or remain planning-only), which routes the agent away from misuse. It does not define what 'approved' means or which sibling to call for the unsupported cases, so it stops short of explicit alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_network_jobB
Read-onlyIdempotent
Inspect

Read one contributor job ticket and its verified progress or final pi_chunk_v1 receipt. Only completed jobs contain a final estimate. Counts aggregate each verified shard once, with expired leases rejected. The publicly visible job contains only synthetic sampling parameters, never private user data.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower, yet the description adds real context: counts de-duplicate each verified shard with expired leases rejected, final estimates exist only for completed jobs, and the payload exposes synthetic sampling parameters only. That privacy and aggregation detail meaningfully exceeds the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action before qualifications. No sentence is filler, though domain jargon like 'pi_chunk_v1' slightly taxes readability without adding clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does describe what comes back: verified progress, a final receipt, a completed-only estimate, and de-duplicated counts. For a single-parameter read tool this is largely sufficient, missing only guidance on retrieving the job_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single job_id parameter has 0% schema description coverage, so the schema supplies only type, format, and a UUID pattern. The description alludes to a 'job ticket' but never explains what job_id identifies or where the caller obtains it, leaving the parameter's meaning largely undocumented in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: read one contributor job ticket plus its verified progress or final receipt. This clearly distinguishes it from the write-oriented submit_network_job and from aggregate status tools. It stops short of naming a sibling explicitly, so the differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance or routing to alternatives. The note that 'only completed jobs contain a final estimate' tells the agent something about state but not about when to call this versus submit_network_job or network_status. Usage must be inferred from the verb 'read'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_capabilitiesC
Read-onlyIdempotent
Inspect

Inspect the thirteen native composition-bank tool contracts (triangulate_tracks is independently callable but not yet in this bank), themes, units, assumptions, execution backends and composition readiness. Optional exact theme filter. The scheduling family compares task selections; approved selection_risk and route_risk chains use explicit adapters and simulate_plan. All tools remain independently callable.

ParametersJSON Schema
NameRequiredDescriptionDefault
themeNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description's genuinely additive facts are that the bank contains thirteen contracts, that triangulate_tracks is independently callable but excluded from this bank, and that all tools remain independently callable — useful scope boundaries, though no auth, rate-limit, or return-shape detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening clause is front-loaded and the theme filter follows, which is good. But the trailing sentences about the scheduling family, approved selection_risk/route_risk chains, explicit adapters, and simulate_plan read as tangential domain chatter that a caller of a read-only listing tool does not need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must convey what comes back; it does name the categories (contracts, themes, units, assumptions, backends, readiness) but does not describe the response shape, count, or ordering. For a small, single-parameter listing tool this is only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the lone theme parameter, but the description compensates partially by labeling it an "exact theme filter" and marking it optional, which conveys matching semantics the bare schema does not. It still gives no format, examples, or behavior on a non-matching theme, so it falls short of full compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description does state a verb ("Inspect") and a resource ("the thirteen native composition-bank tool contracts... themes, units, assumptions, execution backends and composition readiness"), so the general intent is legible. However, it never distinguishes itself from the sibling list_composition_capabilities, which appears to cover the same territory, and the enumerated scope is dense internal jargon rather than a plain statement of what a caller gets back.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Optional exact theme filter" is parameter detail, not usage guidance. The sentences about the scheduling family, risk chains, and simulate_plan describe the domain rather than telling the agent when this tool is the right choice versus list_composition_capabilities or compose_capabilities, and no exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_composition_capabilitiesB
Read-onlyIdempotent
Inspect

Read the trusted native bank of versioned semantic contracts, engine capabilities, explicit adapters, composite capabilities and reproducible planning examples. Each individual MCP engine remains independently callable. This catalogue contains implemented capabilities only; unsupported model/unit facts explain missing adapters. No computation is executed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that no computation is executed and that only implemented capabilities are listed, which is useful context. However, it doesn't disclose the return format, pagination, or size of the catalogue.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, somewhat dense and jargon-heavy. It front-loads the main purpose but includes abstract phrases that could be simplified. It's not excessively long but could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should explain what the tool returns. It mentions 'catalogue' but doesn't describe the structure or format of the returned data. For a read-only listing tool, this is a moderate gap, though annotations cover safety aspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters (parameter count 0), so per the rules the baseline is 4. The description adds no parameter information, but none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it reads a catalogue of capabilities, which is a specific verb and resource. However, the terminology is abstract and domain-specific ('trusted native bank of versioned semantic contracts'), making it hard to quickly understand what the tool returns. It doesn't clearly differentiate from sibling list_capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like list_capabilities or compose_capabilities. The description implies it's for discovering capabilities but doesn't specify conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_statusA
Read-onlyIdempotent
Inspect

Read the live contributor pilot: available and busy GPUs, explicitly offered CPU threads, stale workers and observed verified pi_chunk_v1 throughput. Hardware is reported, not attested. VRAM is separate per GPU, not a shared memory pool. This does not add the legacy worker inventory a second time.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it read-only, idempotent, and non-destructive, so the safety profile is covered. Beyond that, the description adds real behavioral data: hardware is 'reported, not attested' (a provenance caveat), VRAM is per-GPU rather than a shared pool, and it avoids double-counting legacy inventory. Those are non-obvious traits an agent cannot get from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first clause and packs the qualifiers into short sentences. Every sentence carries meaning, though the final line about the legacy worker inventory is somewhat cryptic and could confuse rather than inform.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must characterize the return content, and it does enumerate the reported fields. Combined with annotations covering the safety profile, an agent has enough to call and interpret this correctly, though the unexplained 'pi_chunk_v1' and legacy-inventory references leave minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to convey and the baseline of 4 applies. The description correctly does not invent parameter detail and instead spends its budget on what the call returns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('live contributor pilot') then enumerates exactly what is reported: available/busy GPUs, offered CPU threads, stale workers, and pi_chunk_v1 throughput. The purpose is clear, though no sibling tool is named for differentiation, and the reference to a 'legacy worker inventory' is not anchored to any listed sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The read-only framing and the closing exclusion ('does not add the legacy worker inventory a second time') imply when this tool is appropriate, but no explicit when-to-use condition or named alternative (e.g. get_network_job) is given. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_tasksA
Read-onlyIdempotent
Inspect

Select independent tasks exactly to maximize their total value within an integer-minute time budget. Each task can be selected once; at most 32 tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
budgetYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so safety is covered. The description adds real behavioral context worth having: each task may be selected at most once and the input is capped at 32 tasks. It does not state that the result is an exact optimum vs. approximation, nor anything about the returned structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, objective stated first, then the two governing constraints. No filler and nothing repeated from structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only solver with no output schema, the description never indicates what is returned (e.g., the chosen task set and total value), leaving the agent to guess the response shape. Objective and constraints are complete, but the result contract is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: 'budget' is framed as integer minutes and 'items' as tasks with name/duration/value that can each be selected once, capped at 32. The nested task fields' exact semantics are left to the schema, which is acceptable given their self-evident names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (select) and resource (tasks) plus the exact objective: maximize total value within an integer-minute budget. This is a precise combinatorial optimization description that an agent can act on immediately, and it is clearly distinct from the unrelated siblings (chess, orbits, markets).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The budget-maximization framing implies when to use it, but there is no explicit when-to-use/when-not statement and no named alternative. Since no sibling is a plausible substitute, the practical risk is low, but the guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_missionB
Read-onlyIdempotent
Inspect

Preview a bounded robust scheduling mission without executing engines. Validate explicit duration assumptions and local tool permissions, then return up to three candidate budgets (nominal, 10% reserve, 20% reserve), bounds on engine calls and samples, or explicit blockers. Does not generate arbitrary tools or execute external code.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
itemsYes
budgetYes
samplesNo
uncertaintyYes
allowed_toolsNo
duration_modelNoindependent_uniform
max_candidatesNo
min_success_rateNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety burden is lifted. The description still adds real behavioral context beyond annotations: it validates duration assumptions and local tool permissions, caps output at three candidates, and can return explicit blockers rather than failing. It does not describe cost, latency, or determinism of the preview.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and immediately followed by what is returned. Terminology is dense but each clause carries information; no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, no-output-schema tool the description reasonably conveys output shape and the non-execution guarantee, so an agent can call it without guessing return format. The large parameter-semantics gap (0% coverage) keeps it from being fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across nine parameters, so the description must carry the load. It only obliquely gestures at a few inputs ('duration assumptions', 'local tool permissions', 'three candidate budgets'), leaving items, budget, seed, samples, uncertainty, duration_model, and min_success_rate entirely unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Preview a bounded robust scheduling mission') and enumerates the artifacts returned (three candidate budgets, engine-call/sample bounds, or blockers). It never names or contrasts with the obvious siblings (build_robust_plan, optimize_tasks, simulate_plan), so an agent still has to infer which related tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without executing engines' and 'Does not generate arbitrary tools or execute external code' imply this is the non-executing preview step, which is useful context. However, no explicit when-to-use rule or named alternative (e.g., build_robust_plan or execute_composition) is given, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

route_optimizerA
Read-onlyIdempotent
Inspect

Order 2–256 supplied points from fixed start index 0, with an optional return. Explicit planar kilometres or great-circle kilometres on a 6371.0088 km sphere. CPU Held–Karp certifies the numerical optimum only after complete search for at most 13 points; larger instances use bounded nearest-neighbour and 2-opt heuristics. Returns every leg, exact flag, stopping reason, measured time and work. Geometric distances only: no roads, traffic or travel-time forecasts. Callable independently of the orchestrator.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricYes
pointsYes
time_budget_msNo
return_to_startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
unitYes
exactYes
orderYes
methodYes
metricYes
pointsYes
tour_legsYes
elapsed_msYes
provenanceYes
work_limitYes
work_unitsYes
limitationsYes
stop_reasonYes
upper_bound_kmYes
return_to_startYes
contract_versionYes
total_distance_kmYes
greedy_reference_kmYes
improvement_vs_greedy_pctYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld=false/destructive=false, so the safety profile is covered; the description goes further by disclosing the exact-vs-heuristic cutover, that it returns an exact flag and stopping reason, and a hard limitation ('geometric distances only: no roads, traffic or travel-time forecasts'). This is meaningful context beyond the annotations, though it does not quantify the time_budget effect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core ordering constraint is front-loaded, followed by metric semantics, algorithm behaviour, and output content in a logical order. It is dense but each sentence carries distinct information; the only mild excess is restating return fields already covered by the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with an output schema, the description covers inputs (metric, point count, return flag), algorithmic guarantees/limits, and domain restrictions, which is nearly everything needed to call it correctly. The unaddressed time_budget_ms semantics is the one remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It defines the metric semantics precisely (planar kilometres vs great-circle on a 6371.0088 km sphere) and clarifies the point-count range and fixed start index. It does not explicitly explain time_budget_ms, leaving one of four parameters undocumented beyond its schema default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a precise verb+resource+scope: 'Order 2–256 supplied points from fixed start index 0, with an optional return.' No sibling tool (optimize_tasks, build_robust_plan, plan_mission) does geometric point ordering, so the agent can distinguish it immediately without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the algorithmic selection rule (exact Held–Karp at ≤13 points, heuristics above) and that it is 'callable independently of the orchestrator,' which implies usage context. However, it never states when to prefer this over sibling tools like optimize_tasks or build_robust_plan, nor any when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screen_wash_tradingA
Read-onlyIdempotent
Inspect

Screen a supplied trade series for repeated volume at nearly identical prices within a bounded time window. Report recycled-volume ratio and cycle count. This is a pattern to examine, not proof of wash trading or an accusation.

ParametersJSON Schema
NameRequiredDescriptionDefault
tradesYes
window_sNo
price_band_pctNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds real interpretive value the annotations cannot: it specifies the metrics returned and explicitly warns the result is a pattern to examine, not proof or an accusation — important for how an agent should phrase its findings. It does not, however, disclose input-size limits or cost/performance traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the detection logic before the reported outputs and the interpretive caveat. Every sentence earns its place and nothing is redundant with the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly names the return values (recycled-volume ratio, cycle count), and the 'not proof' framing is genuinely useful. It omits the input constraints (20-5000 trades) that would prevent call failures, which is the only meaningful gap for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden, yet it never names window_s or price_band_pct or explains their units, defaults, or bounds. 'Within a bounded time window' and 'nearly identical prices' gesture at the concepts but give no mapping to parameters, leaving two of three inputs semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (screen) and a specific resource (a supplied trade series), and states precisely the pattern sought: repeated volume at nearly identical prices within a bounded window. It also names the reported outputs (recycled-volume ratio, cycle count), so an agent knows exactly what the tool does without opening the schema. None of the siblings overlap in domain, so no differentiation is needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'screen a supplied trade series' implies the input context (you must already have a trade series to hand), but there is no explicit when-to-use, when-not-to-use, or alternative tool guidance. The closing caveat clarifies interpretation, not invocation. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_piA
Read-onlyIdempotent
Inspect

Estimate pi with reproducible Monte Carlo sampling and a 95% Wilson interval. Select CPU or CUDA; at most 5 million points. CUDA acceleration is available for this tool only.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedYes
backendNocpu
samplesYes
device_idsNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/non-openWorld, so the burden is light. The description still adds real behavioral context beyond them: results are reproducible via seed, the output includes a 95% Wilson interval, the sample cap is 5 million, and CUDA is the only accelerated path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with the core purpose front-loaded and no filler. The second sentence packs backend choice and the sample cap efficiently, though the semicolon-joined clause is slightly compressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the returned estimate and Wilson interval, and it covers the reproducibility and cap constraints. It is still incomplete for a 4-parameter tool: device_ids is undocumented and the required/optional status of seed and samples is left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it partially does: seed implies reproducibility and the backend enum values are called out, with the 5M cap restating the samples maximum. The device_ids array is entirely unexplained (how many GPUs, and whether it only matters under CUDA), leaving a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ("Estimate pi") plus the algorithm and statistical method, so an agent knows exactly what it computes. It does not differentiate against siblings, but the sibling list (chess, orbits, shadows, trading) is unrelated enough that no differentiation is needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Select CPU or CUDA" plus "CUDA acceleration is available for this tool only" gives a backend-choice heuristic and a relative capability note versus siblings. However, there is no explicit when-to-use/when-not guidance beyond backend selection, and nothing about when a Monte Carlo pi estimate is the appropriate call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_planB
Read-onlyIdempotent
Inspect

Test a task plan, including one produced by optimize_tasks, using independent bounded uniform duration variations. Return the probability of meeting the budget, mean duration and empirical 95th percentile. This illustrative uncertainty model is not a forecast guarantee.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedYes
budgetYes
samplesYes
durationsYes
uncertaintyYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the burden is lower. The description adds genuinely useful behavioral context beyond that: the simulation model (independent bounded uniform variations on durations) and an explicit caveat that the result is illustrative rather than a forecast guarantee. It does not, however, mention that output is deterministic given the seed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler, with the action and mechanism front-loaded and the output list and caveat following. The caveat sentence earns its place by clarifying model limitations, though the wording is slightly hedged and could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description usefully enumerates the three returned values. However, with five required, undocumented parameters and no output schema, an agent still lacks the parameter semantics and output format details needed to invoke this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across five required parameters, so the description carries the full semantic burden and fails to do so. It never explains what 'uncertainty' means (a fraction? absolute bound?), what 'samples' controls, or how 'seed' affects reproducibility, even though 'durations' and 'budget' are at least inferable from context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Test a task plan') and the mechanism ('independent bounded uniform duration variations'), and it names the three quantities returned. It also explicitly ties itself to the sibling optimize_tasks as a source of plans, which aids disambiguation. It stops short of a fully crisp one-line scope statement, but an agent can tell what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentioning that the plan may come from optimize_tasks implies the natural workflow (optimize, then simulate), but there is no explicit when-to-use or when-not-to-use guidance and no comparison to alternatives. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

solve_equilibriumB
Read-onlyIdempotent
Inspect

Solve Kuhn poker with CFR+ self-play and measure exploitability using pure best responses. Compare the game value against the analytical reference -1/18. An illustrative game-theory solver, not gambling advice.

ParametersJSON Schema
NameRequiredDescriptionDefault
iterationsNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds real methodological context (CFR+ self-play, exploitability measurement, reference value -1/18) but omits the significant computational cost implied by allowing up to 5,000,000 training iterations, which is the main behavioral caveat an agent would want to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and method, with the reference value in the next line. The disclaimer sentence is arguably marginal but short and earns its place as a misuse guardrail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully signals what a result contains (exploitability and game value against -1/18), and annotations cover safety. However it leaves the sole tunable parameter and the runtime implications of large iteration counts entirely unexplained, which is a meaningful gap for a compute-heavy solver.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'iterations' has 0% schema description coverage, so the description carries the full burden of explaining it and does not mention it at all. An agent gets no guidance on the default (200000), the 1000-5000000 range, or how iteration count trades off against accuracy and runtime.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource (solve Kuhn poker equilibrium) and specifies the method (CFR+ self-play) and the measured quantity (exploitability via pure best responses). No sibling tool covers game-theoretic solving, so an agent can select this unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, nor any prerequisite or context for invoking it. The closing 'not gambling advice' line is a disclaimer, not usage guidance, and the illustrative/solver framing is implied at best.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_network_jobAInspect

Queue a public synthetic Monte Carlo pi_chunk_v1 job on consenting contributor clients. This is the fixed 31-bit integer model, distinct from simulate_pi. No files or private datasets are accepted. Returns a ticket: queued is not completed. Poll get_network_job for the result. Every accepted shard is recomputed on coordinator CPU for verification; no net acceleration claim. Workers may be unavailable and jobs may expire.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
backendNoauto
samplesYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: returns a ticket that does not mean completion, requires polling via get_network_job, every accepted shard is recomputed on coordinator CPU with no net acceleration claim, and workers may be unavailable or jobs may expire. This is unusually rich disclosure for an async distributed job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Six dense sentences, all earning their place, with the core action and the simulate_pi distinction front-loaded and operational caveats trailing. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains the return value (a ticket, queued != completed) and how to retrieve results. Combined with the constraint and caveats, an agent has enough to call and follow through correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters, yet it says nothing about seed, backend, or samples. The names are self-explanatory and the schema carries min/max bounds, defaults, and an enum, but the description adds no semantic value on any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Queue) and resource (public synthetic Monte Carlo pi_chunk_v1 job on contributor clients). It explicitly distinguishes itself from the sibling simulate_pi and notes the fixed 31-bit integer model, so an agent can tell the two apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (simulate_pi) and the follow-up tool (get_network_job) for polling, and states an exclusion ('No files or private datasets are accepted'). It gives clear context but never states the decision condition for choosing this over simulate_pi.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

triangulate_tracksA
Read-onlyIdempotent
Inspect

Estimate a single moving object’s 3D position and velocity from geometric azimuth/elevation tracks supplied by 2–4 stationary observers. Bounded Rust CPU weighted least squares assumes constant velocity in a spherical Earth-fixed frame (R=6371.0088 km), declared angular standard deviations and supplied fixed clock offsets. Returns estimated, inconsistent, degenerate or no_convergence, component standard deviations, trajectory and observer residuals. No refraction correction or automatic image analysis; a reduced chi-square cutoff of 9 is a screening convention. Local uncertainties are conditional on the model, not calibrated real-world accuracy. Does not identify the object or its origin. Independently callable; not yet registered in the native composition bank.

ParametersJSON Schema
NameRequiredDescriptionDefault
observersYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
gridYes
stateYes
statusYes
reasonsYes
observersYes
residualsYes
elapsed_msYes
provenanceYes
trajectoryYes
baseline_kmYes
diagnosticsYes
limitationsYes
contract_versionYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive safety, and the description goes well beyond them: the estimator is bounded Rust CPU weighted least squares, assumes constant velocity in a spherical Earth-fixed frame with R=6371.0088 km, applies a reduced chi-square cutoff of 9, and emits estimated/inconsistent/degenerate/no_convergence statuses with residuals. It also honestly caveats that local uncertainties are conditional on the model, not calibrated real-world accuracy.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense, front-loaded and information-rich: purpose, method, assumptions, return statuses, limitations, scope, and integration status all appear in order. It runs long and reads as a block of clauses, but nearly every clause conveys distinct signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex kinematic estimator, the description supplies the model assumptions, convergence outcomes, and accuracy caveats an agent needs. An output schema exists, so the explicit return-value enumeration is a bonus rather than a necessity, and nothing critical to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported as 0% and there is effectively one top-level parameter with deep nesting, so the description must compensate. It does add meaning by naming angular standard deviations and fixed clock offsets and restating the 2–4 observer range, but it does not explain the per-sample or per-observer fields the caller must populate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Estimate a single moving object's 3D position and velocity from geometric azimuth/elevation tracks'. The observer-count constraint (2–4 stationary observers) and the fact it estimates kinematics distinguish it from siblings like compute_orbit, analyze_kinematics, and check_ephemeris.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context (stationary observers, 2–4 stations, angle tracks) and exclusions ('No refraction correction or automatic image analysis'), plus integration status ('independently callable; not yet registered in the native composition bank'). It does not, however, explicitly route the agent to a sibling tool when those conditions are unmet.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_shadowsA
Read-onlyIdempotent
Inspect

Check solar-shadow consistency from supplied image measurements: NOAA solar position, ground-plane homography from 4 to 16 control points, and Monte Carlo uncertainty. Returns compatible, contredit or intestable under the stated assumptions; never an image-authenticity verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
shadowYes
samplesNo
date_utcYes
latitudeYes
longitudeYes
pixel_sigmaNo
second_shadowNo
control_pointsYes
time_sigma_minutesNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered. The description adds meaningful behavior beyond that: the method (NOAA position, homography, Monte Carlo) and the three-valued verdict vocabulary with the caveat that results are assumption-dependent and not an authenticity judgment. A note on seed-driven reproducibility vs. the idempotentHint would have completed it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and its inputs, then the return contract and boundary. Dense but no filler; a typo ('intestable') and heavy clause packing are the only minor costs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, nested-object tool with no output schema, the description usefully declares the return values, which is what the missing output schema would otherwise supply. But it leaves most parameters undocumented and gives no guidance on coordinate formats, sampling/seed semantics, or how control-point count affects reliability, so the definition is only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, including nested structures, so the description carries the full explanatory burden — and it only alludes to control points (4–16) and the shadow geometry. Key inputs such as pixel_sigma, samples, seed, time_sigma_minutes, second_shadow, and the px/world coordinate convention go entirely unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check solar-shadow consistency') and names the exact inputs it consumes: NOAA solar position, a ground-plane homography from 4–16 control points, and Monte Carlo uncertainty. It even specifies the categorical outputs and the scope boundary, so an agent can distinguish it from siblings like check_shadow_track without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly frames when the tool applies (image measurements in hand) and explicitly excludes one use case ('never an image-authenticity verdict'), which is valuable negative guidance. It does not, however, route the agent to or away from specific sibling tools such as check_shadow_track, so sibling selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • Addedget_network_job
    • Addednetwork_status
    • Addedsubmit_network_job
    • Addedtriangulate_tracks
  2. 5 tool updates
    • Changedbuild_robust_plan1 field changed
      • changedInput schema / properties / allowed_tools / maxItems
        Previous value: -12New value: +13
    • Changedcompose_capabilities2 fields changed
      • addedInput schema / properties / adapters
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "route_durations.v1": {
        +      "additionalProperties": false,
        +      "properties": {
        +        "speed_kmh": {
        +          "maximum": 150,
        +          "minimum": 1,
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "speed_kmh"
        +      ],
        +      "type": "object"
        +    }
        +  },
        +  "type": "object"
        +}
      • changedInput schema / properties / constraints / properties / allowed_tools / maxItems
        Previous value: -13New value: +14
    • Changedexecute_composition2 fields changed
      • addedInput schema / properties / adapters
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "route_durations.v1": {
        +      "additionalProperties": false,
        +      "properties": {
        +        "speed_kmh": {
        +          "maximum": 150,
        +          "minimum": 1,
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "speed_kmh"
        +      ],
        +      "type": "object"
        +    }
        +  },
        +  "type": "object"
        +}
      • changedInput schema / properties / constraints / properties / allowed_tools / maxItems
        Previous value: -13New value: +14
    • Changedplan_mission1 field changed
      • changedInput schema / properties / allowed_tools / maxItems
        Previous value: -12New value: +13
    • Addedroute_optimizer
  3. 3 tool updates
    • Addedcompose_capabilities
    • Addedexecute_composition
    • Addedlist_composition_capabilities
  4. 3 tool updates
    • Addedbuild_robust_plan
    • Addedlist_capabilities
    • Addedplan_mission
  5. 12 tool updates
    • First observedanalyze_chess_game
    • First observedanalyze_kinematics
    • First observedaudit_statistics
    • First observedcheck_ephemeris
    • First observedcheck_shadow_track
    • First observedcompute_orbit
    • First observedoptimize_tasks
    • First observedscreen_wash_trading
    • First observedsimulate_pi
    • First observedsimulate_plan
    • First observedsolve_equilibrium
    • First observedverify_shadows

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Dependency-free, locally sandboxed MCP stdio server. Tools run as WASM inside a capability-based WASI sandbox (wasmtime); every tools/call is attested with cost/policy and signed manifests fail-closed.
    3
    408 PyPI
    45
    Apache 2.0
  • F
    license
    A
    quality
    B
    maintenance
    Local-first MCP tools for AI-assisted work receipts, workspace maps, routing ledgers, measured verdicts, and shared state verification across the five Project Telos flagships.
    41
    2
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables deterministic, neural-free compilation of natural language into typed intent and canonical receipts, with proof-gated synthesis and governed tool use via read-only compiler functions.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources