Skip to main content
Glama

sentinel_accuracy

Review live precision, recall, Brier score, calibration, degradation status, and training holdout labels over a rolling window before acting on a forecast.

Instructions

Sentinel's published accuracy: live precision, recall, Brier score and calibration over a rolling window, whether the model is marked degraded, and the training holdout labelled as such. Read this before acting on a forecast. Read-only public GET.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
window_daysNoRolling window in days, 1-365 (default 30)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv3.0.2

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningful work: it discloses that this is a read-only public GET (no auth/scope concerns) and that the training holdout is labelled as such, so the agent won't conflate holdout numbers with live performance. It still omits rate limits, pagination, and behaviour on out-of-range windows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: what is returned, the when-to-use trigger, and the safety profile. The most actionable clause (read before acting on a forecast) is front-loaded near the top.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter read tool with no output schema and no annotations, the description covers purpose, usage trigger, and safety profile adequately, and enumerating the returned fields compensates for the missing output schema. Minor gaps remain around error/edge behaviour for the window parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — window_days is documented with type, range 1-365 and default 30 — so the schema already does the heavy lifting. The description only alludes to a "rolling window" without adding format or boundary behaviour, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (Sentinel's published accuracy) and enumerates the exact metrics returned: live precision, recall, Brier score, calibration, degraded status, and labelled training holdout. It is clearly distinct from sibling metric tools like sentinel_calibration_history and get_risk_forecast, though it never names them to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Read this before acting on a forecast" gives an explicit precondition for use, which is real routing guidance rather than an implied one. It stops short of the top band because it names no alternative tool or when-not-to-use condition (e.g. historical vs current accuracy).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools