Skip to main content
Glama

Server Details

Smarter Weather MCP: forecasts, alerts, outlooks, observations, AQI, grids, and map imagery.

Ownership verified
Status
Healthy
Uptime
92.0% over 47 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
smarterweather/developer
GitHub Stars
1

TDQS

A4.1/5.0

Scored across 35 tools

Disambiguation5/5

Despite the large number of tools, each has a clearly distinct purpose, and the descriptions explicitly steer between similar-sounding options such as get_forecast, get_hourly_forecast, and get_current_conditions. Aggregate tools are distinguished from lower-level/raw tools, and data/image pairs are clearly separated.

Naming Consistency5/5

All tool names use consistent snake_case with predictable action prefixes such as get_, list_, query_, search_, describe_, compare_, and find_. There are no mixed conventions or confusing abbreviations.

Tool Count2/5

At 35 tools, the set is substantially larger than the typical well-scoped range and imposes high selection overhead on an agent. While each tool has a coherent niche, several specialized tools could likely be consolidated without losing meaningful capability.

Completeness4/5

Coverage is broad across forecasts, observations, climate, hazards, tropical weather, exposure, maps, datasets, and geocoding. Minor gaps remain for some specialized domains such as explicit TAF/aviation forecasts and marine/hydrological forecast products, but core weather workflows are well covered.

Available Tools

35 tools
compare_locationsCompare locationsA
Read-onlyIdempotent
Inspect

Compare forecast variables across multiple locations side-by-side in one batched call. Returns a distilled per-location series matrix for direct comparison -- prefer this over N sequential forecast calls. Locations accept place names directly. Example: {"locations": [{"location": "Denver"}, {"location": "Boulder, CO"}], "variables": ["temperature_2m", "precipitation_probability"], "hours": 48}.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNoForecast hours. Default 24.
locationsYesLocations to compare (2-10). Each takes location OR lat/lon, optional label.
variablesYesStandard variable names (e.g. temperature_2m, precipitation).
dataset_idNoDataset override. Default: auto-resolved NBM per location.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hoursYes
variablesYes
comparisonsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it is a batched call, returns a distilled per-location matrix, and accepts place names directly. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: main capability first, output behavior second, usage note third, and a concrete example last. Every sentence contributes useful information without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema, output schema, and read-only/idempotent annotations, the description covers the core behavioral contract completely: batched multi-location comparison, direct place-name input, matrix output, and preference over sequential calls. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all parameters already documented clearly. The description's example reinforces the expected shape but adds little semantic meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: compare forecast variables across multiple locations in one batched call. It also differentiates itself from sequential forecast calls by emphasizing the 'distilled per-location series matrix' return, which separates it from sibling tools like get_forecast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to prefer this tool over N sequential forecast calls, which conveys the primary comparison use case. It does not name a specific sibling alternative or state when to use a single-location tool instead, but the guidance is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_datasetDescribe datasetA
Read-onlyIdempotent
Inspect

Variables available in a dataset, with standard names, units, descriptions, and the time range of available data. Use before query_dataset to discover valid variable names. Example: {"dataset_id": "nbm_conus"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataset_idYesPipeline dataset identifier (e.g. "nbm_conus", "mrms_reflectivity_conus"). Discover valid values with list_datasets; inspect variables with describe_dataset.

Output Schema

ParametersJSON Schema
NameRequiredDescription
detailYes
dataset_idYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds useful context about what information is returned (names, units, descriptions, time range) but does not disclose additional behavioral traits such as response shape nuances or potential edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the core output defined first and a concrete example second. Every sentence earns its place, and the example is directly actionable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one simple parameter, rich annotations, an output schema, and a description that explains both the returned content and the intended usage sequence. Nothing essential is missing for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter dataset_id is already well documented with examples and cross-references to list_datasets and describe_dataset. The description adds an example value ('nbm_conus') but does not materially expand beyond what the schema already provides; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource (dataset) and a specific purpose: listing available variables with standard names, units, descriptions, and time range. It also distinguishes itself from query_dataset by saying it is meant to discover valid variable names before querying.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use before query_dataset to discover valid variable names,' giving clear when-to-use guidance and naming the key alternative. It does not enumerate when not to use it or contrast with list_datasets, but the core usage direction is present and unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_best_windowFind best weather windowA
Read-onlyIdempotent
Inspect

Find the optimal time window for an activity based on weather criteria. Scans the forecast and returns daylight-aware periods matching all conditions. Criteria are expressed in the selected units system (default imperial: °F, mph, miles, feet). Example: {"location": "Boulder, CO", "criteria": {"min_temperature": 55, "max_wind_speed": 15, "max_precipitation_probability": 20}, "hours": 72, "activity_duration_hours": 3}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
hoursNoHours to search. Default 72.
unitsNoUnit system for all values in the request and response: imperial (°F, mph, inches), metric (°C, km/h, mm), or si (K, m/s, mm). Defaults to imperial.imperial
activityNoFree-form activity to rate the windows for (e.g. "afternoon golf"). Accepted even when rating is off so clients can send it.
criteriaYesWeather criteria defining acceptable conditions (all optional).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
daylight_onlyNoOnly consider daylight hours (sunrise to sunset). Default true.
activity_duration_hoursNoMinimum consecutive hours meeting criteria. Default 2.

Output Schema

ParametersJSON Schema
NameRequiredDescription
unitsYes
messageNo
windowsYes
locationYes
sun_timesNo
daylight_onlyYes
criteria_appliedYes
activity_duration_hoursYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the safety profile is covered. The description adds behavioral context beyond annotations: it discloses that it scans the forecast, returns daylight-aware periods, requires all conditions to match, and specifies the units system. It doesn't describe edge cases like no matching window or the exact optimization objective, but given annotation coverage, the added detail is sufficient for a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a one-sentence purpose, a behavior statement, a units note, and an example. It is front-loaded with the core purpose, and every sentence contributes meaning. The example is compact yet informative, covering key parameter usage. No filler words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 parameters and a nested object, the description covers the essential concepts (criteria, units, daylight-awareness, example) without needing to explain every field since the schema is fully documented. An output schema exists, so return-value details are not required. The description is complete enough for an agent to understand the tool's capabilities, though it could briefly clarify what 'optimal' means in terms of ranking windows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value by explicitly noting that criteria are expressed in the selected units system and by providing a concrete example covering multiple criteria and parameters. This helps agents understand how to structure the `criteria` object and how units apply, going beyond the schema's per-field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Find the optimal time window for an activity based on weather criteria.' It names a specific verb ('find') and resource ('optimal time window'), and distinguishes itself from siblings by focusing on window selection rather than raw data retrieval. The additional phrase 'Scans the forecast and returns daylight-aware periods matching all conditions' further clarifies the unique functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: it is for finding time windows based on weather criteria, which implies when an agent would choose this over raw forecast tools. It doesn't explicitly mention alternatives or when-not-to-use, but the stated purpose is distinct enough that an agent can infer appropriate usage. No exclusions are given, but the clarity of the purpose earns a solid 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_air_qualityGet air qualityA
Read-onlyIdempotent
Inspect

AirNow air quality at a location (CONUS): current overall AQI plus per-pollutant detail (PM2.5, ozone, PM10 concentrations) and the AirNow AQI forecast. AQI scale: 0-50 good, 51-100 moderate, 101-150 unhealthy for sensitive groups, 151-200 unhealthy, 201-300 very unhealthy, 301+ hazardous. pollutants=["aqi"] (default) is the cheap headline call; add pollutant keys or include_forecast=true when the user digs in. Example: {"location": "Boise", "pollutants": ["aqi", "pm25"], "include_forecast": true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
pollutantsNoWhich measurements to return. aqi = combined AQI index. Default: ["aqi"].
include_forecastNoAlso return the AirNow next-day AQI forecast.

Output Schema

ParametersJSON Schema
NameRequiredDescription
widgetNosw-ui-spec widget block rendered by the MCP Apps weather widget (ui://weather-widget/v1/index.html). Additive; safe to ignore.
currentYes
forecastNo
locationYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior; the description adds genuine behavioral context: the AirNow source, the CONUS geographic limitation, the AQI band scale, and that forecast is opt-in. No statement in the description contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Each of the four sentences earns its place: resource and return scope, AQI scale, default-versus-expanded usage guidance, and a concrete example. The identifying scope is front-loaded, and there is no filler or repetition of schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-format details need not be repeated. The description covers data source, geography, default behavior, optional expansions, and an example, while the schema covers coordinate/location validation and constraints. An agent has what it needs to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes all five parameters at 100% coverage, so the baseline is 3. The description adds strategy on top of the schema by labeling the default call as 'cheap' and showing a complete example combining location, pm25, and include_forecast, which helps an agent compose parameter combinations rather than merely filling names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete resource and scope: 'AirNow air quality at a location (CONUS)' and lists exactly what is returned: current overall AQI, per-pollutant detail, and forecast. It is clearly distinct from sibling weather tools because it identifies the AirNow source and AQI domain rather than generic current conditions or observations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (for AirNow AQI at a location) and provides parameter-selection guidance: the default ['aqi'] is the 'cheap headline call', and additional pollutants or include_forecast should be added 'when the user digs in'. It does not explicitly name alternative sibling tools or state exclusion criteria, which keeps it from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_alertsGet NWS alertsA
Read-onlyIdempotent
Inspect

NWS watches, warnings, advisories. Point (city/ZIP/lat+lon): containing polygons. BBox or US state/DC/CONUS (codes, full names, US/national): intersecting polygons. A city miss is not a statewide all-clear — query the state or a bbox; never say regional inventory is impossible. NY/WA and "New York State"/"Washington State" are states; "New York"/"Washington" stay cities. Omit at for now; at (ISO-8601 UTC) is the snapshot then. Empty = all-clear or purged (~24h). alert_id = detail+geometry, ignores at. Ex: {"location":"WI","events":["Tornado Warning"]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
atNoISO-8601 UTC past instant for the in-effect snapshot. Ignored with alert_id.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
bboxNoBounding box {west,south,east,north}. Skips geocoding; intersecting polygons.
eventsNoOptional event-name filter, e.g. ["Tornado Warning"].
alert_idNoAlert identifier for detail mode. When set, location is ignored.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
alertNo
alertsNo
widgetNosw-ui-spec widget block rendered by the MCP Apps weather widget (ui://weather-widget/v1/index.html). Additive; safe to ignore.
alert_idNo
locationNo
valid_timeNoEcho of at when an as-of snapshot was requested.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/openWorld/idempotent annotations, the description reveals important behavioral traits: empty results mean all-clear or data purged around 24 hours, alert_id switches to detail mode and ignores at, and location parsing treats NY/WA as states while New York/Washington remain cities. It also warns against claiming regional inventory is impossible, reinforcing the open-world hint with concrete guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient, packing multiple query modes, edge cases, state-name disambiguation, temporal semantics, empty-result meaning, and an example into a few telegraphic sentences. Nothing is wasted, and the most important product-level information comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and all parameters already documented, the description covers the remaining ambiguities an agent would face: location naming pitfalls, bbox and point semantics, alert_id detail mode, at behavior, and the meaning of empty results. There is no obvious missing context needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already documents all parameters at 100% coverage, the description adds real semantic value: it explains the point-versus-bbox containment distinction, clarifies that at is a snapshot selector and should be omitted for current alerts, defines alert_id behavior, and gives concrete location disambiguation rules plus a full example. This goes well beyond the baseline for fully covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title "Get NWS alerts" plus the opening phrase 'NWS watches, warnings, advisories' makes the verb and resource unmistakable. It is clearly distinct from sibling forecast, observation, and climate tools because it specifically targets alert products, not general weather data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong operational guidance: point queries return containing polygons, bbox/state/CONUS return intersecting polygons, city misses must not be treated as statewide all-clear, and at should be omitted for current snapshots. It does not explicitly name sibling alternatives, but the tool's domain is unique enough that the routing is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_climate_normalsGet climate normalsA
Read-onlyIdempotent
Inspect

Day-of-year climate normals (NCEI 1991-2020 30-year averages) for a US location, from the nearest station with a record. Returns normal high, normal low, and normal mean for each date in the window, plus the station and how far away it is. Use this whenever a question needs a baseline rather than a forecast: "is this warm for October?", "what is a typical high here in January?", "how does this week compare to normal?". Pair it with get_forecast to say how far above or below normal the coming days run. Covers dates by day of year, so it answers for any date, past or future -- these are long-period averages, not a forecast and not observed history for a specific year.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoInclusive range end as YYYY-MM-DD. Must be supplied with start.
latYesLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonYesLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
daysNoRange length in days when start/end are omitted; the window opens today (default 14).
unitNoUnit system for the normals. Default imperial (°F).
startNoInclusive range start as YYYY-MM-DD. Must be supplied with end.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only, idempotent, and non-destructive, so the description's job is to add behavioral context. It does: returns data from the nearest station with a record, includes station distance, covers any date by day of year, and explicitly states these are long-period averages, not forecasts or single-year observations. This goes well beyond the annotation metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: definition, return values, usage trigger, pairing guidance, and the day-of-year caveat. The most important scoping information is front-loaded in the first sentence, and despite its length, the description remains efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, the description fully covers what an agent needs to select and call it correctly: what it returns, when to use it, how it relates to forecasts, and the key conceptual caveat about day-of-year averaging. No critical decision-relevant context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds temporal semantics not present in the schema: dates are interpreted by day of year, making any past or future date valid. It also clarifies the meaning of the returned 'normal' values in terms of 30-year averages, which helps the agent reason about start/end/days parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific definition: 'Day-of-year climate normals (NCEI 1991-2020 30-year averages) for a US location, from the nearest station with a record.' It names the resource (climate normals), the scope (US location), and the return values (normal high, low, mean, station and distance). It also distinguishes itself from siblings by noting it is not a forecast and not observed history for a specific year, which differentiates it from get_forecast and get_climate_records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it: 'Use this whenever a question needs a baseline rather than a forecast,' with concrete examples. It also gives pairing guidance: 'Pair it with get_forecast to say how far above or below normal the coming days run.' This gives the agent clear decision criteria versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_climate_recordsGet climate reports and recordsA
Read-onlyIdempotent
Inspect

NWS daily climate data: type=reports returns CLI daily climate reports (observed high/low/precip vs normals per station); type=records returns RER record event reports (record highs/lows/rainfall actually set). Filter by wfo (3-letter office, e.g. DMX), station, date (YYYY-MM-DD), start/end range, or hours lookback. Examples: {"type": "records", "hours": 48} or {"type": "reports", "wfo": "DMX", "date": "2026-07-04"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoRange end date, YYYY-MM-DD.
wfoNoWFO office filter (e.g. DMX, OUN).
dateNoSingle date, YYYY-MM-DD.
typeYesreports = CLI daily climate reports; records = RER record events.
hoursNoLookback window in hours (1-168) when no date/range is given.
startNoRange start date, YYYY-MM-DD.
stationNoStation identifier filter (reports only).
record_typeNoRecord type filter (records only), e.g. HIGH, LOW, RAIN.

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
resultsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint, idempotentHint, openWorldHint, and non-destructive behavior. The description adds meaningful behavioral context by explaining what each type actually returns (observed vs normals for reports; actually set records for records) and by showing example filter combinations. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it starts with the core resource, immediately distinguishes the two type modes, then lists filters, and finishes with concrete JSON examples. Every sentence contributes value and the examples make invocation behavior unambiguous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema with full parameter documentation, strong annotations, and an output schema, the description is complete enough for an agent to call this tool correctly. It covers the key selection logic between reports and records, available filters, and representative example calls. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has 100% description coverage for all eight parameters, including formats, constraints, and type-specific applicability. The description reinforces wfo format and date format and gives examples, but it does not add substantial meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that this tool returns NWS daily climate data, and explicitly breaks out the two modes: type=reports returns CLI daily climate reports (observed values vs normals) and type=records returns RER record event reports. This makes the resource and the distinction between report types clear, and it differentiates the tool from climate-normals-focused siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context by listing allowed filters (wfo, station, date, start/end range, hours lookback) and provides two concrete examples showing valid payloads. It does not explicitly name alternatives or state when not to use this tool, but the scope is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_current_conditionsGet current conditionsA
Read-onlyIdempotent
Inspect

Current weather right now at a location from two independent sources in one call: the RTMA gridded analysis (exact-point values, updated sub-hourly) and the nearest METAR station observation (ground truth with raw METAR, flight category). Use the analysis for point-accurate values and the station for verification. Analysis fields: temperature_2m, dew_point_2m, relative_humidity_2m (derived here from temperature and dew point; listed in analysis.derived), wind_speed_10m, wind_direction_10m, wind_gusts_10m, surface_pressure (station pressure at ground level, not sea-level pressure), visibility, cloud_cover, cloud_ceiling. Analysis values are SI by default; pass units "imperial" or "metric" to convert them (labels in analysis.unit_labels; get_forecast defaults to imperial). nearest_station is the station's own report, unconverted. For a forecast, use get_forecast. Example: {"location": "Pella, IA"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
unitsNoUnit system for the analysis values: si (default; K, m/s, Pa, m, as RTMA serves them), imperial (°F, mph, inHg, mi, ft) or metric (°C, km/h, hPa, km, m). Unlike get_forecast, which defaults to imperial. Does not apply to nearest_station.si
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
unitsYes
analysisYes
locationYes
nearest_stationYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds valuable context beyond annotations: sub-hourly updates, that RH is derived, that surface_pressure is station-level not sea-level, and unit conversion behavior. It doesn't mention rate limits or auth, but for a read-only weather tool that's minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded and starts strong, but it becomes a dense wall of field names and unit details that could be streamlined since the schema already covers parameters. It's informative but verges on over-specification, and the example is helpful but placed at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (dual data sources, many fields, unit handling) and the presence of an output schema, the description provides enough context to invoke it correctly. It covers sources, fields, units, and alternatives. It doesn't explain the output schema's structure, but that's acceptable since the output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents lat, lon, units, and location thoroughly. The description adds a small clarification about default units (si) and the distinction from get_forecast's imperial default, but mostly repeats schema content. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Current weather right now at a location') and then distinguishes itself from siblings by naming the two data sources (RTMA analysis, nearest METAR station) and explicitly routing forecasts to get_forecast. This level of specificity lets an agent select it correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear guidance on when to use analysis vs station ('Use the analysis for point-accurate values and the station for verification') and explicitly names get_forecast as the alternative for forecasts. It doesn't mention siblings like get_observations or get_hourly_forecast, which could be ambiguous, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forecastGet forecastA
Read-onlyIdempotent
Inspect

Complete weather overview for a location: current conditions, daily forecast (day/night periods, SPC threats, severity, CAPE, UV), active alerts, and convective outlooks in one call. Data is pre-aggregated across NBM, HRRR, GFS, RTMA, and SPC and unit-converted server-side. This is the primary weather tool; reach for lower-level tools only when you need raw observations or a specific dataset. Accepts a place name directly. Examples: {"location": "Denver"} or {"location": "Portland, OR", "days": 5} or {"lat": 41.4, "lon": -92.9}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
daysNoNumber of forecast days (1-14). Default 10.
unitsNoUnit system for all values in the request and response: imperial (°F, mph, inches), metric (°C, km/h, mm), or si (K, m/s, mm). Defaults to imperial.imperial
includeNoComma-separated sections: current, daily, hourly, alerts, outlooks. Default "current,daily,alerts,outlooks". Use get_hourly_forecast for hourly detail.current,daily,alerts,outlooks
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
detail_levelNostandard: compact response (~5-10KB); daily includes day_precip_probability / night_precip_probability when available (precip_probability is max of day/night). detailed: also preserves CAPE, UV, full day/night period objects, extra hourly fields (~12-20KB).standard

Output Schema

ParametersJSON Schema
NameRequiredDescription
unitsYes
widgetNosw-ui-spec widget block rendered by the MCP Apps weather widget (ui://weather-widget/v1/index.html). Additive; safe to ignore.
forecastYes
locationYes
data_statusNoPresent only when the platform reports degraded/outage data sources: overall state, a caveat note, and the affected sources. Absent means no advisory was available -- not a freshness guarantee. See get_platform_status for the full document.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, non-destructive, idempotent behavior. The description adds useful behavioral context by explaining data is pre-aggregated across NBM, HRRR, GFS, RTMA, and SPC and unit-converted server-side. Minor flaw: listing CAPE and UV as part of the daily forecast slightly overstates what the default standard detail_level may include.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description front-loads the core purpose, then adds behavioral context, routing guidance, and examples in a compact format. Every sentence contributes meaning; the JSON examples are especially efficient for showing valid parameter combinations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich output schema, 100% parameter coverage, and helpful annotations, the description fully covers what the tool does, when to use it, and how to call it. It also provides enough sibling differentiation to guide tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters. The description adds value with concrete usage examples, clarification that location accepts a place name directly, and the note about server-side unit conversion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: it returns a complete weather overview for a location, explicitly listing current conditions, daily forecast, alerts, and outlooks. It also distinguishes itself as the primary weather tool, separating it from lower-level siblings like get_current_conditions and get_alerts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly positions itself as the default weather tool and tells agents to reach for lower-level tools only when raw observations or a specific dataset is needed. It also cross-references get_hourly_forecast for hourly detail, giving concrete routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forecast_discussionGet forecast discussionA
Read-onlyIdempotent
Inspect

Expert forecaster text products. type=afd: Area Forecast Discussion. type=hwo: Hazardous Weather Outlook. type=now: WFO short-term NOW. type=fwf/hls/esf: local fire weather / hurricane local statement / hydrologic discussion. type=mcd: SPC Mesoscale Discussion. type=mpd: WPC Mesoscale Precipitation Discussion (flash flood). type=swo/fwd/ero: national outlook discussions. type=tcd/tcp/tcm/twd/two: NHC tropical text (type=two is the text TWO, not GIS nhc_two). type=pmd: WPC/CPC desk discussion (pass awips_id for a specific desk, e.g. PMDSPD). type=pwo: SPC public weather outlook. National types (swo/fwd/ero/tcd/tcp/tcm/twd/two/pmd/pwo) need no location; day selects the outlook day for swo and fwd. summary_only=true returns the pipeline LLM summary without the full body. Examples: {"location": "Des Moines", "type": "afd"} or {"type": "swo", "day": 2, "summary_only": true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dayNoOutlook day for type=swo or type=fwd (1-8). Ignored for the other types; WPC files ERO days 1-3 under one product.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
wfoNoWFO identifier override (e.g. BOU). Default: resolved from the location.
typeYesProduct type: afd (WFO discussion), hwo (hazard outlook), now (short-term NOW), fwf/hls/esf (local WFO), mcd (SPC mesoscale), mpd (WPC precipitation discussion), swo (SPC convective outlook), fwd (SPC fire weather), ero (WPC excessive rainfall), tcd/tcp/tcm/twd/two (NHC tropical text; two is text TWO not GIS), pmd (desk discussion), pwo (SPC public outlook).
limitNoNumber of recent products (1-10). Default 1 (latest).
awips_idNoFull AWIPS identifier (e.g. TCDAT1, PMDSPD). More specific than type + location. Exact source_ref match.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
summary_onlyNoReturn only the LLM summary + sections, omitting the full body text.

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
locationNo
productsYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses several non-obvious behaviors: type=two is the text TWO not the GIS product, pmd requires an awips_id for a specific desk, WPC files ERO days 1-3 under one product, summary_only returns the pipeline LLM summary, and national types accept no location. These details go well beyond the annotations (readOnly, idempotent, non-destructive) and materially improve correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but necessarily so, since it must document 18 product type codes. It is front-loaded with the core purpose and uses a compact semicolon-separated style with concrete examples at the end. A small deduction for redundancy: much of the type-by-type explanation is duplicated in the input schema's enum descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters, 1 required, a rich enum, and an output schema, the description covers all critical invocation aspects: type selection, location options, day semantics, summary_only, and representative examples. The presence of an output schema means the response format need not be described. No notable gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by giving concrete invocation examples, clarifying that national types require no location, and explaining the pmd desk-desk behavior with a sample awips_id. It partially repeats schema descriptions for the type enum, which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Expert forecaster text products' and enumerates every product type with a one-line explanation, making it clear this tool retrieves NWS/SPC/WPC/NHC text forecast discussions, not gridded forecast data. The name and title align with the described content, and the level of detail distinguishes it from sibling tools like get_forecast and get_current_conditions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool and how to select among its many type values, including the rule that national types need no location and that `day` only applies to swo/fwd. It does not explicitly name sibling alternatives or state when not to use it, but the context is concrete enough that an agent can infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forecast_distributionGet forecast distributionA
Read-onlyIdempotent
Inspect

Probabilistic forecast guidance from NBM for one aspect of the weather: percentile ranges (p10-p90), exceedance probabilities, and ensemble spread. Use this for any question about odds, ranges, potential or confidence ("how much could we get", "worst case for the wind", "how sure is this") -- a deterministic forecast value cannot answer one. Reading the percentiles: p50 is the most likely outcome, p90 is the reasonable worst case when the risk is the high end (snow totals, wind, rainfall), and p10 is the reasonable worst case when the risk is the low end (cold, minimum visibility, ceiling). A single percentile is not the forecast -- report the likely value with the tail that matters, and label which is which. Aspects: precip (PoP, QPF + percentiles), snow (accumulation percentiles, >1/2/4in probabilities, snow level, snow-type probability), ice (freezing-rain ice accretion, freezing-rain / ice-pellet type probabilities), temperature (temp/dewpoint + stddev), wind (speed/gust percentiles), severe (thunderstorm probability plus NBM CWASP, the Craven-Wiedenfeld Aggregate Severe Parameter, a severe-environment index, both in percent: severe_weather_prob_p50 is the median CWASP value and severe_weather_prob_above_75 the probability CWASP exceeds 75; NBM has no tornado, hail or damaging-wind probabilities, so use get_outlooks hazard=severe for SPC's), aviation (LIFR/IFR/MVFR visibility + ceiling probabilities), confidence (ensemble stddev; low spread = settled forecast, high spread = details still in play). Examples: {"location": "Denver", "aspect": "snow", "hours": 72} or {"lat": 32.9, "lon": -97.0, "aspect": "severe"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
hoursNoForecast hours (1-264). Default varies by aspect (48-72).
aspectYesWhich distribution family to return (see tool description).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hoursYes
aspectYes
seriesYes
locationYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/openWorld/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: the NBM source, that p50/p90/p10 should be reported together rather than singly, and a genuine limitation (NBM has no tornado, hail or damaging-wind probabilities). It stops short of noting any latency, rate-limit, or update-cadence behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then interpretation rules, then the aspect catalog, then examples – a sensible order with no filler sentences. The aspect enumeration is dense and pushes the block long, but each clause carries information an agent needs to select the right aspect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-format explanation is unnecessary, and the description still covers everything else: when to reach for it, how to read the percentiles, what each aspect yields, and where the data falls short. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description goes well past the field names: every value of the required `aspect` enum is explained in terms of what it actually returns (PoP/QPF, snow-level and >1/2/4in probabilities, CWASP semantics, LIFR/IFR/MVFR, stddev). It also supplies worked call examples that clarify the location-vs-lat/lon choice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (probabilistic NBM forecast distribution) and immediately scopes it to percentile ranges, exceedance probabilities and ensemble spread. It explicitly distinguishes itself from the deterministic sibling by saying a deterministic forecast value cannot answer these questions, and it carves out the severe-hazard case to get_outlooks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggering conditions ('any question about odds, ranges, potential or confidence') with quoted user phrasings, and names the alternative for one sub-case ('use get_outlooks hazard=severe for SPC's'). When-to-use, when-not, and alternatives are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forecast_skillGet forecast skillA
Read-onlyIdempotent
Inspect

How accurate our forecasts have actually been near a location, measured against observed analysis truth. Returns bias (positive = the model runs high), mean absolute error, RMSE, and a skill score against local climatology, per model, weather variable, and forecast lead time; continuous and vector entries also carry persistenceSkillScore, skill against the analysis at forecast issue time (null means not enough persist pairs, not zero skill -- do not compare it to skillScore as if they shared a denominator), and analysisDisagreementMae, the analyses' own disagreement at that lead -- a floor on how good the forecast can look, not a skill score and not an excuse (null means the sibling row is missing or below minimumSamples); for probability forecasts, the Brier score and a reliability breakdown. Use this to qualify a forecast rather than assert it -- "NBM has been running 1.8F warm at 3-day leads near you, so treat that 72 as around 70" -- and to answer "how much should I trust this forecast", "is the model biased here", or "how accurate were you last month". Evidence is reported at three scopes side by side: the exact point (strongest, slowest to accumulate), the ~50km neighborhood, and the ~300km region. Prefer the most specific scope that has samples. Metrics below minimumSamples observations are withheld and listed under insufficientHistory with their count -- say that history is still accumulating rather than treating thin numbers as evidence. Coverage is a rolling recent window over verified US variables, not all of history. Entries are per model and their samples are not matched, so never conclude that one model beats another by comparing their numbers here. Each entry states the truth field it was measured against -- one designated analysis per variable -- so never compare numbers carrying different truth values either. Each entry also states the regime it was measured under: ALL for every observation regardless of weather, or a conditioned tier such as SEA:DJF (winter), SCN1:WINDY / SCN1:WET / SCN1:QUIET (what the forecast was showing), or JC1:NW (a circulation pattern). Pass the regime parameter to ask for a conditioned track record. It falls back, so asking for SCN1:WINDY and getting back regime ALL is a successful answer, not a missing one -- always read the regime field and qualify the claim with it, because "NBM runs warm here when it shows windy" and "NBM runs warm here" are different statements. Regimes overlap by construction across families, so entries under different regimes are alternative answers to one question and must never be compared or added; within SCN1: the labels are mutually exclusive. Entries with a categorical block answer a yes/no question instead of an error magnitude -- did it rain, at the thresholdMm stated on the entry -- with pod (of the times it happened, how often we called it), far (of the times we called it, how often it did not happen), and frequencyBias (above 1 = we call it too often). Use these for "will it actually rain" questions, where a small average error means nothing if the rain lands in the wrong hour. A null rate means the sample cannot answer it -- the event has not happened, or been forecast, enough times to divide by -- and must be reported as unknown, never as zero. The counts beside it are still evidence, and for a rare event they are often the whole answer: "it has only rained twice here in the record" is a useful thing to say.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
unitNoUnits for the error magnitudes. Default imperial (bias/MAE/RMSE in °F, mph, in).
modelNoNarrow to one model, e.g. nbm or rrfs.
truthNoMeasure against a named truth source instead of the default one for each variable, e.g. urma. Only pass this if the user asked which analysis was used or named one; the default is already the designated source, and the analyses disagree, so switching changes the numbers.
regimeNoAsk for a track record measured only under particular conditions, as a comma-separated preference chain, most specific first, e.g. "SCN1:WINDY,SEA:JJA". SEA: is the meteorological season (DJF, MAM, JJA, SON); SCN1: is a forecast-conditioned scenario (WINDY, WET, QUIET — mutually exclusive within the family); JC1: is a circulation pattern. The most specific tier with enough observations answers and the unconditioned record is the last resort, so this never empties a result the way truth does -- it degrades. Read the regime field on each entry to see which tier actually answered. Pass this when the question is conditional ("is it worse in winter", "how does it do when the model shows windy"); omit it otherwise, since conditioned tiers are thinner and slower to earn numbers.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
variableNoNarrow to one variable, e.g. temperature_2m, dew_point_2m, wind_speed_10m, wind_gusts_10m, wind_vector_10m, cloud_cover, precipitation, precipitation_probability, or a thresholded rain event such as precipitation_gt_0p254mm (any measurable rain) or precipitation_gt_2p54mm. Omit for everything measured at the location.
lead_hoursNoNarrow to the lead time being asked about, in hours; the containing lead bucket is selected for you (60 gives the 48-72h bucket). Use the lead of the forecast you are qualifying: ~24 for tomorrow, ~72 for three days out. Never approximated -- a lead we have not verified returns no entries rather than a nearby bucket, so an empty result means we cannot speak to that range.

Output Schema

ParametersJSON Schema
NameRequiredDescription
cellsNo
skillYes
unitsYes
trackedYesWhether this exact coordinate is one the verification pipeline tracks.
locationYes
minimumSamplesYes
insufficientHistoryYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint/openWorldHint/idempotentHint and the description agrees with no contradiction. Beyond annotations it discloses exceptional behavioral depth: null means 'not enough persist pairs, not zero skill — do not compare it to skillScore'; regime queries degrade to ALL rather than erroring; metrics below minimumSamples are withheld into insufficientHistory; coverage is a rolling recent window over verified US variables, not all history; samples are not matched across models so cross-model comparison is invalid; entries under different regimes are alternative answers that must never be compared. This far exceeds what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose, return values, and the usage example are front-loaded before the caveat sections, and nearly every sentence carries real interpretive weight given the tool's complexity. It loses a point for sheer length — roughly 700 words, at the extreme end of what an agent can scan efficiently — where tighter bulleted structure would improve parsability. Every section earns its place, but the whole is longer than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no required parameters and an output schema present, the description covers everything needed for correct invocation and result interpretation: three scopes with a preference rule ('prefer the most specific scope that has samples'), minimumSamples withholding, null semantics for multiple field types, regime fallback behavior, and three separate incomparability constraints (model, truth, regime). The output schema carries return structure, so nothing critical for correct use is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds genuine interpretive semantics beyond the field-level text: model and truth carry comparability constraints ('never conclude that one model beats another', 'never compare numbers carrying different truth values'), regime gains cross-family overlap and within-family mutual exclusivity ('within SCN1: the labels are mutually exclusive'), and lead_hours gains the 'never approximated — a lead we have not verified returns no entries' warning. This is meaningful added value, though not organized parameter-by-parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource — measuring forecast accuracy near a location against observed analysis truth — and enumerates the concrete outputs (bias, MAE, RMSE, skill score). The phrase 'qualify a forecast rather than assert it' plus the worked example ('NBM has been running 1.8F warm at 3-day leads') sharply separates this from forecast-content siblings like get_forecast or get_hourly_forecast. An agent can identify what this tool is for without inspecting sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use signals: qualify a forecast rather than assert it, and answer 'how much should I trust this forecast', 'is the model biased here', or 'how accurate were you last month'. It also provides an example of how the answer should be phrased. However, it never names alternatives or states when-not-to-use conditions (e.g., nothing says 'for the actual forecast values, use get_forecast'), so routing relies on clear context rather than explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forecast_skill_mapGet forecast skill mapA
Read-onlyIdempotent
Inspect

Measured forecast skill as a map of res-4 hexes inside a bounding box, sliced from the public skill map. Use this when the question is spatial ("where is NBM temperature skill weak across Colorado"), not "how good is the forecast at this point" — that is get_forecast_skill. Requires model, variable, lead_hours, and bbox {west,south,east,north}. lead_hours selects the containing published bucket (20 → NBM 24); a lead we have not published returns no slices rather than a nearby one. A hex is included when its center is inside the box; edge-overlapping hexes are dropped. The box cannot wrap the dateline. Each cell is [h3, samples, skill_score]; skill_score is unitless (fraction of climatological variance explained) and already gated at n ≥ 30. withheld lists hexes still accumulating. Do not compare slices across models or truths. This is CELL# only — never treat a hex as a person.

ParametersJSON Schema
NameRequiredDescriptionDefault
bboxYesBounding box {west,south,east,north}. Hexes whose H3 center falls inside are returned.
modelYesModel id, e.g. nbm, gfs, hrrr, or rrfs.
domainNoOptional domain (conus, alaska, hawaii). Omit unless you already know it.
variableYesVerified variable, e.g. temperature_2m, dew_point_2m, precipitation, or precipitation_gt_0p254mm.
lead_hoursYesLead time in hours; the containing published bucket is selected. Never approximated.

Output Schema

ParametersJSON Schema
NameRequiredDescription
bboxYes
slicesYes
generated_atYes
atom_resolutionYes
minimum_samplesYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses lead-time bucket selection, hex-center inclusion rules, dateline restrictions, cell shape [h3, samples, skill_score], the n ≥ 30 gating, withheld hexes, and the warning not to compare slices across models or truths. This is substantial behavioral context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core purpose and spatial-vs-point routing. Nearly every sentence carries behavioral or usage value; only the line "Requires model, variable, lead_hours, and bbox" is somewhat redundant with the required fields in the schema, and "CELL#" is slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the nested bbox parameter, output schema, and nuanced slicing behavior, the description covers all the important call-time concerns: spatial interpretation, lead bucketing, inclusion rules, dateline restriction, return cell format, and gating. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: lead_hours "selects the containing published bucket (20 → NBM 24)", no-slice behavior for unpublished leads, and bbox interpretation via hex center. These details go beyond the schema's property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: "Measured forecast skill as a map of res-4 hexes inside a bounding box," which clearly distinguishes it from the point-based sibling. It also explicitly contrasts itself with get_forecast_skill, so an agent can disambiguate without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states exactly when to use this tool: "Use this when the question is spatial ... not 'how good is the forecast at this point' — that is get_forecast_skill." It also gives exclusionary edge cases, such as "returns no slices rather than a nearby one" and "The box cannot wrap the dateline," which tell an agent when the tool will not behave as expected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_growing_degree_daysGet growing degree daysA
Read-onlyIdempotent
Inspect

Growing Degree Units (GDU / GDD) for a US location (CONUS, Alaska, Hawaii), computed from daily max/min temperatures. Pass a crop id (e.g. "corn", "soybean", "wheat") to use calibrated base/upper thresholds, or crop="custom" with base_temp_c (and optional upper_temp_c / method). Without season_start you get per-day GDU across the forecast horizon; WITH season_start (YYYY-MM-DD) you get the cumulative season-to-date total (observed history + today + forecast) plus a per-day cumulative series -- the number a grower tracks against crop milestones. Answers "how many growing degree days has my corn accumulated since May 1?" and "what's the GDU forecast this week?".

ParametersJSON Schema
NameRequiredDescriptionDefault
latYesLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonYesLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
cropYesCrop id from the catalog (e.g. "corn", "soybean", "wheat") or "custom" to supply your own thresholds via base_temp_c.
daysNoForecast horizon in days (default 10).
unitNoUnit system for GDU + temps. Default imperial (°F-days).
methodNoGDU method for custom crops. Defaults from whether upper_temp_c is set.
base_temp_cNoCustom base threshold in °C. Required when crop="custom".
season_startNoSeason/planting start as YYYY-MM-DD (local date). Presence switches the response to a cumulative season-to-date GDU total. Must be within the ~180-day observed window.
upper_temp_cNoCustom upper cutoff in °C (enables the modified method). Optional.
day_definitionNoDaily boundary: "nws" (default; NBM MaxT/MinT period extremes) or "local_calendar" (midnight-to-midnight local day).
include_milestonesNoInclude the crop's growth-stage GDU milestones in the response.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already mark this as read-only, idempotent, and non-destructive, the description adds substantial behavioral context: the response differs based on season_start, the cumulative mode includes observed history plus today plus forecast, and the tool supports calibrated crop thresholds. This goes well beyond what annotations convey and is particularly valuable with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph with no wasted words. It front-loads the tool's core function and geographic scope, then explains the two modes and ends with concrete example questions. Every sentence contributes useful decision-making information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing both response shapes: per-day GDU across the forecast horizon versus cumulative season-to-date totals with a per-day cumulative series. It also covers default crop behavior, custom thresholds, optional milestones, and the USA-only geographic limitation. For an 11-parameter tool, this is complete enough to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the crop='custom' plus base_temp_c flow, the behavior switch triggered by season_start, and the effect of upper_temp_c on method selection. Not every parameter is elaborated in the description, but the most semantically complex ones are clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Growing Degree Units (GDU / GDD) for a US location (CONUS, Alaska, Hawaii), computed from daily max/min temperatures.' It also gives two concrete representative questions, 'how many growing degree days has my corn accumulated since May 1?' and 'what's the GDU forecast this week?', making the tool's purpose unmistakable and distinct from siblings like get_forecast or get_climate_normals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains when to use each mode: without season_start you get per-day GDU, with season_start you get cumulative season-to-date totals, and it explains when to use a predefined crop id versus crop='custom'. It does not explicitly name sibling tools or when-not-to-use cases, but the mode-based guidance is strong enough that an agent can decide how to invoke it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hourly_forecastGet hourly forecastA
Read-onlyIdempotent
Inspect

Blended hourly forecast: temperature, feels-like, humidity, wind, precipitation probability/amount, conditions, and icon per hour. Snapped to the current hour so hourly[0] is "now". Timestamps are UTC ISO 8601; convert to the local timezone before presenting. Ask for the days you need up front -- one call with days: 4 beats four calls. ALWAYS check hourly_coverage before answering about a specific hour: it reports first_time and last_time (the window the rows actually span), sample_interval_hours (past the first day rows are every 2-3h, not every hour), and truncated: true when upstream returned less than you asked for. If the hour the user cares about is after last_time, say the forecast does not reach that far yet rather than answering from the nearest row you do have. For one stretch of time ask for that stretch with hours_from/hours_to: it comes back hour by hour even where the full range would be sampled. Accepts a place name or coordinates. Examples: {"location": "Portland, OR", "days": 2} or {"lat": 41.88, "lon": -87.63, "hours_from": 36, "hours_to": 48}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
daysNoDays of hourly data (1-7). Default 2. Widened when hours_to reaches further.
unitsNoUnit system for all values in the request and response: imperial (°F, mph, inches), metric (°C, km/h, mm), or si (K, m/s, mm). Defaults to imperial.imperial
hours_toNoWindow end, in hours from now, exclusive. 36 to 48 is hours 36-47.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
hours_fromNoWindow start, in hours from now (0 = the current hour).
detail_levelNostandard: compact hourly data (sampled past 24h). detailed: preserves CAPE, ceiling, UV, gust, thunderstorm probability for the first 48h.standard

Output Schema

ParametersJSON Schema
NameRequiredDescription
unitsYes
widgetNosw-ui-spec widget block rendered by the MCP Apps weather widget (ui://weather-widget/v1/index.html). Additive; safe to ignore.
forecastYes
locationYes
data_statusNoPresent only when the platform reports degraded/outage data sources: overall state, a caveat note, and the affected sources. Absent means no advisory was available -- not a freshness guarantee. See get_platform_status for the full document.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses important behaviors: rows are snapped to the current hour, timestamps are UTC ISO 8601, hourly_coverage reports the actual window, sampling drops to every 2-3 hours after the first day, and truncated signals upstream shortfalls. This is exactly the kind of context an agent needs to avoid answering incorrectly from a sparse or partial result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: the first sentence defines the payload, the second defines temporal snapping and timezone handling, and the remaining sentences cover the coverage/truncation pitfalls and parameter combinations. It is front-loaded with the most important facts and contains no filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations cover safety/idempotency, the description covers the remaining gaps: timezone conversion expectations, sampling intervals, truncated responses, how to interpret hourly_coverage, and how to request either a full multi-day range or a specific hour window. An agent calling this tool has the information needed to interpret results correctly and choose sensible parameter combinations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all parameters at 100% coverage, so the baseline is 3. The description adds meaningful beyond-schema semantics: the relationship between days and hours_from/hours_to, the 'one call with days: 4 beats four calls' guidance, the sampling behavior difference when using a window, and concrete JSON examples showing locations and hour ranges. This elevates it above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource precisely: 'Blended hourly forecast: temperature, feels-like, humidity, wind, precipitation probability/amount, conditions, and icon per hour.' This is a specific, informative statement of what the tool returns. It does not explicitly contrast with sibling tools like get_forecast or get_current_conditions, but the emphasis on hourly rows and 'Snapped to the current hour' makes the tool's role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage advice: ask for days up front to avoid multiple calls, check hourly_coverage before answering about a specific hour, use hours_from/hours_to for a contiguous stretch, and treat truncation explicitly. It lacks explicit when-not-to-use guidance against sibling tools, but the usage context is strong and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_lightning_activityGet lightning activityA
Read-onlyIdempotent
Inspect

Real-time lightning near a location: GLM satellite flash count (30km/10min) and MRMS ground-truth lightning density + 30-minute probability. The summary field is ready-to-use. A zero flash count means no lightning inside that window -- report it as a quiet observation scoped to the window in scope, never as a data gap. Only call when storms may be active or the user asks about lightning. Example: {"location": "Tampa"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scopeYesArea and time window searched, so a zero count is unambiguous to report.
locationYes
lightningYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already providing readOnly, idempotent, and non-destructive hints, the description adds valuable interpretation context: a zero flash count means a quiet observation within the window, not a data gap, and the `summary` field is ready to use. It also explains the temporal/spatial window of the data, which is not in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and front-loaded, leading with the core function and data types before adding interpretive guidance and usage conditions. Every sentence earns its place, and the example is compact and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given rich annotations, a fully described input schema, and an output schema, the description covers the remaining context an agent needs: what the data represents, how to interpret zero values, the appropriate invocation window, and an example. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has 100% coverage with clear descriptions for lat, lon, and location. The description adds a practical example ({"location": "Tampa"}) and reinforces that the tool is location-oriented, helping agents choose between coordinate and place-name inputs, though it does not introduce deep semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly identifies the tool as retrieving real-time lightning activity near a location, with specific data sources (GLM satellite flash count, MRMS lightning density) and a 30-minute probability. It distinguishes itself from sibling weather-data tools by naming the exact resource and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States an explicit trigger condition: 'Only call when storms may be active or the user asks about lightning.' This tells an agent when the tool is appropriate, but it does not name alternative sibling tools for non-lightning weather queries, so it stops short of full alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_map_snapshotGet map snapshotA
Read-onlyIdempotent
Inspect

Render a weather map image for visual analysis. Simple form: pass product (a viz-catalog product_id like "mrms_qpe_01h_pass2_conus", "goes_truecolor_conus", "spc_day1_categorical", "hrrr_precip_hybrid_derived_conus" (future radar), "hrrr_subhourly_conus" (15-min Future Radar), "mrms_radar_nowcast_conus", "rtma_conus", "nbm_daily_temps", or "nexrad_l3:{SITE}:{PRODUCT}" for single-site radar, e.g. "nexrad_l3:TLX:N0B") plus a location and zoom (5=regional, 8=metro, 10=city). Composed form: pass scene -- a declarative scene document layering basemap + multiple weather products + active alerts + storm features + inline GeoJSON in one image (layers draw bottom-to-top, under basemap labels). Example scene: {"scene":"1.0","view":{"center":{"lat":43.8,"lon":-91.2},"zoom":8},"layers":[{"type":"weather","product":"goes_truecolor_conus"},{"type":"weather","product":"nexrad_l3:ARX:N0B"},{"type":"alerts","filter":{"events":["Tornado Warning"]},"onError":"skip"}]}. Alert filters (all optional, AND-combined): ids (specific alerts), events, severities, minSeverity (Extreme>Severe>Moderate>Minor>Unknown). Single-site radar keys: the address is nexrad_l3:{SITE}:{KEY} where KEY is N{tilt}{measurement} and tilt 0 is the 0.5 degree sweep -- N0B reflectivity (dBZ, where and how heavy), N0G base velocity (knots toward/away from the radar), N0S storm-relative velocity (storm motion removed, so a couplet is rotation rather than translation -- prefer it for rotation questions), N0C correlation coefficient (0-1, debris and hail), N0X differential reflectivity (dB). Legacy codes (N0V, N0R, N0Q) are accepted as aliases. Not every site produces every key; when a render reports which keys a site has, retry with one of those. Optional time (unix seconds): closest frame. Forecast (HRRR/nowcast/NBM) honors future times; analysis (MRMS/NEXRAD/RTMA/GOES) clamps to latest past. Pass time for future-radar asks — do not claim that capability is missing. Product ids must be real viz-catalog entries -- shorthand like "radar" or "reflectivity" is not one. Omit product for the default hybrid precip still. For Alaska and Hawaii prefer a local site or mrms_precip_hybrid_derived_alaska over CONUS mosaics, which do not cover them. Returns the rendered image plus per-layer resolved valid times.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
timeNoUnix seconds; closest frame (default: latest). Forecasts honor future times.
zoomNoMap zoom (simple form)
sceneNoFull scene document (composed form). When set, product/location/zoom are ignored.
widthNo
heightNo
opacityNoWeather layer opacity
productNoviz-catalog product_id or nexrad_l3:{SITE}:{KEY} (simple form)
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description adds substantial behavior beyond that: forecast products honor future times while analysis products clamp to the latest past, the return includes per-layer resolved valid times, layers draw bottom-to-top under basemap labels, and NEXRAD site-key availability may require retrying. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every sentence earns its place given 10 parameters and no output schema. It is front-loaded with the purpose and the simplest form, then moves to composed scenes, then edge cases and region-specific guidance. The structure makes the dense detail navigable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 10-parameter tool with nested scene objects and no output schema, the description is exceptionally complete: it documents the return value ('rendered image plus per-layer resolved valid times'), parameter interactions, product-id validity requirements, time semantics, and regional caveats. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 80%, the description greatly expands meaning: it gives concrete product_id examples, explains the entire NEXRAD key format (tilt, measurement, and legacy aliases), defines alert filter fields and AND-combination semantics, provides a full scene example, and clarifies interactions like `scene` overriding product/location/zoom. The schema descriptions alone would not enable correct invocation nearly as well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Render a weather map image for visual analysis.' It clearly distinguishes the simple form (single `product` plus location and zoom) from the composed form (`scene` document), which separates it from all sibling data/forecast tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to guidance: use the composed scene form for layering multiple products/alerts, pass `time` for future-radar asks, prefer local sites or Alaska-specific products over CONUS mosaics, and omit `product` for the default hybrid precip. It even warns against claiming future-radar capability is missing. While it doesn't name sibling tools, the map-rendering purpose is unambiguous among the listed siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_observationsGet station observationsA
Read-onlyIdempotent
Inspect

METAR surface observations from weather stations: temperature, wind, visibility, ceiling, flight category, raw METAR. Nearest mode (default) returns the closest N stations to a location; station mode returns history for a specific ICAO identifier. Examples: {"location": "Denver", "n": 3} or {"station": "KJFK", "hours": 6}.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoNumber of nearest stations (1-10). Default 1. Ignored in station mode.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
hoursNoHours of history in station mode (1-24).
stationNoICAO station identifier (e.g. KJFK). Switches to station-history mode.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
stationNo
locationNo
observationsYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/destructive hints, so the bar is lower. The description adds behavioral detail beyond annotations by explaining the default nearest mode, the switch to station-history mode when an ICAO is provided, and what fields the observations contain. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, information-dense, and front-loaded with the core resource before moving to modes and examples. Every sentence earns its place and the two short examples make the schema concrete without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a 100%-covered schema, an output schema, and safety-relevant annotations, the description fills the remaining gap by explaining mode semantics and typical use cases. It is complete enough for an agent to invoke the tool correctly in either mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining how the parameters interact (nearest mode vs station mode) and by giving working examples for each mode, which helps an agent choose between location/station and n/hours. This goes beyond the individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise definition: METAR surface observations from weather stations, listing the data fields (temperature, wind, visibility, ceiling, flight category, raw METAR). It clearly distinguishes the tool's two modes and differentiates it from the many forecast/alerts siblings by anchoring on surface observations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear mode-selection guidance: nearest mode for closest N stations and station mode with hours for ICAO history, along with concrete JSON examples. It does not explicitly contrast the tool with sibling tools such as get_current_conditions, but the mode guidance is enough for correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_observed_precipitationObserved precipitationA
Read-onlyIdempotent
Inspect

How much rain fell: observed precipitation totals at a location. Use for "did it rain", "how much rain fell", "rain in the last 24 hours / since yesterday". Source: MRMS MultiSensor QPE Pass2 (radar bias-corrected to rain gauges) over 1/3/6/24/48/72 h windows (CONUS; 1 h only in Alaska and Hawaii). Values have about 2 mm (0.08 in) resolution, so 0 means under about 1 mm; data runs about an hour behind real time. Pass at (within 33 h) for totals ending at a past time (newest frame up to 3 h before it). With include_nowcast (default) it also returns a next-hour radar nowcast (CONUS): peak intensity (light/moderate/heavy), type and start time, not an amount. For forecast totals beyond the next hour use get_period_totals; for station reports use get_observations.

ParametersJSON Schema
NameRequiredDescriptionDefault
atNoRFC3339 end time within the last 33 h. Default: latest.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
unitsNoUnit system for all values in the request and response: imperial (°F, mph, inches), metric (°C, km/h, mm), or si (K, m/s, mm). Defaults to imperial.imperial
windowsNoAccumulation windows ending at the latest frame (or at `at`).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
include_nowcastNoNext-hour radar nowcast (CONUS; ignored with `at`).

Output Schema

ParametersJSON Schema
NameRequiredDescription
as_ofNo
notesNo
unitsYes
derivedYes
nowcastNo
windowsYes
_sourcesYes
locationYes
data_statusNo
quantization_stepYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, non-destructive, openWorld), and the description adds substantial operational context beyond them: data source (MRMS QPE Pass2), coverage limits (CONUS; 1 h only in AK/HI), ~2 mm resolution with the meaning of zero, ~1 h data latency, and the 33 h `at` limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the plain-language purpose before any caveats, and each sentence carries distinct information (source, windows, resolution, latency, `at` behavior, nowcast, alternatives). It is dense and slightly long, but no sentence is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the remaining gaps an agent would care about - spatial/temporal coverage, data latency, measurement resolution, and where to go for adjacent needs - are all addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns extra by explaining semantics the schema cannot: that `at` selects totals ending at a past time with the newest frame up to 3 h before it, and that include_nowcast returns intensity/type/start time rather than an amount.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opening clause states a concrete verb+resource ('observed precipitation totals at a location') and immediately frames it with example questions. It distinguishes itself from named siblings get_period_totals (forecast) and get_observations (station reports) within the same description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use hedges ('did it rain', 'how much rain fell', 'rain in the last 24 hours') plus explicit routing to alternatives for the adjacent cases. An agent can pick this tool over the 30+ siblings without reading any schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_outlooksGet hazard outlooksA
Read-onlyIdempotent
Inspect

Hazard outlooks affecting a location. hazard=severe returns SPC convective outlooks (Day 1-8 categorical risk + tornado/wind/hail probabilities); hazard=fire returns SPC fire weather outlooks; hazard=rain returns WPC Excessive Rainfall Outlook polygons (days 1-3); hazard=heat returns the NWS HeatRisk index at the point (0 none .. 4 extreme, days 1-3). include_narrative=true adds the forecaster discussion for severe (SWO), fire (FWD), or rain (QPF/QPFERD; one PIL for all days). An empty result means no outlook covers the point -- not a failure. Examples: {"location": "Moore, OK", "hazard": "severe", "include_narrative": true} or {"location": "Phoenix", "hazard": "heat"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dayNoOutlook day for the narrative filter (1-8). Default 1.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
hazardNoHazard family: severe = SPC convective, fire = SPC fire weather, rain = WPC excessive rainfall, heat = NWS HeatRisk index. (Winter/WSSI is a planned expansion.)severe
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
include_narrativeNoInclude the forecaster narrative for the requested day (severe, fire, and rain).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dayYes
hazardYes
locationYes
outlooksYes
heat_riskNo
narrativeNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description goes beyond them by disclosing that an empty result means no outlook covers the point, not an error, and by clarifying narrative behavior for severe/fire/rain. This is meaningful behavioral context without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose first, then hazard-specific behavior, then the empty-result caveat, then examples. There is no filler and no restating of what the schema already provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full input schema and an output schema present, the description adds what remains necessary: per-hazard return semantics, narrative behavior, empty-result semantics, and representative usage examples. An agent has enough to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial meaning: severe returns Day 1-8 categorical risk plus tornado/wind/hail probabilities, rain returns Day 1-3 polygons, heat returns a 0-4 index, and narrative returns specific PILs. The examples also make valid parameter combinations concrete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with "Hazard outlooks affecting a location" and then precisely maps each hazard value to a concrete product family (SPC convective, SPC fire weather, WPC Excessive Rainfall, NWS HeatRisk). This is specific enough to distinguish the tool from the forecast, alerts, and observation siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to call the tool: when a user needs hazard outlooks at a point, and it explains how the hazard parameter selects among severe/fire/rain/heat. It does not explicitly name alternatives or when-not conditions, so it misses the top tier, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_path_exposureGet path exposureA
Read-onlyIdempotent
Inspect

Civic POIs (schools, hospitals, airports, …) inside a caller-supplied GeoJSON Polygon or MultiPolygon. Clip first — do not pass a CONUS HeatRisk dissolve. For an NWS alert use get_alerts({alert_id}) (already includes pois). Optional classes filters the civic allowlist. Example: {"geometry":{"type":"Polygon","coordinates":[[[-88.15,41.77],[-88.14,41.77],[-88.14,41.78],[-88.15,41.78],[-88.15,41.77]]]},"classes":["school"]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
classesNoOptional civic classes: airport, station, hospital, school, university, library, museum, stadium, cemetery, place_of_worship, camp_site, golf, attraction.
geometryYesGeoJSON Polygon or MultiPolygon. Clip national outlooks first.

Output Schema

ParametersJSON Schema
NameRequiredDescription
countsYes
featuresYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, so the description doesn't need to repeat those. It adds valuable behavioral context: the tool expects pre-clipped geometry and fails or misbehaves if given a CONUS HeatRisk dissolve, plus it notes that classes filter a civic allowlist. This goes beyond the structured annotations with practical constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: one purpose sentence, one warning/alternative sentence, one classes sentence, and an example. Every sentence adds distinct value—purpose, exclusion, parameter behavior, and a concrete illustration—without repeating schema content or padding. The example is long but earns its place as a callable template.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values need no explanation. Annotations cover safety traits, the schema documents both parameters, and the description adds the prerequisite, an alternative, and an example. Given the modest complexity (two params, one nested object), nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds extra meaning by warning against a specific problematic input ('do not pass a CONUS HeatRisk dissolve') and by providing a full JSON example that demonstrates the geometry and classes structure. The example hands an agent a concrete usage template beyond the schema's field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Civic POIs (schools, hospitals, airports, …) inside a caller-supplied GeoJSON Polygon or MultiPolygon', which names the resource (civic POIs), the operation (find inside a polygon), and the input type. It also distinguishes itself from get_alerts by stating that get_alerts is for NWS alerts and already includes POIs, giving agents a clear basis to choose this tool over a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use and when-not-to-use guidance: 'For an NWS alert use get_alerts({alert_id}) (already includes pois)' is a direct alternative, and 'Clip first — do not pass a CONUS HeatRisk dissolve' gives a concrete prerequisite with a negative example. This tells an agent exactly when to pick this tool versus another and how to prepare input.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_period_totalsGet period totalsA
Read-onlyIdempotent
Inspect

Aggregate a weather variable over one or more time periods. Returns server-computed totals, maxima, minima, or averages per period. Period start/end times should use the user's local timezone boundaries (not UTC midnight). Response includes the converted value and unit per period. Ideal for questions like "total rainfall today and tomorrow" or "peak wind speed this weekend". For rain that already fell, use get_observed_precipitation. Accepts a place name directly. Example: {"location": "Portland, OR", "variable": "precipitation", "aggregation": "sum", "periods": [{"start": "2026-07-08T07:00:00Z", "end": "2026-07-09T07:00:00Z", "label": "Today"}]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
unitsNoUnit system for all values in the request and response: imperial (°F, mph, inches), metric (°C, km/h, mm), or si (K, m/s, mm). Defaults to imperial.imperial
periodsYesTime periods to aggregate over (1-14).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
variableYesStandard variable name (e.g. precipitation, snowfall, temperature_2m, wind_speed_10m, cape).
dataset_idNoDataset override. Default: auto-resolved NBM for the location.
aggregationNoAggregation function. Default sum. Use sum for precipitation/snowfall, max for temperature/wind, min for low temperatures, avg for humidity/cloud cover.sum

Output Schema

ParametersJSON Schema
NameRequiredDescription
messageNo
periodsYes
locationYes
variableYes
aggregationYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely non-structured behavior: periods must use the user's local timezone boundaries rather than UTC midnight, and the response includes a converted value and unit per period. It stops short of covering failure modes or rate limits, so it earns a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then routing guidance, then the timezone caveat, then the example. Every sentence carries information, though the inline JSON example makes the block longer than strictly needed since the schema already defines the shape.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no further explanation, and the description still notes the per-period value/unit response. Combined with the timezone convention, the sibling routing, and the input example, an agent has everything needed to call this correctly for an 8-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real semantics beyond the schema: the local-timezone boundary rule for period start/end (the schema only says 'ISO 8601'), a worked example payload, and the note that a place name is accepted directly. These reduce the chance of malformed period timestamps, which the schema alone would not prevent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Aggregate a weather variable over one or more time periods') and immediately clarifies the computed outputs (totals, maxima, minima, averages) plus the per-period response shape. It explicitly names a sibling it must not be confused with ('For rain that already fell, use get_observed_precipitation'), so an agent can select it without opening sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage contexts ('total rainfall today and tomorrow', 'peak wind speed this weekend') and an explicit exclusion routing to get_observed_precipitation for past rainfall. The distinction between this tool (aggregation over periods) and the observational alternative is stated rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_platform_statusGet platform statusA
Read-onlyIdempotent
Inspect

Current data-freshness status of the weather platform: overall state, per-source states (ok / degraded / outage / no_signal), open incidents with cause attribution (provider outage vs internal processing delay), and active provider advisories. Use this when a user asks whether data is current, when other tools return surprisingly stale data, or before presenting time-critical weather. If a source is degraded or in outage, tell the user their data may be stale rather than presenting it as live. No inputs. Refreshed about every 5 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_okNotrue: list every monitored source including healthy ones. false (default): only sources that are not ok, keeping the response compact.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
overallYesWorst state across customer-facing data sources; "unknown" when status is unavailable.
sourcesNo
advisoriesNo
generated_atNoWhen the status document was generated (UTC).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful behavioral context beyond that: the status refreshes every ~5 minutes, incidents are categorized by cause, and the agent should warn users about potentially stale data. The phrase 'No inputs' is slightly misleading given the optional include_ok parameter, but this is more of a parameter-semantics issue than a behavioral one.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: it front-loads the purpose, then gives usage guidance and an important staleness caveat, and ends with refresh cadence. The 'No inputs' sentence is inaccurate and unnecessary, but the rest is tight and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description does not need to explain return values in depth. It covers when to use the tool, what the status categories mean, and how to handle degraded/outage states. The only notable gap is the misleading 'No inputs' statement, which is offset by the fully described input schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents include_ok with a clear description and default, so schema coverage is 100%. The description's 'No inputs' line is imprecise because the tool does accept an optional boolean, though it may be intended to mean 'no required inputs.' Overall, the description adds no real parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns 'current data-freshness status of the weather platform' and lists the specific contents: overall state, per-source states, open incidents, and provider advisories. This distinguishes it from sibling tools like get_current_conditions or get_observations, which return weather data rather than platform freshness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit use cases: when a user asks whether data is current, when other tools return surprisingly stale data, or before presenting time-critical weather. It does not name alternatives or state when not to use the tool, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_population_exposureGet population exposureA
Read-onlyIdempotent
Inspect

National population-exposure headline for a risk-zone outlook product: how many people are inside risk bands at or above min_level. Powers headlines like "~57M people under major heat risk tomorrow". hazard=heat covers NWS HeatRisk days 1-3 (levels: 1 minor, 2 moderate, 3 major, 4 extreme). Pass product_id directly for other risk-zone products. Example: {"hazard": "heat", "min_level": 3}.

ParametersJSON Schema
NameRequiredDescriptionDefault
hazardNoHazard family (expands the day-1..3 product set). Currently: heat (HeatRisk).
min_levelNoMinimum risk level to count (>=). Default 1 (any elevated risk).
product_idNoExplicit risk-zone product ID (overrides hazard), e.g. heatrisk_day1_conus.

Output Schema

ParametersJSON Schema
NameRequiredDescription
min_levelYes
summariesYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds valuable behavioral context beyond that: it counts people at or above a minimum risk level, covers only NWS HeatRisk days 1-3, and explains that product_id overrides hazard. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences plus a JSON example. The core behavior is front-loaded, the hazard scope is defined, and the product_id path is stated without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, all parameters are documented with 100% schema coverage, and annotations cover safety traits, the description completes the picture: it explains the product context, level semantics, hazard coverage, and how to handle other risk-zone products. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents hazard, min_level, and product_id with clear semantics. The description adds a worked example and reinforces the override relationship, but it does not materially expand on the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: get a national population-exposure headline counting people inside risk bands at or above min_level. It also grounds the purpose with a concrete example headline and separates it from generic forecasting tools by calling out risk-zone outlook products.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: hazard=heat covers NWS HeatRisk days 1-3, and product_id should be passed for other risk-zone products. It does not explicitly name sibling alternatives or state when not to use this tool, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_soundingGet radiosonde soundingA
Read-onlyIdempotent
Inspect

Nearest RAOB (radiosonde) vertical soundings to a point. Each sounding carries: profile (pressure-indexed thermodynamics: pressure_hpa, height_m, temperature_c, dewpoint_c, wind arrays), wind_profile (height-indexed winds for hodographs/shear), and derived indices (sbcape/mucape/mlcape + cin, lifted_index, k_index, total_totals, pwat_mm, freezing_level_m, lcl/lfc/el, bulk_shear_0_6km_kt). Soundings launch at 00Z/12Z so data can be hours old. Example: {"location": "Norman, OK"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoNumber of nearest soundings (1-5). Default 1.
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
locationYes
soundingsYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, openWorld, idempotent, and non-destructive behavior, so the bar is lower. The description adds useful behavioral context beyond annotations: soundings launch only at 00Z/12Z and can be hours old, and the result is the nearest sounding rather than an exact point measurement. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core operation and then gives a dense but useful breakdown of the returned data, followed by a compact example. The long data-fields sentence is somewhat verbose but earns its place by showing the tool's output shape and capabilities.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage and an output schema present, the description is nearly complete: it explains what the tool returns, gives an example, and warns about data freshness. The main gap is the lack of explicit routing between get_sounding and get_sounding_chart, and the n parameter is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with n, lat, lon, and location all described, so the baseline is 3. The description adds a concrete locator example ("Norman, OK") but does not independently clarify parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the nearest RAOB vertical soundings to a point and enumerates the contained data, so the core purpose is unambiguous. It does not explicitly contrast itself with sibling get_sounding_chart, so it falls just short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: call this when you need the nearest radiosonde profile data for a point, and it warns about data being hours old. However, it does not explicitly say when to prefer this over get_sounding_chart or mention any alternative tools or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sounding_chartGet sounding chartA
Read-onlyIdempotent
Inspect

Render the nearest RAOB (radiosonde) sounding as a Skew-T log-P + hodograph chart image for visual analysis: temperature/dewpoint traces, wind barbs, height-banded hodograph, and a derived-indices table (CAPE/CIN, lifted index, PWAT, shear, LCL). Soundings launch at 00Z/12Z so data can be hours old. Use get_sounding for the raw profile numbers. Example: {"location": "Norman, OK"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
unitNoTemperature axis display unitfahrenheit
scaleNoRaster scale factor (2 = retina; higher = larger image payload)
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate this is a read-only, idempotent, non-destructive operation. The description adds useful behavioral context beyond annotations by warning that 'Soundings launch at 00Z/12Z so data can be hours old,' which is critical for interpreting the chart's freshness. It also discloses what the rendered chart contains, giving agents realistic expectations about the output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and efficient: the first sentence defines the tool and output contents, the second provides a critical staleness caveat, and the third routes to the sibling tool. Every sentence contributes actionable information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the tool returns, the data source, the staleness caveat, and the relationship to get_sounding. With no output schema present, the listing of chart elements (temperature/dewpoint traces, wind barbs, hodograph, indices table) gives the agent a solid mental model of the result. Minor details like image format or error behavior are absent but not critical for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter in detail. The description adds only the concrete example '{"location": "Norman, OK"}', which demonstrates valid usage but does not materially enhance parameter understanding beyond the schema. A baseline of 3 is appropriate because the description does not need to compensate for missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and object: 'Render the nearest RAOB (radiosonde) sounding as a Skew-T log-P + hodograph chart image.' It enumerates the visual contents and explicitly separates itself from get_sounding by saying 'Use get_sounding for the raw profile numbers.' This makes the tool's purpose unmistakable and distinct among the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly frames the intended use case ('for visual analysis') and provides an explicit alternative ('Use get_sounding for the raw profile numbers'). It does not fully spell out when not to use this tool, but the visual-vs-raw distinction is enough to guide an agent's selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_storm_cellsGet storm cellsA
Read-onlyIdempotent
Inspect

Radar-identified storm cells near a location, merging NEXRAD Level III algorithm output from the nearest radar site: storm tracks (cell position, movement, forecast positions), hail index (probability of hail/severe hail + max expected size), mesocyclone detections (rotation), and TVS (tornado vortex signatures). Use during active convection to see what the radar algorithms flag. An empty result means no detected cells -- common outside active storms. Example: {"location": "Norman, OK"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
includeNoWhich detection families to include. Default: all.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
tracksNo
summaryYesReady-to-use one-liner. States explicitly when nothing was detected.
locationYes
detectionsYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/openWorld/idempotent annotations, the description discloses that output is merged from the nearest NEXRAD site and that empty results are normal outside storms. This prevents misinterpretation and adds real behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: result contents, usage condition, and empty-result interpretation. The product list is formatted with clear parenthetical details and the example is inline rather than an extra paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description doesn't need to restate return fields. It covers what the tool does, when to use it, and how to read an empty response, leaving no essential gap for an agent invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline applies. The description's mention of product families mirrors the `include` enum and adds an example, but it doesn't materially extend the parameter explanations already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (radar-identified storm cells) and enumerates the algorithm outputs it returns (storm tracks, hail index, mesocyclone detections, TVS). This clearly distinguishes it from weather siblings like get_storm_reports.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit use condition ('Use during active convection') and explains what an empty result means, which is essential for interpreting the tool. It doesn't name alternatives or exclusion criteria, but the context is sufficient for an agent to decide when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_storm_reportsGet storm reportsA
Read-onlyIdempotent
Inspect

Recent NWS Local Storm Reports (LSRs) -- verified reports of tornadoes, hail, damaging winds, flooding near a location. Use to confirm severe weather occurrence or assess reported damage. valid_time is event occurrence (UTC); cite in local time. Example: {"location": "Wichita", "hours": 12, "type": "H"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
typeNoReport type filter: T=tornado, H=hail, W=wind, F=flood, D=damage, S=snow.
hoursNoLookback window in hours (1-24). Default 6.
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hoursYes
reportsYes
locationYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnlyHint=true, idempotentHint=true, non-destructive), so the bar is lower. The description adds value beyond annotations by flagging that valid_time is event occurrence in UTC and should be cited locally, and by characterizing reports as 'verified' — a meaningful data-quality trait. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: definition, use case, time-zone caveat, and a concrete example. The core definition is front-loaded and the example is compact and illustrative without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich output schema and strong annotations, the description correctly focuses on what structure cannot convey: the UTC time semantics and the tool's evidentiary role for confirming severe weather. The only gap is the lack of explicit routing versus siblings such as get_alerts or get_storm_cells, which slightly weakens completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (lat, lon, type, hours, location) are already documented in structure. The description's example ('location: Wichita, hours: 12, type: H') shows composition of the params but adds no new meaning beyond the schema; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a precise resource ('NWS Local Storm Reports (LSRs)') with a specific verb ('get' implied by the name but the description defines the subject matter) and enumerates the content: verified reports of tornadoes, hail, damaging winds, flooding near a location. This clearly distinguishes it from siblings like get_storm_cells (radar storm features) and get_alerts (warnings/advisories), since LSRs are a distinct NWS post-event product.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states an explicit use case: 'Use to confirm severe weather occurrence or assess reported damage.' This tells an agent when this tool is the right choice. It does not name alternative tools or explicit exclusions, so agents might not know to prefer get_alerts for impending warnings or get_storm_cells for active tracking.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_time_contextGet time contextA
Read-onlyIdempotent
Inspect

Complete temporal context for a location: local time, timezone, 14-day calendar with day names and Today/Tomorrow offsets, sunrise/sunset/solar times (from the weather pipeline's astro product), and moon phase. Use whenever you need to reason about dates, times, or daylight for a location -- including "what time is sunset?", "is it dark there now?", or "what day of the week is the 4th-day forecast?". Accepts a place name directly. Example: {"location": "Seattle"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
moonYes
calendarYes
daylightYes
locationYes
current_timeYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, open-world, and non-destructive behavior. The description adds context beyond that: solar times come from the weather pipeline's astro product, it returns a 14-day calendar with Today/Tomorrow offsets, and it accepts a place name directly. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the first sentence states the resource and full contents, the second gives usage guidance with examples, and the third demonstrates the input format. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the read-only/idempotent annotations, 100% schema coverage, and presence of an output schema, the description fully supports tool selection and invocation. It covers what the tool returns, when to use it, and how to pass a location; structured fields handle the remaining return-format details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even without additional parameter detail. The description's 'Accepts a place name directly' and example add marginal nuance, but the schema already documents lat/lon and location semantics thoroughly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns complete temporal context for a location, enumerating local time, timezone, 14-day calendar, solar times, and moon phase. The rich content list and concrete example queries distinguish it clearly from sibling tools about forecasts, observations, alerts, and climate data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use whenever you need to reason about dates, times, or daylight for a location' and gives natural-language trigger examples like 'what time is sunset?' or 'is it dark there now?'. It does not name alternative tools or when-not-to-use conditions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tropicalGet tropical activityA
Read-onlyIdempotent
Inspect

Active NHC (National Hurricane Center) tropical systems: forecast cones, track lines, forecast points, coastal watches/warnings, and 7-day Tropical Weather Outlook formation areas -- Atlantic + East Pacific. Each feature carries a kind (cone | track | points | watch_warning | outlook_area) plus storm name, intensity, and timing properties. include_geometry=true adds full GeoJSON geometries (large). An empty result means no active tropical activity. Aircraft reconnaissance fixes and flight-level observations are in get_tropical_observations. Example: {} or {"include_geometry": true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_geometryNoInclude full GeoJSON geometries (cone/track polygons). Default false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
activeYes
featuresYes
feature_countYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey read-only, idempotent, non-destructive behavior. The description adds valuable context beyond those annotations: include_geometry=true triggers large payloads, empty results mean no active activity, and features carry kind/name/intensity/timing properties. This enriches the agent's mental model of what the tool will do and return.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description packs the resource, content taxonomy, behavior caveats, empty-result semantics, sibling pointer, and an example into four dense sentences with no fluff. Every sentence earns its place and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter and an existing output schema, the description covers all essential decision points: what is returned, how to enlarge geometries, what an empty result means, and where to find a specific related type of data. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single boolean parameter, so the baseline is 3. The description adds a performance caveat (large geometries) and a concrete example invocation, which slightly exceeds what the schema alone provides. The schema already explains the parameter, so the extra value is modest but real.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('Active NHC tropical systems') and clearly enumerates the returned content: forecast cones, track lines, forecast points, watches/warnings, and outlook areas. It also names the sibling tool for reconnaissance data, distinguishing itself without needing to inspect other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly defines its scope (Atlantic + East Pacific, active systems) and gives an exclusion: 'Aircraft reconnaissance fixes and flight-level observations are in get_tropical_observations.' This tells the agent when to choose this tool and when to prefer a sibling. The empty-result interpretation further clarifies the expected output in valid usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tropical_observationsGet tropical observationsA
Read-onlyIdempotent
Inspect

Hurricane Hunter aircraft reconnaissance for active tropical systems: vortex data message (VDM) center fixes — pressure, max flight-level wind, eye — and high-density (HDOB) flight-level observations. These are aircraft measurements, not surface weather stations (use get_observations for METAR). Use get_tropical for NHC cones, tracks, and watches; this tool is what the aircraft measured. kind=hdob returns the newest 200 records (about 1–2 h of flight; truncated says when more exist). Empty when no aircraft has flown in the window.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNovdm = vortex fixes; hdob = flight-level obs; all = both. Default vdm.vdm
hoursNoLookback window in hours (1-72). Default 24.
stormNoATCF id (al14) or storm name (polo). Unnumbered USAF missions never match.

Output Schema

ParametersJSON Schema
NameRequiredDescription
activeYes
stormsYes
featuresYes
truncatedYes
feature_countYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only, idempotent, and non-destructivecars. The description adds genuinely useful behavioral context beyond the schema: kind=hdob returns only the newest 200 records, 'truncated' indicates missing data, and the response is empty when no aircraft flew in the window. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and every sentence earns its place: it explains what the tool returns, redirects to two siblings, and discloses a key result-limiting behavior before the schema details. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with three optional, fully documented parameters and an output schema, the description covers the important operational details: data source, sibling tool boundaries, HDOB limits/truncation, and empty-response behavior. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all three parameters. The description adds valuable runtime semantics for kind=hdob (200-record limit, roughly 1–2 hours of data, truncated indicator), which goes beyond the bare enum definition. It does not add much for hours or storm, but the schema handles those adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource ('Hurricane Hunter aircraft reconnaissance'), names the data products (VDM and HDOB), and explicitly distinguishes itself from get_observations and get_tropical. An agent can immediately tell what this tool returns and what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: use get_observations for METAR/surface stations)Skip and get_tropical for NHC cones/tracks/watches. It also clarifies that this tool returns aircraft measurements. This is clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_datasetsList datasetsA
Read-onlyIdempotent
Inspect

Discover the datasets (model grids, analyses, observations) available at a location, with per-dataset freshness (data age, latest model run). Datasets vary by domain (CONUS/Alaska/Hawaii). Use this to find dataset_id values for query_dataset and describe_dataset, or to assess whether data is current before making decisions. Example: {"location": "Anchorage"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
include_freshnessNoInclude per-dataset data age and run times. Default true.

Output Schema

ParametersJSON Schema
NameRequiredDescription
datasetsYes
locationYes
freshnessNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish safety and idempotence (readOnlyHint, openWorldHint, idempotentHint, destructiveHint: false). The description adds useful behavioral context beyond annotations, including per-dataset freshness, latest model run, and the fact that datasets vary by CONUS/Alaska/Hawaii domain, without contradicting any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences front-load the core behavior, then provide practical usage guidance and a concrete example. Every sentence earns its place and there is no redundant restatement of schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full input schema, an output schema, and annotations covering safety and idempotence, the description adds the remaining needed context: what datasets are listed, why the freshness matters, how to use it with related tools, and a concrete invocation example. This is complete for correct tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with detailed descriptions of lat, lon, location, and include_freshness, so the description is not required to repeat parameter details. The description adds no substantive parameter meaning beyond the example location, which is sufficient given the schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Discover') and a specific resource ('datasets available at a location'), naming the dataset categories (model grids, analyses, observations) and the freshness information returned. It is clearly distinct from siblings like get_forecast or query_dataset, and it also explains its downstream relationship to query_dataset and describe_dataset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: to find dataset_id values for query_dataset/describe_dataset and to assess data currency. It does not explicitly describe when not to use it, but the intended use cases are clear enough to guide selection among a large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_datasetQuery datasetA
Read-onlyIdempotent
Inspect

Raw time series from a specific dataset for specific variables at a point. Power-user access to any gridded product (NBM, HRRR, RRFS, GFS, RTMA, MRMS, air quality, ...). Time modes: hours (next N hours, default 24), time_start+time_end (explicit ISO-8601 window), or latest=true (single most-recent value). reference_time pins a specific model run, and each returned series reports the run that served it (reference_time, or reference_times when a series mixes runs) — check it before comparing two runs, since a run older than about 48 hours may no longer be available. For blended forecasts use get_forecast instead. Examples: {"location": "Denver", "dataset_id": "rrfs_surface", "variables": ["temperature_2m"], "hours": 18} or {"lat": 41.4, "lon": -92.9, "dataset_id": "rtma_conus", "variables": ["temperature_2m"], "latest": true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
latNoLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonNoLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).
hoursNoForecast/lookahead hours from now (1-264). Default 24 when no other time mode set.
latestNoReturn only the most recent value (analysis datasets like RTMA/MRMS).
locationNoFree-text place: city ("Denver"), city+state ("Portland, OR"), US ZIP ("50219"), or "lat,lon" ("39.74,-104.99"). Provide either this OR explicit lat+lon, not both.
time_endNoISO 8601 window end (with time_start).
variablesYesStandard variable names (e.g. temperature_2m, precipitation). Discover with describe_dataset.
dataset_idNoDataset to query. Default: the NBM dataset for the location domain (nbm_conus/nbm_alaska/nbm_hawaii). Discover options with list_datasets.
time_startNoISO 8601 window start (with time_end).
reference_timeNoPin a specific model run (ISO 8601). Default: latest run.

Output Schema

ParametersJSON Schema
NameRequiredDescription
seriesYes
locationYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so no safety concerns need repeating. The description adds valuable behavioral context: run availability (a run older than about 48 hours may no longer be available), returned series reports reference_time or reference_times, and that results may mix runs. This goes beyond the annotations' safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it opens with the core purpose, then covers time modes, run-pinning caveat, sibling routing, and examples. Every sentence earns its place. Slightly dense, but effective for a power-user tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key usage patterns (hours, window, latest), the reference_time caveat, and sibling differentiation. The output schema exists, so return format details are not the description's job. The only minor gap is that it doesn't explain what happens if no time mode is specified beyond hours defaulting to 24, but the schema covers hours default. Overall nearly complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters. The description adds context by explaining time modes and reference_time semantics, but doesn't need to repeat parameter definitions. Baseline 3 is appropriate since the schema carries the heavy lifting and the description supplements with usage context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: query raw time series from a specific dataset for specific variables at a point, and explicitly names the power-user scope (NBM, HRRR, RRFS, GFS, RTMA, MRMS, air quality). It distinguishes itself from get_forecast by saying blended forecasts should use get_forecast instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool vs alternatives: 'For blended forecasts use get_forecast instead.' It also explains the three time modes (hours, time_start+time_end, latest=true) and reference_time behavior, which tells an agent exactly which parameters to set for a given scenario. The examples further illustrate valid usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reverse_geocodeReverse geocodeA
Read-onlyIdempotent
Inspect

Resolve coordinates to a human-readable place (city, state, county, timezone). Use when you have lat/lon but need a display name or the local timezone. Example: {"lat": 39.74, "lon": -104.99} -> Denver, Colorado, America/Denver.

ParametersJSON Schema
NameRequiredDescriptionDefault
latYesLatitude in decimal degrees (-90 to 90). Most tools also accept a `location` place-name string instead of lat/lon.
lonYesLongitude in decimal degrees (-180 to 180). For continental US use negative values (west of the prime meridian).

Output Schema

ParametersJSON Schema
NameRequiredDescription
latYes
lonYes
placeYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is known. The description adds concrete behavioral output (including timezone) and an example, with no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences: purpose, usage condition, example. Every sentence earns its place; no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, read-only tool with an output schema, the description conveys purpose, usage, and a worked example. It does not explicitly correct the misleading 'location' note in the schema or address edge cases, so it falls just short of full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers both parameters with ranges and descriptions (100% coverage), so baseline is 3. The description adds a concrete example mapping lat/lon to a result, though the schema's lat description misleadingly mentions a 'location' string that additionalProperties=false contradicts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Resolve') and resource ('coordinates') and enumerates the output categories (city, state, county, timezone). The example concretely differentiates it from forward geocoding (search_locations) and weather-focused siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction 'Use when you have lat/lon but need a display name or the local timezone' provides an explicit condition for selection. It does not name sibling alternatives or exclusions, but the condition is sufficiently clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_locationsSearch locationsA
Read-onlyIdempotent
Inspect

Resolve a place query to candidate locations with coordinates. Accepts city names ("Denver"), city+state ("Portland, OR" via query), ZIP codes ("50219"), or partial input with fuzzy=true for autosuggest-style matching ("bost" -> Boston). Returns ranked candidates with lat/lon. Most weather tools accept a location string directly and geocode internally -- use this tool only to disambiguate ("which Springfield?") or to present location choices to the user. Example: {"query": "Springfield"} returns all major Springfields ranked by place importance.

ParametersJSON Schema
NameRequiredDescriptionDefault
fuzzyNoAutosuggest mode for partial/misspelled input. Default false (exact search).
limitNoMaximum candidates to return (1-10). Default 5.
queryYesPlace query: city, "city, state", ZIP, or partial text with fuzzy=true.

Output Schema

ParametersJSON Schema
NameRequiredDescription
queryYes
candidatesYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds behavioral detail beyond that: accepted query formats, fuzzy matching behavior, ranked candidates, and lat/lon output. It does not contradict annotations and gives useful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose. Every sentence adds useful information: accepted formats, fuzzy behavior, output type, usage boundaries, and a concrete example. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema, output schema, and annotations, the description covers the remaining context an agent needs: when to invoke it, what inputs are accepted, what output shape to expect, and how it relates to sibling tools. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds extra semantic value by giving concrete examples such as 'Portland, OR' via query, '50219', and 'bost' -> Boston with fuzzy=true, plus an example explaining that Springfields are ranked by importance. This goes beyond the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resolve a place query to candidate locations with coordinates.' It clearly states what the tool does and distinguishes it from siblings by noting that most weather tools geocode internally and this tool is only for disambiguation or presenting choices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'use this tool only to disambiguate' or 'to present location choices to the user,' and contrasts with most weather tools that accept a location string directly. This gives an agent clear routing guidance and prevents unnecessary calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addedget_observed_precipitation
  2. 1 tool update
    • Changedget_current_conditions4 fields changed
      • addedInput schema / properties / units
        Added value: +{
        +  "default": "si",
        +  "description": "Unit system for the analysis values: si (default; K, m/s, Pa, m, as RTMA serves them), imperial (°F, mph, inHg, mi, ft) or metric (°C, km/h, hPa, km, m). Unlike get_forecast, which defaults to imperial. Does not apply to nearest_station.",
        +  "enum": [
        +    "imperial",
        +    "metric",
        +    "si"
        +  ],
        +  "type": "string"
        +}
      • changedOutput schema / properties / analysis / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "data": {
        -        "additionalProperties": {
        -          "items": {
        -            "type": [
        -              "number",
        -              "null"
        -            ]
        -          },
        -          "type": "array"
        -        },
        -        "type": "object"
        -      },
        -      "dataset_id": {
        -        "type": "string"
        -      },
        -      "valid_times": {
        -        "items": {
        -          "type": "string"
        -        },
        -        "type": "array"
        -      }
        -    },
        -    "required": [
        -      "dataset_id",
        -      "valid_times",
        -      "data"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "data": {
        +        "additionalProperties": {
        +          "items": {
        +            "type": [
        +              "number",
        +              "null"
        +            ]
        +          },
        +          "type": "array"
        +        },
        +        "type": "object"
        +      },
        +      "dataset_id": {
        +        "type": "string"
        +      },
        +      "derived": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": "array"
        +      },
        +      "notes": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": "array"
        +      },
        +      "reference_time": {
        +        "type": "string"
        +      },
        +      "reference_times": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": "array"
        +      },
        +      "unit_labels": {
        +        "additionalProperties": {
        +          "type": "string"
        +        },
        +        "type": "object"
        +      },
        +      "valid_times": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": "array"
        +      }
        +    },
        +    "required": [
        +      "dataset_id",
        +      "valid_times",
        +      "data"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedOutput schema / properties / units
        Added value: +{
        +  "enum": [
        +    "imperial",
        +    "metric",
        +    "si"
        +  ],
        +  "type": "string"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "location",
        -  "analysis",
        -  "nearest_station"
        -]New value: +[
        +  "location",
        +  "units",
        +  "analysis",
        +  "nearest_station"
        +]
  3. 1 tool update
    • Addedget_tropical_observations
  4. 1 tool update
    • Changedfind_best_window3 fields changed
      • addedInput schema / properties / activity
        Added value: +{
        +  "description": "Free-form activity to rate the windows for (e.g. \"afternoon golf\"). Accepted even when rating is off so clients can send it.",
        +  "maxLength": 200,
        +  "type": "string"
        +}
      • addedOutput schema / properties / windows / items / properties / limiting_factor
        Added value: +{
        +  "enum": [
        +    "wind",
        +    "precip",
        +    "temperature",
        +    "visibility",
        +    "storms",
        +    "none"
        +  ],
        +  "type": "string"
        +}
      • addedOutput schema / properties / windows / items / properties / rating
        Added value: +{
        +  "enum": [
        +    "unworkable",
        +    "poor",
        +    "workable",
        +    "good",
        +    "ideal"
        +  ],
        +  "type": "string"
        +}
  5. 1 tool update
    • Addedget_path_exposure
  6. 1 tool update
    • Addedget_forecast_skill_map

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables to interact with comprehensive weather data through the MCP protocol, including current conditions, multi-day forecasts, hourly forecasts, and geocoding.
    16 npm
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Provides real-time US weather data for AI assistants via MCP, including current conditions, forecasts, alerts, severe weather outlooks, radar, upper-air analysis, and surface analysis. Supports optional personal weather station integration.
    9
    4
    ISC
  • F
    license
    Not graded
    quality
    C
    maintenance
    MCP server that provides current weather, multi-day forecasts, umbrella recommendations, severe weather alerts (US), and side-by-side city comparisons using Open-Meteo and NWS APIs, with no API key required.
    -
  • F
    license
    A
    quality
    D
    maintenance
    A comprehensive MCP server providing tools for real-time, forecast, and historical weather data, alongside air quality, marine conditions, and climate projections. It also includes geocoding services to search for locations and retrieve precise coordinates for environmental analysis.
    7
    -
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.