Weather That Doesn't Suck
Server Details
Weather decisions for outdoor plans: should I go, when, what to wear, what you will meet en route.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 11 tools
Several tools overlap meaningfully: get_trail_weather vs route_weather both cover trail/route conditions, get_weather vs get_forecast both return general weather, and clothing_recommendation duplicates the 'what to wear' section built into outdoor_activity_weather and get_trail_weather. The descriptions do explicitly tell the agent when to prefer each (e.g. 'should I?' vs 'what is it?'), which saves it from a lower score, but misselection is still likely.
There is a solid verb_noun cluster (get_weather, get_forecast, get_trail_weather, list_trails, find_best_weather_window), but it is mixed with bare verbs (search, fetch, explain_weather) and noun-phrase tools (clothing_recommendation, outdoor_activity_weather, route_weather). Readable but not a single predictable pattern.
Eleven tools is well within the sensible range for a domain this rich (current, forecast, activity decisions, routes, trails, clothing). Only minor bloat: fetch and search exist purely as connector-compatibility shims that duplicate get_weather.
Coverage is broad and lifecycle-like: current conditions, forecasts, activity verdicts, best-window finding, route/trail sampling, clothing, and plain-language explanation. Gaps are minor (no historical weather, no standalone alerts tool), and nothing essential to the stated purpose is missing.
Available Tools
11 toolsclothing_recommendationWhat should I wear?ARead-onlyIdempotentInspect
What to wear for an activity at a time and place, short and practical, followed by the weather it rests on. It dresses you for the effort, not just the temperature: a runner is warmer than a person standing still at the same feels-like, and the same rules as the website's outfit card apply. Optional: intensity, and preferences such as { "feel": "runs_cold" } or { "avoid": ["shorts"] }.
| Name | Required | Description | Default |
|---|---|---|---|
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| activity | No | What the person is doing, if known. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| intensity | No | How hard the effort is, which changes what to wear. Defaults to the usual effort for the activity. | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. | |
| start_time | No | When. Defaults to now. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| preferences | No | Optional: { "feel": "runs_cold" | "runs_warm", "avoid": ["shorts"] }. | |
| duration_hours | No | How long they will be out. Defaults to the usual length for the activity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read-only, idempotent, non-destructive profile, so the description is free to add substantive behavior and does: it explains the effort-adjustment model (a runner is warmer than someone standing still at the same feels-like) and describes the output shape (short practical advice followed by the underlying weather). The opaque 'same rules as the website's outfit card' reference is the one weak spot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the core purpose and kept to three sentences with no padding. The middle 'website's outfit card' clause is the only sentence that does not clearly earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 9 parameters (none required) and no output schema, the description tells the agent what comes back — practical advice plus the weather it depends on — and the schema fully documents every input including defaults and time-parsing behavior. Nothing essential to calling the tool correctly is missing, and the remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter — including the nested preferences object, units, and time decoding rules — is already documented in the schema. The description only restates intensity and preference examples, adding little beyond that. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (what to wear) and scopes it to an activity, time, and place, with the outfit returned alongside the weather it rests on. It is clearly differentiated from weather siblings like get_forecast or outdoor_activity_weather because it produces clothing advice, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer it should be called when a wear recommendation is wanted, and it notes what is optional (intensity, preferences). There is no explicit when-to-use, when-not-to-use, or routing against the weather siblings, and the reference to 'the website's outfit card' points outside the toolset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_weatherExplain the weatherARead-onlyIdempotentInspect
Explains weather in plain terms for conversation: why it feels colder than the temperature, whether rain will matter on a hike, whether wind matters, whether the trail will be muddy, humidity, fog, frost, pressure. It separates what the forecast says (facts, each labelled forecast or formula) from what we are reasoning (inferences, each with its basis and how sure we are) and lists what we cannot see (unknowns). Pass the question as asked, or a topic. If it cannot tell what was asked it says so and gives the general picture.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | The topic, if you know it. Otherwise it is read from the question. | |
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| activity | No | What the person is doing, if known. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| question | No | The question as the person asked it, for example "Why does it feel colder than the temperature?". | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. | |
| start_time | No | The time the question is about. Defaults to now. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| duration_hours | No | How many hours to look at. Default 6. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, idempotent, open-world safety profile, so the description's value is in the behavior it adds: it discloses the response structure (facts labelled forecast or formula, inferences each with basis and confidence, an explicit list of unknowns) and failure behavior ('if it cannot tell what was asked it says so and gives the general picture'). That is substantive disclosure beyond the structured fields, and none of it contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the opening clause and every sentence carries information about capability, output shape, or fallback. It runs long for a tool description, but the length is earned by the breadth of topics covered; there is little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 optional parameters, no output schema, and full schema coverage, the description supplies the missing return-value semantics (facts vs inferences vs unknowns) and the ambiguity fallback, which an agent needs to interpret results. Minor gaps remain, such as how it resolves location vs latitude/longitude conflicts, but those are documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including enums, units defaults, and time formats. The description only adds 'pass the question as asked, or a topic' and the topic list, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — explaining weather in plain conversation terms — and enumerates the kinds of questions it answers (feels-like, rain impact on a hike, mud, frost, pressure). This clearly separates it from data-fetching siblings like get_weather or get_forecast, which the description never names but implicitly contrasts by emphasizing reasoning over raw values.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the conversational framing ('for conversation', 'pass the question as asked, or a topic') signals when this tool fits, but there is no explicit when-to-use/when-not guidance and no named alternatives (get_weather, get_trail_weather, clothing_recommendation) that an agent should prefer for simple lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchGet the weather for a place found by searchARead-onlyIdempotentInspect
Compatibility tool for ChatGPT connectors. Returns the weather for an id from search. For real use prefer get_weather.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | An id from `search`, like place:43.79,-88.05. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description adds the non-obvious behavioral fact that this is a legacy compatibility shim whose real-world recommendation is to use another tool, which is exactly the kind of context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with zero waste; the compatibility framing is front-loaded so the agent immediately understands the tool's status before reading the mechanics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool the description covers purpose, origin of the id, and preferred alternative. It stops short of describing the returned weather payload, which is a minor gap given no output schema exists, but the annotations already cover the safety and idempotency profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single `id` parameter is fully documented in the schema, including its format example. The description merely restates that the id comes from `search`, adding no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Returns the weather for an id from `search`') and explicitly positions itself against the sibling `get_weather`, so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the use case ('Compatibility tool for ChatGPT connectors') and the exclusion with the preferred alternative ('For real use prefer get_weather'), leaving nothing to inference about when to pick this tool over the sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_best_weather_windowWhen is the best time?ARead-onlyIdempotentInspect
Finds the best time on a day for an activity. Returns the best window, an alternate, and any clearly bad stretch, each with a short reason. Windows are stretches, not exact hours: when several start times are about equally good they are reported as one stretch, and when the whole day is about the same it says so. Use it for "when should I go?", "what's the best time Saturday?", "Kilian and I want to hike Saturday". Give a date such as "Saturday", "tomorrow" or 2026-10-03. If nothing on the day is good it says so and names the least bad stretch.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | The day to search. Defaults to today. A part of a day ("tomorrow morning") narrows the search to it, and "this weekend" compares both days. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| activity | Yes | What the person is doing. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| intensity | No | How hard the effort is, which changes what to wear. Defaults to the usual effort for the activity. | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. | |
| duration_hours | No | How long the activity lasts. Defaults to the usual length for the activity, and says so. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description goes well beyond them: it explains that windows are stretches rather than exact hours, that equally-good start times are collapsed, that a uniform day is reported as such, and that a bad day yields a least-bad stretch with a reason. This is exactly the behavioral context an agent needs to interpret the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and output shape are front-loaded, then usage examples, then the date input hint. It is a bit long and repeats the 'says so when the day is uniform / nothing is good' idea, but every sentence still carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain the return shape, and it does: best window, alternate, bad stretch, each with a short reason, plus fallback behavior on poor days. With 8 parameters all schema-documented and annotations covering safety, nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, including date syntax, units, and activity enums. The description largely restates the date guidance ('Give a date such as Saturday, tomorrow or 2026-10-03') rather than adding new meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Finds the best time on a day for an activity') and immediately enumerates what it returns (best window, alternate, bad stretch). This clearly separates it from forecast/weather siblings like get_forecast or outdoor_activity_weather, which report conditions rather than a recommended time window.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete trigger phrases ('when should I go?', 'what's the best time Saturday?') and the acceptable date forms, which tells the agent exactly when this tool is the right pick. It stops short of naming an alternative sibling or stating when not to use it, so it does not reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_forecastForecast for a time periodARead-onlyIdempotentInspect
The forecast for a stretch of time: "What's Saturday looking like?", "What will conditions be between 6 PM and midnight?", "this weekend". Give start_time (and optionally end_time or duration_hours). Returns a one-line summary, the temperature range, rain, wind, any conditions worth acting on, and the hourly and daily rows. The forecast reaches about seven days. Beyond that it says so and does not guess.
| Name | Required | Description | Default |
|---|---|---|---|
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| end_time | No | When the period ends, in the same formats as start_time. "midnight" means the next midnight. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. | |
| start_time | No | When the period starts. Defaults to now. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| duration_hours | No | Instead of end_time: how many hours to cover. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, openWorld), and the description still adds real behavioral context: the ~7-day horizon, that the tool admits rather than guesses beyond it, and that a bare hour without AM/PM is refused instead of inferred. This is exactly the kind of guardrail behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose with illustrative phrasing, then moves to parameters, return shape, and the horizon limit. The example phrases cost a little space but earn it by clarifying accepted time formats; overall tight and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of describing returns and it does so (one-line summary, temperature range, rain, wind, actionable conditions, hourly and daily rows). Combined with the horizon caveat and time-format guidance, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value by specifying the intended invocation pattern (start_time required in practice, end_time or duration_hours as optional alternatives). It complements rather than repeats the rich per-parameter schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (forecast over a time period) and illustrates scope with concrete query examples ('this weekend', 'between 6 PM and midnight'). It implies a distinction from the current-conditions sibling get_weather via the time-range framing, but never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for invocation ('What's Saturday looking like?', 'this weekend') and tells the agent the canonical call shape: give start_time, optionally end_time or duration_hours. It stops short of naming when to prefer sibling tools like find_best_weather_window or get_trail_weather.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trail_weatherWeather on a trailARead-onlyIdempotentInspect
Weather for a trail over time, sampled along the trail rather than at one point: starting conditions, how the temperature changes through the trip, rain windows, wind and gusts, humidity and dew point, wind chill and heat index, sunrise, sunset, twilight, the moon, significant changes, alerts, and what to wear and carry. Give a trail_id (call list_trails) or the route itself, plus start_time and a pace or total_duration_hours. Returns a timeline every two hours and a bottom line. It does not know the terrain, shade or exposure unless the trail record says so, and it says that.
| Name | Required | Description | Default |
|---|---|---|---|
| gpx | No | The text of a GPX file (track points). Up to about 450 KB. | |
| pace | No | How fast the person moves, as a value and a unit. | |
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| geojson | No | GeoJSON as an object or a string: a LineString, MultiLineString, or Features holding them. Longitude first, as GeoJSON has it. | |
| activity | No | What the person is doing, if known. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| polyline | No | An encoded polyline (Google format). | |
| progress | No | For someone already on the route: how far along they are right now. The answer is for the rest of the route, starting now. Give distance_miles or distance_km, or a latitude and longitude. The position is used for this answer only and is never stored or logged. | |
| trail_id | No | A trail we know by name. Call list_trails to see them. Use this OR one of the route formats below. | |
| intensity | No | How hard the effort is, which changes what to wear. Defaults to the usual effort for the activity. | |
| waypoints | No | Two to twelve places to travel between, each a name like "Glenbeulah, WI" or { "latitude": 43.79, "longitude": -88.05 }. They are joined by straight lines, so distances are approximate. | |
| start_time | No | When the trip starts. Defaults to now. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| step_hours | No | Hours between timeline rows. Default 2. | |
| coordinates | No | A list of [latitude, longitude] pairs, latitude FIRST. | |
| preferences | No | Optional: { "feel": "runs_cold" | "runs_warm", "avoid": ["shorts"] }. | |
| out_and_back | No | True if the route is walked out and then back along the same line. | |
| polyline_precision | No | Digits of precision in the polyline. 5 is usual, 6 is what some routers send. | |
| total_duration_hours | No | Instead of a pace: how long the whole trip is expected to take. The pace is worked out from the length of the route. | |
| break_minutes_per_hour | No | Minutes of stops per hour of moving, which stretches arrival times. | |
| rest_at_turnaround_hours | No | Hours spent resting at the turn of an out-and-back (a night's sleep, say). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description adds useful context: it returns a two-hourly timeline plus a bottom line, and it flags that it does not know terrain/shade/exposure and says so. It does not discuss rate limits, auth, or cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first clause, then enumerates outputs and inputs. The long output enumeration is dense but earns its place given there is no output schema; the only mild cost is a slightly run-on feel.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-parameter, no-output-schema tool, the description covers input modes, temporal coverage, timeline granularity, and caveats well. Return-format explanation is necessarily present since no output schema exists, and the coverage is adequate though it does not discuss missing-data or error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 19 parameters in detail. The description adds only marginal framing, restating the trail_id-or-route choice and the pace-or-duration choice that the schema already spells out. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("weather for a trail over time") and immediately distinguishes the sampling model from a point forecast ("sampled along the trail rather than at one point"). It is clear what the tool produces, though it never names the very similar sibling route_weather, so the agent must infer that distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete entry conditions: supply trail_id (routing to list_trails) or a route, plus start_time and either a pace or total_duration_hours. It also states a limitation (no terrain/shade/exposure knowledge). It stops short of explicit when-not guidance or naming an alternative tool for point forecasts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_weatherWeather now and the days aheadARead-onlyIdempotentInspect
Plain-language weather for a place: one sentence first, then current conditions, the next 24 hours, the next 7 days, and official alerts. Use it for "what's the weather", "do I need a jacket", "is it going to rain". For a decision about doing something outdoors (a run, a hike, staining a deck), use outdoor_activity_weather instead: it answers "should I?", not just "what is it?". data.current says whether it is a station observation, a model estimate, or a forecast hour. data.alerts.status says whether "no alerts" can be believed: only "ok" means we checked. Answers carry their sources, how old the data is, and say so when it came from cache.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | Hours of hourly detail to include. Default 24. | |
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld/non-destructive, so the description is free to add the harder context and does: `data.current` distinguishes station observation, model estimate, and forecast hour, `data.alerts.status` distinguishes a verified "no alerts" from an unchecked one, and results carry sources, data age, and cache status. These are trust-critical traits no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and scoped tightly, with no filler paragraphs. It runs long across several clauses and quote-heavy examples, but each sentence carries distinct routing or behavioral information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining the meaning of key response fields (`data.current`, `data.alerts.status`) and the provenance/age/cache behavior of answers. An agent knows what to expect back and how much to trust it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents hours, units, location, latitude, and longitude. The description adds no parameter-level detail beyond that, which matches the baseline 3 when structured fields do the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource (plain-language weather for a place) and enumerates exactly what the answer contains: one sentence, current conditions, 24 hours, 7 days, alerts. It also distinguishes itself from the nearest sibling by naming outdoor_activity_weather and contrasting "what is it?" with "should I?".
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies trigger phrases ("what's the weather", "do I need a jacket", "is it going to rain") and an explicit exclusion with a named alternative for outdoor-decision queries. This is exactly the when/when-not/alternative structure the rubric rewards.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_trailsTrails we knowARead-onlyIdempotentInspect
Lists the trails WTDS knows by name, with length, bounds and whether we have data on which sections are exposed. Use a trail's id as trail_id in get_trail_weather or route_weather. If a trail is not listed, send the route itself.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, openWorld, so the safety profile is covered. The description adds meaningful behavioral context by naming the exact fields returned and the 'known trails' scope, which is what makes the empty-result fallback necessary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: purpose and return fields first, then the downstream id usage, then the fallback. No filler and the most decision-relevant information leads.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description compensates by naming the returned fields (name, length, bounds, exposure coverage) and explaining what to do when a trail is missing. Complete for a zero-parameter discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description still adds value by establishing the id/trail_id naming contract consumed by downstream tools, linking this call's output to other tools' inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the trails WTDS knows by name') and immediately enumerates the returned attributes (length, bounds, exposure data). This distinguishes it from the sibling search/fetch tools that return arbitrary web results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent onward: use a trail's id as trail_id in get_trail_weather or route_weather, and if a trail is absent, send the route itself. Both the when-to-use and the fallback are named, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outdoor_activity_weatherShould I do this outside?ARead-onlyIdempotentInspect
Our flagship tool. Turns weather into a decision for a specific outdoor activity: a practical verdict (go, go_with_caveats, marginal, no_go), a one-sentence reason, the temperature and feels-like range, rain, wind, sunset, what to wear, and the concerns. Activities: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Give a start_time and a duration_hours, or a phrase like "tomorrow morning". If you name a whole day ("Saturday"), it finds and judges the best window inside it. It keeps the forecast apart from our advice, says what it is inferring (mud, ice, fog) and what it cannot know (the trail surface), and an official warning overrides everything. To find WHEN to go, use find_best_weather_window.
| Name | Required | Description | Default |
|---|---|---|---|
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| activity | Yes | What the person is doing. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| latitude | No | Latitude in degrees. Use with longitude instead of location. | |
| location | No | A place name such as "Glenbeulah, WI" or "Greenbush, Wisconsin", or coordinates as "43.79,-88.05". Add the state or country for small towns, because a bare "Greenbush" matches several places. | |
| intensity | No | How hard the effort is, which changes what to wear. Defaults to the usual effort for the activity. | |
| longitude | No | Longitude in degrees, west negative. Use with latitude instead of location. | |
| start_time | No | When it starts. Defaults to now. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| duration_hours | No | How long it lasts. Defaults to the usual length for the activity, and says so. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), and the description goes well beyond them: it separates forecast from advice, discloses what it infers (mud, ice, fog) versus what it cannot know (trail surface), and states that an official weather warning overrides everything. It also effectively documents the return shape in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the output contract before the input guidance, and almost every clause carries information. The opening 'Our flagship tool' is marketing filler that does not help selection, and the middle run-on sentence about inference versus unknowables is dense, but overall the length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden by enumerating the verdict values and the returned fields. Combined with the input guidance and the explicit warning-override rule, an agent has everything needed to call this correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics on top: naming a whole day makes the tool find and judge the best window inside it, and the relationship between start_time and duration_hours as an alternative to a natural-language phrase. The activity enum is restated from the schema rather than extended, which caps this at 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific transformation of a specific resource: weather turned into a go/no_go verdict for a named outdoor activity, and enumerates the returned fields (reason, temp/feels-like, rain, wind, sunset, what to wear, concerns). An agent can distinguish it from get_forecast or get_weather without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the WHEN-to-go case to a named sibling: 'To find WHEN to go, use find_best_weather_window.' It also explains the accepted forms of the temporal input (start_time + duration_hours, or a phrase, or a whole day which triggers best-window selection), so the calling conditions are fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
route_weatherWeather along a routeARead-onlyIdempotentInspect
What weather you will meet as you move along a route, at the moment you reach each part of it. Not "the weather on the trail" but "what will I encounter, where, and when". Accepts a known trail_id, a GPX file, GeoJSON, an encoded polyline, coordinates, or place names. With a start_time and a pace it works out when you reach each checkpoint, reads the forecast for that place at that time, and reports what changes: rain starting, the cold coming in, the light going. For someone already out there, pass progress (how far they are, or where) and it answers for the rest of the route from now. Returns checkpoints with arrival time, place, weather and light, the changes, and a bottom line. Use get_trail_weather for an hour-by-hour timeline and clothing advice.
| Name | Required | Description | Default |
|---|---|---|---|
| gpx | No | The text of a GPX file (track points). Up to about 450 KB. | |
| pace | No | How fast the person moves, as a value and a unit. | |
| units | No | imperial (F, mph, inches) or metric (C, km/h, mm). Defaults to imperial in the United States and metric elsewhere. | |
| geojson | No | GeoJSON as an object or a string: a LineString, MultiLineString, or Features holding them. Longitude first, as GeoJSON has it. | |
| activity | No | What the person is doing, if known. One of: walking, hiking, trail_running, road_running, cycling, camping, hunting, fishing, yard_work, construction, painting_finishing, general_outdoor. Plain words work too ("run", "hike", "stain the deck"). "Running" with no more words is treated as road_running. | |
| polyline | No | An encoded polyline (Google format). | |
| progress | No | For someone already on the route: how far along they are right now. The answer is for the rest of the route, starting now. Give distance_miles or distance_km, or a latitude and longitude. The position is used for this answer only and is never stored or logged. | |
| trail_id | No | A trail we know by name. Call list_trails to see them. Use this OR one of the route formats below. | |
| intensity | No | How hard the effort is, which changes what to wear. Defaults to the usual effort for the activity. | |
| waypoints | No | Two to twelve places to travel between, each a name like "Glenbeulah, WI" or { "latitude": 43.79, "longitude": -88.05 }. They are joined by straight lines, so distances are approximate. | |
| start_time | No | When the trip starts. Defaults to now. Ignored when progress is given. An ISO 8601 time such as 2026-10-03T18:00:00-05:00 (read in the place's own time zone when it has no offset), or plain words: "now", "tonight", "tomorrow morning", "Saturday", "6 PM", "between 6 PM and midnight". A bare hour with no AM or PM is refused rather than guessed. | |
| coordinates | No | A list of [latitude, longitude] pairs, latitude FIRST. | |
| preferences | No | Optional: { "feel": "runs_cold" | "runs_warm", "avoid": ["shorts"] }. | |
| out_and_back | No | True if the route is walked out and then back along the same line. | |
| polyline_precision | No | Digits of precision in the polyline. 5 is usual, 6 is what some routers send. | |
| total_duration_hours | No | Instead of a pace: how long the whole trip is expected to take. The pace is worked out from the length of the route. | |
| break_minutes_per_hour | No | Minutes of stops per hour of moving, which stretches arrival times. | |
| rest_at_turnaround_hours | No | Hours spent resting at the turn of an out-and-back (a night's sleep, say). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the safety profile is covered; the description goes further with real behavior: progress position is never stored or logged, start_time is ignored when progress is given, a bare hour is refused rather than guessed, GPX input is capped near 450 KB, and units default by region. These are non-obvious traits an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the routing decision before any parameter talk, and the return summary and sibling pointer come last. The quoted 'Not "the weather on the trail" but...' sentence is rhetorically useful but slightly padded against an otherwise dense, well-ordered description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter, zero-required, nested-input tool with no output schema, the description covers the input modes, the two execution modes, timing semantics, and even the return shape (checkpoints with arrival time, place, weather and light, the changes, and a bottom line). Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema cannot: it enumerates the mutually exclusive route-input families (trail_id, GPX, GeoJSON, polyline, coordinates, place names) and states that trail_id is used OR one of the route formats. It also explains the pace-vs-total_duration_hours relationship and how break/rest inputs stretch arrival times.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — weather encountered along a route at each checkpoint's arrival time — and sharpens it with the contrast between 'the weather on the trail' and 'what will I encounter, where, and when'. It also names the sibling get_trail_weather and the condition that selects it, so an agent can route between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers two usage modes: a planned trip (start_time + pace) and a person already on the route (progress, answering for the remainder from now). It routes to get_trail_weather for an hour-by-hour timeline. It doesn't distinguish itself from other siblings like find_best_weather_window or outdoor_activity_weather, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchFind a place to get the weather forARead-onlyIdempotentInspect
Compatibility tool for ChatGPT connectors. Finds places matching a query and returns ids to pass to fetch. For real use prefer get_weather, outdoor_activity_weather and the other tools, which answer directly.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | A place name, such as "Greenbush, WI". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description adds genuinely useful pipeline context: it returns ids that must be passed to fetch, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the compatibility role and the correct alternative, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter lookup with no output schema, the description supplies the missing return-value behavior (ids for fetch) plus routing guidance. Nothing an agent needs to call or avoid it is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single query parameter is fully documented in the schema with an example. The description reframes it as a place-matching query but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (finds) and resource (places matching a query), and immediately distinguishes itself from fetch by noting it returns ids to pass along. The role as a compatibility shim is unambiguous, so an agent can tell it apart from the direct-answer weather siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says this is a compatibility tool and instructs the agent to prefer get_weather, outdoor_activity_weather and others for real use. Both the when-to-use (compatibility path) and when-not-to-use conditions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
- First observed
clothing_recommendation - First observed
explain_weather - First observed
fetch - First observed
find_best_weather_window - First observed
get_forecast - First observed
get_trail_weather - First observed
get_weather - First observed
list_trails - First observed
outdoor_activity_weather - First observed
route_weather - First observed
search
Related MCP Connectors
Marine and outdoor weather: multi-model forecasts, official bulletins, tides, anchorages.
Forecasts, climate history, severe alerts by location — for outdoor-event planners.
When to go and where to book: 87 places scored month by month from measured weather.
MCP server for weather with reasoning — umbrella advice, outdoor checks, city comparisons.
Related MCP Servers
- FlicenseAqualityDmaintenanceCombines UK Met Office weather forecasts with travel routing to provide outfit recommendations for walking, cycling, or driving journeys.5-
- AlicenseAqualityCmaintenanceProvides weather forecasts and climate history via Open-Meteo, including daily and hourly outlooks, typical past weather for future trip dates, location disambiguation, and ranked recommendations for pleasant outdoor days.5MIT
- AlicenseNot gradedqualityDmaintenanceProvides personalized recommendations for optimal outdoor exercise times by integrating weather data, Garmin Connect training schedules, and user performance metrics.2Apache 2.0
- AlicenseNot gradedqualityDmaintenanceProvides real-time weather information and multi-day forecasts for global locations using city names, coordinates, or ZIP codes. It includes tools for current conditions, forecasting, and weather summaries designed for activity planning.19 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.