Veabird: ski & beach trip planner
Server Details
Ski & beach trip planning: resort matching, snow and sargassum forecasts, lodging and flights.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- alexchouck-hash/bluebird
- GitHub Stars
- 0
TDQS
Scored across 15 tools
Most tools have clearly distinct purposes: search vs rank/recommend vs conditions vs booking. Some overlap exists because plan_trip bundles functions also available through find_flights, find_lodging, and the conditions tools, and beach_conditions includes sargassum data that sargassum_forecast covers in more depth. These overlaps are clarified by descriptions, so misselection is limited.
All names use snake_case and are readable, but the set mixes verb_noun names (find_flights, rank_beaches, search_ski_resorts) with noun_noun names (beach_conditions, ski_conditions, hurricane_threats, tides). The convention is not perfectly uniform, though the pattern is still predictable and domain-oriented.
15 tools is reasonable for a planner covering both ski and beach domains plus conditions, lodging, flights, advisories, and specialized forecasts. It sits at the upper end of the ideal range, and sargassum_model_skill is somewhat niche for a trip-planning server, but the count is not excessive.
The beach side is well-covered with search, ranking, conditions, tides, sargassum, hurricane, and advisory tools; the ski side covers search, ranking, and conditions but lacks equivalent depth for hazards such as avalanche forecasts or live lift/trail status. Flights, lodging, and plan_trip provide workable booking coverage, so gaps are minor rather than blocking.
Available Tools
15 toolsbeach_conditionsConditions at a beach resort on a dateARead-onlyIdempotentInspect
Full conditions for one beach resort on one date: 0-100 beach score with per-metric scores, sargassum, weather, water temperature, waves and rip currents, UV, crowds, hurricane risk, holidays. For today, also live conditions and alerts. Days 0-15 are forecasts, later dates climate outlooks.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | YYYY-MM-DD, today to ~365 days out (default today). | |
| resort | Yes | Resort id or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: date-dependent output character (live+alerts for today, forecasts for days 0-15, climate outlooks beyond), which an agent could not derive from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with scope before the metric list and temporal behavior, with no filler. The metric enumeration is long but each item conveys distinct returned content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so describing the returned payload (score with per-metric detail, sargassum, weather, waves, UV, crowds, hurricane risk, holidays) is exactly the missing information, and the temporal forecast/live distinction completes the picture for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only 2 params, so the schema fully documents 'date' format/range and 'resort' as id-or-name. The description reinforces the one-resort/one-date constraint but adds no syntax or format detail beyond the schema, matching the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: full conditions for one beach resort on one date, with an explicit enumeration of the returned metrics. It clearly reads as a superset tool versus narrow siblings like sargassum_forecast or tides, but it never names an alternative to make the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides temporal usage guidance ('For today, also live conditions and alerts. Days 0-15 are forecasts, later dates climate outlooks'), implying when a date is worth querying. However, it gives no when-to-use/when-not versus siblings such as rank_beaches or better_nearby_beaches, leaving alternative selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
better_nearby_beachesBetter nearby beachesBRead-onlyIdempotentInspect
For a beach resort and date (within 16 days), nearby resorts with a meaningfully higher beach score: swap suggestions when sargassum or weather is bad.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | YYYY-MM-DD | |
| resort | Yes | ||
| radiusMiles | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and openWorld, so safety is covered. The description adds useful behavior beyond that: the 16-day date ceiling and the 'meaningfully higher score' selection criterion. It still omits return format, pagination, and how radius defaults behave.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core scope front-loaded ('For a beach resort and date... nearby resorts with a meaningfully higher beach score'). It wastes nothing, though the colon-clause tail reads as slightly clipped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should carry return-value expectations; it gestures at the result (nearby higher-scoring resorts) but leaves 'meaningfully higher' undefined and says nothing about how many results or how radius affects them. Adequate but with clear gaps for a 3-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (just the date format), so the description must compensate. It does add the 16-day constraint on date, but resort and radiusMiles receive no semantics in the schema or the description, and radiusMiles bounds (25-1000) go entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete output (nearby resorts with a meaningfully higher beach score) for a given resort and date, so an agent can tell it apart from search_beach_resorts or rank_beaches. It does not explicitly name a sibling it is not, and 'swap suggestions' is slightly colloquial, but the resource and purpose are clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a triggering scenario ('when sargassum or weather is bad') and the 16-day horizon, which is real usage context. However, it names no alternative tools (e.g. sargassum_forecast, beach_conditions) and gives no when-not guidance, so the routing is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_flightsFlight searches to a resortARead-onlyIdempotentInspect
Pre-filled flight searches from the traveler's home airport to the best gateway airports for a ski or beach resort (or any airport code), with the drive from each airport to the resort.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Ski or beach resort name/id, or a destination airport code. | |
| from | Yes | Traveler's home airport code (e.g. DFW), city (e.g. 'Dallas') or 'lat,lon'. | |
| adults | No | Adults (default 2). | |
| depart | Yes | YYYY-MM-DD | |
| return | No | Return date (optional for one-way). | |
| childrenAges | No | Ages of children traveling, e.g. [6, 9]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, covering the safety profile. The description adds useful non-obvious behavior by stating that searches come pre-filled from the traveler's home airport and include the drive from each gateway airport to the resort. It does not, however, disclose what the response contains (fares, airlines, drive times).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that conveys the tool's scope and its distinguishing feature (gateway airports plus drive time) without filler. Slightly clause-heavy but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent search tool with full schema coverage and annotations carrying the safety profile, the description supplies the essential purpose and the notable pre-filled/drive-time behavior. Only return-value expectations are left unspecified, which is acceptable since no output schema exists but is not strictly required here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (including adults, depart, return, childrenAges) are already documented in the schema. The description only restates the from/to flexibility ('or any airport code') and adds no new syntax or format detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (searches flights) and resource (pre-filled flight searches to gateway airports for a resort), and clarifies the destination can be a resort or any airport code. It is clearly distinguishable from search_beach_resorts or find_lodging, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the content (flight search to a resort) and the input schema makes the trip-planning context obvious, but there is no explicit when-to-use guidance, no mention of prerequisites, and no routing against alternatives like find_lodging or plan_trip.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_lodgingFind places to stayARead-onlyIdempotentInspect
Places to stay near a ski or beach resort (or any 'lat,lon') for given dates and party: live nightly prices, ratings and photos where available, filtered by budget, plus hotel and vacation-rental searches.
| Name | Required | Description | Default |
|---|---|---|---|
| near | Yes | Ski or beach resort name/id, or 'lat,lon'. | |
| limit | No | ||
| rooms | No | Rooms needed (default 1). | |
| adults | No | Adults (default 2). | |
| checkin | Yes | YYYY-MM-DD | |
| checkout | Yes | YYYY-MM-DD | |
| radiusMiles | No | Search radius (default 5). | |
| childrenAges | No | Ages of children traveling, e.g. [6, 9]. | |
| maxPricePerNight | No | Lodging budget per night for the whole party, in USD. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/openWorld and non-destructive behavior, so the bar is lower. The description still adds real value beyond them: live (i.e., dynamically priced) results, coverage limits via 'where available', budget filtering, and that both hotel and vacation-rental inventories are queried.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the anchor and return contents front-loaded; nothing is wasted. The parenthetical placement of the 'lat,lon' alternative is slightly awkward but doesn't hurt comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-output-schema tool, the description usefully sketches the return payload (prices, ratings, photos) and the budget filter. It stops short of describing pagination, result ordering, or what happens when no lodging is found, which would be needed for a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 89%, so the parameters are already well documented in the schema. The description restates the key dimensions (dates, party, budget, location) without adding syntax or format detail, which earns the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (places to stay) and the anchor (ski/beach resort or 'lat,lon'), plus what it returns: live nightly prices, ratings, photos. That is enough to separate it from search_beach_resorts/search_ski_resorts, though it never names those siblings, so maximum clarity isn't reached.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'near a ski or beach resort' framing implies a context, but there is no explicit when-to-use or when-not-to-use guidance, and no routing against the several nearby-relevant siblings (search_beach_resorts, search_ski_resorts, plan_trip). An agent gets no help choosing between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_ski_resortsFind the best ski resorts for a travelerARead-onlyIdempotentInspect
Ranks ski resorts for one traveler or group: terrain that fits their ability (beginner/intermediate/advanced/expert, or a mixed group), snow and crowds for their dates, travel from their home airport or city (drive or fly), and their Epic/Ikon pass. Returns the reasons for each pick plus lodging, flight and rental links. Use for 'where should I ski' questions.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | Limit to a country, state or range, e.g. 'colorado', 'utah', 'alps', 'japan', 'US'. | |
| from | No | Traveler's home airport code (e.g. DFW), city (e.g. 'Dallas') or 'lat,lon'. | |
| pass | No | Season pass they hold. | |
| limit | No | How many resorts (default 5). | |
| budget | No | Value weights pass coverage and driving more heavily. Use find_lodging for nightly prices. | |
| depart | No | First ski day, YYYY-MM-DD (default today). | |
| nights | No | Trip length in nights (default 3). | |
| travel | No | How they want to get there (default either). | |
| ability | No | Ability of each skier or the group, e.g. ['intermediate'] or ['advanced','beginner']. | |
| childrenAges | No | Ages of children traveling, e.g. [6, 9]. | |
| maxDriveHours | No | Longest acceptable drive (default 5). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds value beyond that by disclosing the return content ('the reasons for each pick plus lodging, flight and rental links'), which tells the agent this is an explainable recommendation rather than a bare list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the ranking behavior and closing with the usage trigger. The first sentence is dense but every clause maps to a real parameter dimension; nothing is wasted, though it is near the upper size limit for a single sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 optional parameters and no output schema, the agent needs to know both the scope and the response shape; the description covers both ('ranks ... for one traveler or group' and 'Returns the reasons for each pick plus lodging, flight and rental links'). Minor gap: no mention of pagination/limit behavior or how defaults interact, though those are documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 11 parameters are already documented (including enum meanings and defaults). The description only restates the ability/dates/origin/pass axes at a high level, adding no syntax or weighting detail beyond the schema's own notes. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('ranks') plus resource ('ski resorts') and enumerates the ranking dimensions (ability, snow/crowds, travel origin, pass). An agent can distinguish it from search_ski_resorts (search/filter) and ski_conditions (conditions lookup) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear trigger language: "Use for 'where should I ski' questions." It establishes the recommendation context well, but never names an alternative or a when-not condition (e.g. use search_ski_resorts/ski_conditions when you already know the resort).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hurricane_threatsActive hurricanes and threatened resortsARead-onlyIdempotentInspect
Active Atlantic tropical storms and hurricanes (NOAA NHC) with position, intensity and track, and which Veabird beach resorts are threatened.
| Name | Required | Description | Default |
|---|---|---|---|
| resort | No | Optional resort to check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description adds meaningful context beyond them: the upstream source (NOAA NHC), the geographic scope (Atlantic only), and what the response contains (position, intensity, track, at-risk resorts), which helps an agent judge freshness and coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core resource and appends the resort-threat twist; no filler or repetition. It is slightly overloaded with fields in one clause but earns every word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only query tool with no output schema and full annotation coverage, the description is close to sufficient: it states the data source, scope, and returned fields. Only per-parameter semantics and refresh cadence are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional 'resort' parameter and the schema already documents it at 100% coverage, so the baseline is 3. The description's mention of 'threatened resorts' aligns with the parameter but adds no format or matching semantics (e.g., exact resort name vs. free text).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (active Atlantic tropical storms and hurricanes), a data source (NOAA NHC), and the payload (position, intensity, track, plus threatened Veabird resorts). That is well beyond a restated title, though it never explicitly distinguishes itself from domain-adjacent siblings like travel_advisory or beach_conditions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool for hurricane-threat queries, but there is no explicit when-to-use, no exclusions, and no pointer to an alternative such as travel_advisory for non-hurricane risk.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_tripPlan a ski or beach tripARead-onlyIdempotentInspect
Everything to book one trip to a ski or beach resort: conditions for the dates, places to stay (live nightly prices with photos where available, plus rental and hotel searches), flights from the traveler's airport to the best gateway, and extras (lift tickets, ski rentals, tours, car hire, insurance).
| Name | Required | Description | Default |
|---|---|---|---|
| from | No | Traveler's home airport code (e.g. DFW), city (e.g. 'Dallas') or 'lat,lon'. | |
| kind | No | Only if the name is ambiguous. | |
| adults | No | Adults (default 2). | |
| depart | Yes | Arrival date, YYYY-MM-DD. | |
| resort | Yes | Ski or beach resort name or id. | |
| return | No | Departure date (default 4 nights later). | |
| childrenAges | No | Ages of children traveling, e.g. [6, 9]. | |
| maxPricePerNight | No | Lodging budget per night for the whole party, in USD. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, and non-destructive behavior, so the safety profile is covered. The description adds useful behavioral context by enumerating what the tool returns: conditions, live nightly prices with photos, rental and hotel searches, flights, and extras like lift tickets and insurance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with a colon-separated list that is front-loaded with the core purpose: 'Everything to book one trip.' It is dense but every listed component earns its place by clarifying the tool's breadth.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an eight-parameter aggregator with no output schema, the description sufficiently communicates the breadth of returned information and the annotations cover safety. It could go further by explicitly relating itself to sibling tools, but the overall context is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all eight parameters thoroughly. The description reinforces some semantics, such as flights from the traveler's airport and lodging budget, but adds no new syntax or constraints beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: it plans a comprehensive ski or beach resort trip, enumerating conditions, lodging, flights, and extras. It functions as an umbrella aggregator, so an agent can infer its scope even without explicit sibling naming, though it does not call out alternatives by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: 'Everything to book one trip' signals a comprehensive trip-planning scenario. It does not explicitly list when not to use it or name sibling tools, but the context is strong enough for an agent to distinguish it from single-purpose searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rank_beachesBest beach resorts for a dateARead-onlyIdempotentInspect
Rank beach resorts by 0-100 beach score for a date, optionally within a region, island or country, or a radius around a resort. With from, also gives each resort's airports and a flight search. Use for 'where should I go to the beach' questions.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | YYYY-MM-DD (default today). Beyond 16 days needs a narrow query or near. | |
| from | No | Traveler's home airport code (e.g. DFW), city (e.g. 'Dallas') or 'lat,lon'. | |
| near | No | Resort id or name to rank around. | |
| limit | No | Default 8. | |
| query | No | Region, island, town or country, e.g. 'mexico', 'bahamas', 'punta cana'. | |
| radiusMiles | No | With near: max distance (default 300). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/open-world, so the safety profile is covered. The description adds real behavioral context beyond them: `from` additionally returns airports and a flight search, and dates beyond 16 days require a narrow query or `near`. Return shape (ranked list with scores) is only loosely implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core ranking semantics before the scope modifiers and the usage cue. No filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description sketches what comes back (ranked resorts with 0-100 scores, plus airports and a flight search when `from` is set), which is adequate for a read-only ranking tool. The default behavior when no filters are supplied could be stated more explicitly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds cross-parameter semantics the schema does not: the interaction of `from` with airport/flight output, and the query-vs-near vs date-window tradeoff. It still doesn't clarify how `query` and `near`/`radiusMiles` combine.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('rank beach resorts by 0-100 beach score') plus the filtering dimensions (region/island/country/radius). It implicitly distinguishes ranking from plain discovery, but never names the obvious sibling search_beach_resorts, so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the triggering question type ('where should I go to the beach'), which is clear usage context. It stops short of stating when NOT to use it or naming alternatives such as search_beach_resorts or better_nearby_beaches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sargassum_forecastSargassum forecast for a beach resortARead-onlyIdempotentInspect
Sargassum (seaweed) outlook for a beach resort: the next 16 days in detail plus a weekly outlook for coming months, barrier and cleanup programs, and the model's source and held-out skill.
| Name | Required | Description | Default |
|---|---|---|---|
| weeks | No | Weekly outlook length (default 12) | |
| resort | Yes | Resort id or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe read-only, idempotent, non-destructive profile, so the description's burden is lighter. It usefully adds what the caller actually receives — a 16-day daily detail plus a weekly monthly outlook, barrier/cleanup program information, and model source and held-out skill — which is meaningful behavioral context in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that is front-loaded with the core purpose (forecast horizon) before enumerating supplementary returns. It is efficient, though the trailing item list ('barrier and cleanup programs, the model's source and held-out skill') is somewhat tacked on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with no output schema, the description covers the return content well enough that an agent knows what it will get, and annotations cover the safety profile. The main gap is the missing tie-breaker against the closely related sargassum_model_skill sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both 'resort' and 'weeks' are already documented in the schema, making the baseline 3. The description's '16 days in detail plus a weekly outlook' implicitly clarifies the daily-vs-weekly granularity that the 'weeks' parameter controls, adding marginal value but no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (sargassum/seaweed outlook) and scope (next 16 days in detail plus a weekly outlook for coming months) plus supplementary content like barrier/cleanup programs and model skill. It is clear what the tool returns, but it does not distinguish itself from the sibling sargassum_model_skill, which the 'model's source and held-out skill' clause appears to overlap with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the description of output content; there is no explicit when-to-use statement, no exclusions, and no routing guidance versus alternatives such as beach_conditions, sargassum_model_skill, or plan_trip. The mention of model source/skill even creates ambiguity about which sibling to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sargassum_model_skillSargassum model accuracyARead-onlyIdempotentInspect
How well Veabird's sargassum model performs on held-out years (AUC, Brier score, calibration) versus baselines.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds that the evaluation is on held-out years and compared against baselines, which is useful content, but it says nothing about output shape or interpretation (e.g., what a good AUC looks like) beyond the metric names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the subject (the model) and lists the metrics compactly. No filler sentences, no redundancy with the title beyond what is necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only informational tool with no output schema, the description is sufficient: it names the metrics and the comparison baseline an agent will receive. It stops short of explaining how the metrics should be interpreted or the time span of held-out years.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no argument semantics to document; the baseline for a parameterless tool is 4. The description correctly implies no inputs are needed to obtain the report.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (Veabird's sargassum model performance) and enumerates the reported measures (AUC, Brier score, calibration) against baselines, so an agent understands it retrieves an evaluation report rather than a forecast. It does not explicitly distinguish itself from the sibling sargassum_forecast, which is the main ambiguity an agent could face.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: by describing held-out-year metrics, it signals this is the tool for inspecting model quality rather than getting a prediction. There is no explicit when-to-use, when-not-to-use, or pointer to sargassum_forecast as the alternative for actual forecasts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_beach_resortsSearch beach resortsARead-onlyIdempotentInspect
Search Veabird's 1,400+ beach resorts (Caribbean, Gulf of Mexico, Mexico's Pacific coast, Hawaii) by name, town, island, region or country. Returns ids, today's headline conditions, photos and page links.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | e.g. 'tulum', 'punta cana', 'barbados' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive). The description adds useful context beyond that: the dataset size and geographic coverage, and the shape of results (ids, today's headline conditions, photos, page links). It omits pagination/limit behavior, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste; the corpus scope is front-loaded and the return summary follows. Every clause carries information an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully summarizes return contents, and no parameters are required, so a bare query is valid. The only gap is the unexplained limit parameter and any result-cap behavior, which an agent would need when paginating.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: query has an example description while limit is undocumented. The description expands on query by listing the searchable facets (name, town, island, region, country), which adds meaning, but says nothing about limit or expected result counts, so it only marginally exceeds the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) plus resource (beach resorts), scopes the corpus (1,400+ resorts across named regions), and enumerates the searchable facets (name, town, island, region, country). It is immediately distinguishable from the parallel sibling search_ski_resorts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The name and facets imply that this tool is for locating resorts, but the description never states when to use it versus siblings like beach_conditions, rank_beaches, better_nearby_beaches, or find_lodging, nor does it name any exclusion or alternative. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_ski_resortsSearch ski resortsARead-onlyIdempotentInspect
Search Veabird's ~300 ski resorts by name, country, state or range, optionally only those on a pass. Returns ids, terrain mix, passes, gateway airports and forecast links.
| Name | Required | Description | Default |
|---|---|---|---|
| pass | No | ||
| limit | No | ||
| query | No | e.g. 'utah', 'alps', 'whistler', 'japan' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds value by disclosing the fields returned (ids, terrain mix, passes, gateway airports, forecast links), but says nothing about result caps or pagination behavior despite a limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with no filler; the core capability is front-loaded and the optional pass filter and return payload follow in a single clause each.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields and even mentions forecast links, which is the right level of detail for a search tool. The missing piece is any note on result limiting or default behavior, which matters for a tool with a limit parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate. It partially does by naming the facets query searches (name, country, state, range) and explaining the pass filter, but the limit parameter and its 1-100 bounds are never explained, and pass enum semantics (epic/ikon) are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search), resource (ski resorts), dataset size, and the searchable facets (name, country, state, range, pass). It does not, however, differentiate itself from the sibling find_ski_resorts, leaving the agent to guess which lookup is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'optionally only those on a pass' implies a filtering use case, and listing the searchable facets implies when the tool applies. But there is no explicit when-to-use guidance, no statement of when not to use it, and no routing to find_ski_resorts or ski_conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ski_conditionsSnow conditions at a ski resortARead-onlyIdempotentInspect
Today's 0-100 'ski today' score (fresh snow, base depth, weather, crowds), the next 7 days of summit snowfall and temperature, crowd outlook, terrain mix, passes and gateway airport for one ski resort.
| Name | Required | Description | Default |
|---|---|---|---|
| resort | Yes | Ski resort name or id, e.g. 'Breckenridge'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description goes beyond them by detailing the payload: a 0-100 score with named inputs (fresh snow, base depth, weather, crowds), a 7-day forecast window, crowd outlook, terrain mix, passes and gateway airport. It does not mention freshness/update cadence or how the score is sourced, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the headline output (today's score) before the supporting data. Every clause corresponds to a distinct returned data group; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so comprehensively — score, forecast horizon, crowd, terrain, passes, airport. Combined with annotations covering the safety profile and a fully documented single parameter, an agent has everything needed to invoke and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter and schema coverage is 100%, so the schema already documents 'resort' with an example. The description only reinforces the cardinality ('for one ski resort') and adds no format, aliasing, or disambiguation guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (one ski resort) and enumerates exactly what the tool produces: today's 0-100 score with its components, 7-day summit snowfall/temperature, crowds, terrain, passes and gateway airport. It is clearly a conditions-report tool rather than a search tool, though it never explicitly contrasts itself with siblings like find_ski_resorts or search_ski_resorts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied — you call it for a specific resort when you want current and near-term snow conditions — but there is no explicit when-to-use statement, no prerequisites, and no pointer to find_ski_resorts/search_ski_resorts for resort discovery. The agent must infer the routing itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tidesTide timesBRead-onlyIdempotentInspect
High and low tide times and heights at a beach resort for the next days (Open-Meteo sea level model, heights relative to mean sea level).
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | ||
| resort | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive, so safety is covered. The description adds genuinely useful non-annotation context: the Open-Meteo sea level model as source and that heights are relative to mean sea level, plus a forecast horizon. It does not cover rate limits, timezone of the times, or refresh cadence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the model source and datum are packed into a parenthetical. Efficient, though the parenthetical slightly buries operational info.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read-only forecast with no output schema, the description covers what the data is and its reference datum. The gap is the input side: neither parameter's expected format or limits are conveyed, leaving the agent to guess at the resort identifier.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does not. 'resort' is never explained (what form of name/ID?), and the 'days' parameter is only alluded to as 'the next days' with no mention of the 1-7 maximum enforced by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: high/low tide times and heights at a beach resort, with a forecast horizon. It is clearly distinguishable from siblings like beach_conditions or sargassum_forecast, though it does not name any sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, no pointer to alternatives. Usage is only implied by the tide/forecast framing; an agent gets no help deciding between this and beach_conditions for coastal info.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
travel_advisoryTravel advisory (crime and safety)ARead-onlyIdempotentInspect
US State Department travel advisory level (1-4) and summary for a country or a beach resort's country, plus Veabird's local safety notes.
| Name | Required | Description | Default |
|---|---|---|---|
| resort | No | ||
| country | No | ISO2 code, e.g. MX, JM, DO |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, open-world, and non-destructive, so the safety profile is covered. The description adds valuable content beyond that: the advisory is a 1-4 level with a summary plus Veabird's local safety notes, telling the agent what the payload contains. It stops short of disclosing freshness or caching behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler, front-loading the primary return value (advisory level and summary) before the supplementary local notes. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by naming the level, summary, and safety notes. For a two-parameter read tool this is largely complete; the only small gap is not stating what happens if both or neither optional parameter is supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only 'country' is documented as an ISO2 code). The description compensates by clarifying the 'resort' parameter's role: it resolves to the resort's country, so the agent understands the two parameters are alternative input paths to the same underlying lookup. This adds meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (US State Department travel advisory) and scope (crime and safety, level 1-4 plus summary and local notes), so the agent knows exactly what it returns. It does not explicitly contrast with adjacent siblings like hurricane_threats or sargassum_forecast, but the 'crime and safety' framing makes the domain distinction clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent can infer this is for looking up a country's or resort's safety advisory, and the phrase 'a beach resort's country' hints at the resort path. There is no explicit when-to-use, when-not-to-use, or alternative named, so guidance stays at the implied level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
find_ski_resorts1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"How many resorts (default 6)."New value: +"How many resorts (default 5)."
15 tool updates
- First observed
beach_conditions - First observed
better_nearby_beaches - First observed
find_flights - First observed
find_lodging - First observed
find_ski_resorts - First observed
hurricane_threats - First observed
plan_trip - First observed
rank_beaches - First observed
sargassum_forecast - First observed
sargassum_model_skill - First observed
search_beach_resorts - First observed
search_ski_resorts - First observed
ski_conditions - First observed
tides - First observed
travel_advisory
Related MCP Connectors
Live verified resort snow, forecasts, powder search, trip planning & grounded Q&A for 430+ resorts.
Live ski snow, multi-model forecasts, powder rankings & a grounded Answer Engine for 500+ resorts.
Ski resort conditions, forecasts, avalanche danger, sun, routes and snow history for 5,000+ resorts.
Snow forecasts, lift status, season history, costs and AI ski trip planning for Europe
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables users to find the best ski resort snow conditions worldwide and search for flights to get there.-
- AlicenseAqualityBmaintenanceLive ski & snow data for AI agents: 14-day multi-model forecasts, powder rankings, resort guides, webcams, ski-pass intelligence, and avalanche/road safety across 500+ resorts. Hosted streamable-HTTP — no install, no auth.40MIT
- FlicenseNot gradedqualityDmaintenanceProvides personalized travel recommendations by analyzing climate, currency exchange rates, safety, and budget constraints through AI agents and multiple public APIs.-
- AlicenseNot gradedqualityBmaintenanceEnables users to translate natural-language vacation or business travel goals into multi-leg flight segments, lodging selections, layover buffers, baggage-fee optimization, and consolidated reservation manifests.7MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.