Hamptons Verified
Server Details
Verified Hamptons & North Fork data: open now, events, beaches, permits, closures.
- Status
- Healthy
- Uptime
- 100.0% over 43 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 19 tools
Each tool is scoped to a distinct domain (beaches, galas, lodging, sport, transit, water) and the descriptions go to unusual lengths to cross-reference and hand off to the right sibling tool. The residual overlaps — search_places vs whats_open_now, things_to_do vs play_sport vs work_out, and the recently_closed vs search_places boundary — are explicitly addressed in the text but still leave a few judgment calls for an agent.
All names are consistent snake_case with no camelCase or style clashes. There is a mild split between verb-led (search_places, plan_my_day, play_sport) and noun-phrase (beach_info, golden_hour, picnic_spots) forms, but every name reads predictably and maps to one obvious domain.
19 tools is on the heavy side for a local-guide server, but each covers a genuinely separate slice of the East End surface (beaches, galas, shops, transit, lodging, safety, etc.) with no obvious redundant entries. It is slightly over the ideal band yet every tool clearly earns its place.
The surface covers the full breadth of a regional guide: dining/venues, shopping, lodging, activities, sports, workouts, events, galas, transit, parking, water safety, emergency care, and community services. Deliberate omissions (shop hours, weather, live incident feeds) are documented in-line rather than hidden, so agents hit no silent dead ends.
Available Tools
19 toolsbeach_infoBeach: dog rules today, parking permit, lifeguardsAInspect
One named beach (or every verified beach in a town or hamlet — 'Amagansett, NY' and 'Sag Harbour' resolve, and readAs reports which place was answered for) with the answers people actually arrive with: whether dogs are allowed TODAY, what parking permit the lot needs, whether you can get there without a car, lifeguard hours and rip-current notes. 49 beaches on both forks, each with its governing town/village/state source and the date we last read it. Call this for any dog question — the rule is per-month and splits morning from daytime, so most South Fork ocean beaches are off-leash at dawn and closed to dogs in daylight from May to September, and no model can know today's answer. Also call it before saying anyone can just park at a beach. For the full permit map across all 7 jurisdictions and non-resident day passes, use parking_permit_rules instead. Returns up to 5.
| Name | Required | Description | Default |
|---|---|---|---|
| beach | Yes | Beach name, town, hamlet or village, e.g. 'Georgica', 'Main Beach' or 'Amagansett' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that town/hamlet queries return multiple beaches, that readAs reports the resolved place, that results include source and date, and that output is capped at 5. It doesn't cover every edge case, but the main behavioral traits are clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is fairly long but dense with necessary detail. It is front-loaded with the main purpose, then adds resolution behavior, usage guidance, and a sibling alternative. Every sentence contributes, though some tightening could make it even more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple single-parameter schema and lack of annotations or output schema, the description is remarkably complete: it covers what results are returned, the dynamic dog-rule aspect, the parking permit scope, source/date metadata, output limit, and when to use an alternative. The agent has enough to invoke the tool correctly in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the parameter name and example inputs, but the description adds meaning beyond that by explaining how town/hamlet names resolve, that readAs reports the matched place, and that multiple results (up to 5) are possible. This is useful context the schema alone does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool returns beach-specific information (dog rules today, parking permit, car-free access, lifeguards, rip currents) for one named beach or all verified beaches in a town. It explicitly distinguishes itself from sibling tools by naming parking_permit_rules as the alternative for permit maps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives direct invocation guidance: 'Call this for any dog question' and 'Also call it before saying anyone can just park at a beach.' It also states when to use parking_permit_rules instead, providing a clear exclusion and alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
benefit_galasBenefit galas: the dates, the dress code, and who the money goes toAInspect
The East End's summer benefit season: 15 dated charity galas, each joined to the nonprofit it raises for. Call this for any gala, benefit, fundraiser, charity, black-tie, what-should-I-wear, where-can-I-donate or where-can-I-volunteer question. Three things here are unknowable from your weights. First the DATES: this is the 2026 season, and a benefit invented for the right-sounding weekend sends someone to an empty field. Second dressCode, which is the organiser's own and is the thing people actually ask — it is 'Cocktail' on 16 of the 15 and 'Black Tie' on exactly ONE, so the black-tie assumption you would reach for is wrong twenty-six times out of twenty-seven; 'Beach Formal', 'Farm Festive' and 'Garden Chic' are real answers you cannot guess. Third, the PRICE, which cuts both ways and is why you read tickets.publishedRate per record instead of generalising. On 14 of the 15 it is false, and that is a verified finding rather than a missing field — the organisers hand off to a ticketing platform and print no rate, and a sweep of all 15 ticket links found 11 dead and several resolving to a prior year. For those, NEVER state or estimate a price; say it is not published and give the phone. On 1 it is true, and tickets.published carries the organiser's own figure with the URL an editor read it from and the date — quote and cite that, keeping the word 'from', because it is the entry rate in a ladder and not the only price on sale. tickets.tier is a relative band across these events, not a figure, and nothing may be inferred from it either way. Where tickets.publishedTiers appears an editor read a real published line: quote it verbatim, including the ones whose content is that no price is published and here is the number — that IS the answer to what it costs. beneficiary carries the charity's mission, phone, donateUrl and volunteerUrl, all verified more recently than the ticket link, so give the phone as the route that works and answer 'I can't make it, can I still help' without leaving the data. alsoThisSeason holds the marquee fixtures with no gala record — the Hampton Classic, Authors Night, the art fair — and is where several more genuinely published prices in this vertical live. Past benefits never appear unless you pass from. Returns up to 10.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ISO date. Omit for the rest of the season. | |
| from | No | ISO date. Default is today (East End local), so only benefits still to come are returned. Pass an earlier date to reach ones already held. | |
| town | No | A hamlet ('East Hampton', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| cause | No | Filter to one cause, when the question is plainly about one. | |
| query | No | What they are after, e.g. 'animal', 'ARF', 'black tie', 'film', 'land trust', or a benefit's name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it is remarkably forthcoming about quirks: dead ticket links, unpublished rates, the need to quote verbatim from `publishedTiers`, and the 'from' parameter gating past events. However, a few numerical inconsistencies (e.g., 'Cocktail' on 16 of 15 events, 'wrong twenty-six times out of twenty-seven') introduce confusion in an otherwise transparent account.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is densely packed with essential caveats, but it is also overlong and at times convoluted (e.g., 'unknowable from your weights', 'cuts both ways', 'the black-tie assumption you would reach for is wrong twenty-six times out of twenty-seven'). While front-loaded with purpose and usage, several sentences could be tightened without losing meaning, and the internal contradictions undermine clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complex data model (ticket pricing nuances, dress codes, beneficiary info, seasonal fixtures) and the absence of an output schema, the description is remarkably complete. It tells the agent what fields exist, what to quote, what to avoid estimating, how to handle donation/volunteer questions, and how to interpret edge cases like dead links and unpublished rates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already covers all five parameters, the description adds substantial semantic value beyond the field names: it explains the meaning of `tickets.publishedRate`, `tickets.tier`, `tickets.publishedTiers`, `beneficiary`, and `alsoThisSeason`, and it operationalizes the `from` parameter ('Past benefits never appear unless you pass `from`'). This goes far beyond the schema's minimal descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's domain: '15 dated charity galas, each joined to the nonprofit it raises for.' It goes beyond the title by explicitly enumerating the question types it answers ('gala, benefit, fundraiser, charity, black-tie, what-should-I-wear, where-can-I-donate or where-can-I-volunteer'), making it easy to distinguish from sibling tools like `upcoming_events` or `community_help`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit, actionable guidance on exactly when to invoke the tool: 'Call this for any gala, benefit, fundraiser, charity, black-tie, what-should-I-wear, where-can-I-donate or where-can-I-volunteer question.' It also provides strong negative guidance, e.g., 'NEVER state or estimate a price' when `tickets.publishedRate` is false, and explains the `from` parameter's role for accessing past benefits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
community_helpCommunity help: free and low-cost services, in Spanish and EnglishAInspect
The guide's bilingual community directory: free and low-cost services for people who LIVE and WORK on the East End. Immigration legal advocacy, sliding-scale and no-status-asked clinics, food pantries, the winter shelter, tenant-rights and wage-theft hotlines, ESL / GED / citizenship classes, childcare and youth programs, money transfer, Spanish-language Mass times, the Latino markets, and the Sunday soccer. Call this for any I-need-help question — unpaid wages, eviction, a doctor without insurance, English classes, food, a lawyer for an immigration case — and for anything asked in Spanish. Answering from memory here means naming a national hotline or a legal aid society that does not serve this county, and the asker loses a day's work finding that out. Every record carries phone, source, sourceUrl and lastVerified, and the guide's own text in BOTH languages (description.es / description.en): quote the Spanish as written rather than translating the English. Two things are computed rather than stored: nextSpanishMass / nextEnglishMass, the next service from right now on the East End's clock (America/New_York) — 'today at 7:00 PM' means today here — and inSeasonNow on the soccer leagues for this month. offEastEnd: true marks a real referral to Bay Shore, Hempstead or Central Islip: still the right answer sometimes, but say the drive out loud. NEVER volunteer that a service is free, confidential, or safe regarding immigration status unless the record's own words say so — quote those verbatim where they appear and attribute them, and otherwise give the phone number. Never name a provider that is not in the results. Pass query in the asker's own words, in either language; need is a shortcut vocabulary; town takes the asker's own place words. With no argument nothing is looked up — you get the needs and towns covered, so ask. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| need | No | Shortcut category, when the question is plainly one of these. | |
| town | No | A hamlet ('East Hampton', 'Riverhead') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| query | No | What they need, in their words and either language, e.g. 'my boss hasn't paid me', 'clases de inglés', 'food pantry', 'misa en español', 'dentist', 'help with rent'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It discloses the risk of answering from memory ('the asker loses a day's work finding that out'), explains computed fields with timezone handling, describes the offEastEnd flag and its required communication, and sets strict rules about not naming unlisted providers. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but justified given the tool's complexity. Every sentence contributes, from the list of services to the rules about quoting Spanish and not naming out-of-results providers. It could be slightly more structured with bullet points, but it remains efficient and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully covers the tool's behavior: returns up to 8 records each containing phone, source, sourceUrl, lastVerified, and bilingual text; explains computed fields (next masses, inSeasonNow) with timezone; and handles edge cases like offEastEnd and no-argument calls. Since there is no output schema, this level of detail is necessary and sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful guidance beyond the schema: passing the asker's own words for query and town, 'need' as a shortcut vocabulary, and the no-argument behavior where nothing is looked up but needs/towns are returned. This enriches understanding without being exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is a bilingual community directory for free and low-cost services on the East End, and explicitly says 'Call this for any I-need-help question' with a list of examples. It distinguishes itself from sibling tools by covering a broad range of social services rather than specific topics like beaches or parking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use guidance: 'Call this for any I-need-help question' and lists concrete scenarios such as unpaid wages, eviction, medical care, and English classes. It also gives contextual rules like quoting Spanish verbatim and never volunteering safety claims unless the record states them, which clarifies appropriate usage versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_shopsShops: the verified East End stores, and where each one isAInspect
Verified East End SHOPS. 74 stores across 9 hamlets on the South Fork, North Fork and Shelter Island: fashion, jewelry, swimwear, activewear, kids, toys, home, books, beauty, consignment, accessories, general-store, florist. Call this for any shop, shopping, store, boutique, browse, bookshop, gift or where-can-I-buy question, and call it for the shopping half of a day plan. Every address here was read off the brand's OWN store page at sourceUrl — never a homepage, never a third party — and coordSource records what corroborated the coordinate, because three brands publish a schema.org geo on their own store page that is simply wrong (Athleta's Bridgehampton store publishes a point 3.6 km away, in East Hampton village). The one thing this tool does NOT have is hours: the guide holds no opening schedule and no published hours line for any shop, in any form. So never say a store is open, closed, or open until a time, and never borrow an hour from a neighbouring venue — give phone where there is one and the address otherwise. That is also why shops are not in search_places, which promises an open/closed verdict on every row it returns. query reads the name, the brand, the category and the one-line what the guide keeps, so 'bookshop', 'toys', 'flowers' and 'The Wharf Shop' all land. Pass town in the asker's own words. With no argument nothing is looked up — you get the categories and towns covered, so ask which. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Sag Harbor', 'East Hampton') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| query | No | What they want, e.g. 'bookshop', 'toys', 'flowers', 'swimwear', 'gift', or a store's name ('The Wharf Shop'). | |
| category | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: provenance of addresses (brand's own store page, never a homepage or third party), coordSource corroboration, three brands publishing wrong schema.org geo, a hard 'no hours' constraint with instructions never to claim open/closed, a fallback to `phone`, and a result cap of 8. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb, resource, and coverage counts before any caveats. The provenance and hours passages are long but each prevents a concrete failure (hallucinated addresses, fabricated opening hours), so they earn their place; the text could still be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must cover returns and limits — it states 'Returns up to 8' and the no-argument fallback. For a 3-parameter lookup tool it leaves no material gap an agent would need before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description adds real meaning: `query` 'reads the name, the brand, the category and the one-line `what` the guide keeps', confirming fuzzy multi-field matching beyond the schema's example list. It also documents the no-argument behavior (returns categories and towns), which is parameter-adjacent semantics the schema cannot express.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (verified East End shops) with concrete scope — 74 stores across 9 hamlets on the South Fork, North Fork and Shelter Island — and enumerates the covered categories. It also explicitly distinguishes itself from the sibling `search_places`, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger vocabulary ('any shop, shopping, store, boutique, browse, bookshop, gift or where-can-I-buy question') and a second use case ('the shopping half of a day plan'). It also names the condition that excludes it from a sibling: shops lack hours, which is why `search_places` is used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
getting_hereGetting here: fares, ferries, and whether the road is movingAInspect
How to get to the East End and around it: Hampton Jitney, the LIRR Cannonball and off-peak Montauk Branch, Blade, seaplanes, the three ferries, cycling, the ride-hail reality, and driving — each with the operator's published fare, season and stop list, plus when the guide last read it. Call this for any getting-there, how-long, how-much or when-should-I-leave question, and for 'without a car'. Two things here are unknowable from your weights and worth the call on their own: current fares (a prepaid Jitney seat is a different number from one paid on board, and Cross Sound adds a fuel surcharge to every published fare), and roadConditionsNow — whether this instant falls inside one of the published NY-27 crawl windows, computed against the East End's own clock (America/New_York), with the next window's start time so you can tell someone when to leave instead of guessing, and all three published windows with their days and hours so a departure time that is not now still gets a real answer. Pass to with the asker's own destination words: the answer then reports which Jitney stops actually serve it and the drive-time row for that run, with the guide's typical AND summer minutes rather than one averaged number. Pass mode: 'car-free' for someone without a car. Every figure carries its source; where sourceCaveat appears, the record is filed against a link that does not publish it — attribute it to the guide and do not cite that URL. There is no live incident feed behind this: it is the published pattern, not what the road is doing this minute.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Where they are going — a hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. Omit for the whole picture. | |
| mode | No | 'car-free' when the asker has no car (drops driving); 'driving' for road answers only. Omit for every option. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: it uses published fares with source timestamps, computes `roadConditionsNow` against America/New_York, explains source caveats for links that don't publish data, and clearly states the tool provides published patterns not live incidents. This is thorough and honest about limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph of ~300 words. It is front-loaded with an overview and contains valuable details, but it lacks clear structure or brevity. Some sentences could be trimmed (e.g., the repeated emphasis on 'without a car'). It earns its place content-wise but is not concise or easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description gives a comprehensive picture: what modes are covered, how fare data is sourced, how `roadConditionsNow` works, how `to` affects output, and explicit caveats about source attribution and lack of live data. For a tool with this complexity, the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for both parameters, and the description adds meaningful details beyond them. For `to`, it explains the output effect (reports Jitney stops and drive-time rows with typical AND summer minutes). For `mode`, it mostly restates the schema but reinforces 'car-free' as the right choice for without a car. This is helpful but not a huge leap beyond the structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: how to get to the East End and around it, covering specific transit modes (Hampton Jitney, LIRR, ferries, etc.) and driving. It explicitly says to call this for any getting-there, how-long, how-much, or when-should-I-leave question, distinguishing it from sibling tools about beaches, parking, or events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: call for any travel question, pass `to` with the asker's own destination words, pass `mode: 'car-free'` for someone without a car. It also clarifies when NOT to use the tool ('There is no live incident feed behind this'), preventing misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
golden_hourSunset and sunrise: where to stand, and what time to be thereAInspect
Verified East End sunrise and sunset viewpoints — public overlooks, beaches, vineyard lawns, hotel terraces and the bars people actually book for it — each with the TIME. Call this for any sunset, sunrise, golden hour, blue hour, best-views or where-should-we-watch question, and call it even when you think you know the spot, because the half that decides the evening is the clock. time is computed from that spot's own latitude and longitude for that date on the East End's clock (America/New_York): sunset here swings over three hours across the year and Orient and Montauk Point differ by about two minutes, so a remembered time is wrong twice over. You also get goldenHour (the hour the light is worth the drive — INTO sunset, OUT OF sunrise), blueHour, and arriveBy, which is 30 minutes earlier because that is when the lot fills; parking and walkDistance say why. startsIn counts down when it has not happened yet today and alreadyPassedToday says when it has — do not offer tonight's number as though it were still coming. Pass when: 'sunrise' for the morning side; the default is sunset. Pass date (YYYY-MM-DD) to plan ahead, and read inBestMonth, because several of these only work in the months the record names. The field no model has is wrongWayRound: verified spots in the same towns that face the OTHER way — Ditch Plains reads like a sunset beach and is east-facing, and sending someone there is the mistake this replaces. There is no weather here, so never promise a clear sky. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ISO date (YYYY-MM-DD) to compute for. Omit for today, East End local. | |
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| when | No | 'sunset' (default) or 'sunrise'. | |
| query | No | What they want, e.g. 'dinner', 'drinks', 'no walk', 'wheelchair', 'dog', 'kids', 'vineyard', or a spot's name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses computation logic ('time is computed from that spot's own latitude and longitude'), derived fields ('goldenHour', 'blueHour', 'arriveBy'), countdown semantics ('startsIn counts down... alreadyPassedToday'), a unique warning field ('wrongWayRound'), a hard cap ('Returns up to 8'), and an exclusion (weather). This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with necessary nuance; every sentence introduces behavior or guidance that the schema and annotations do not provide. It is not tautological or padded. Minor redundancy exists (multiple mentions of the clock/timing), but the structure is front-loaded and each clause serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no output schema, and no annotations, the description covers the essential behavioral context: return fields, field meanings, timing semantics, edge cases ('alreadyPassedToday'), and a warning against a common mistake ('wrongWayRound'). It even cautions about weather and seasonal availability. This is complete enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real value beyond the schema by explaining the implications of `when` ('Pass when: 'sunrise' for the morning side; the default is sunset'), `date` ('to plan ahead... read inBestMonth'), and `query` via examples. It stops short of enumerating every parameter, but its guidance enriches the schema sufficiently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific scope: 'Verified East End sunrise and sunset viewpoints' with named venue types. It clearly identifies the tool's job via 'Call this for any sunset, sunrise, golden hour, blue hour, best-views or where-should-we-watch question', differentiating it from sibling tools like beach_info or whats_open_now.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit invocation triggers are given: 'Call this for any sunset, sunrise, golden hour, blue hour...' and even 'call it even when you think you know the spot'. It also states an exclusion: 'There is no weather here, so never promise a clear sky,' giving clear when-not-to-use guidance. No alternative tools are named, but the context is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parking_permit_rulesCan I park here? Permits, day passes, and whether it's enforced nowAInspect
The East End beach-parking permit map — all 7 jurisdictions — resolved against the beach or town asked about. Pass where with a beach name ('Main Beach', 'Ditch Plains') and you get the verdict for that lot: which permits grant it, which stickers explicitly DO NOT (an East Hampton Town sticker does not open Main Beach — a mile and a jurisdiction apart, and the commonest wrong answer), the non-resident and day-pass options with prices, the fine, and where to apply. Crucially each carries enforcement.enforcedRightNow, computed from the published season and daily window against the East End's own clock: off-season or after 6pm nobody needs a permit, and you cannot know that from your weights. Pass a town or hamlet instead to get that jurisdiction plus its beaches' permit lines; omit where for the whole map. 16 beaches are state, county or DEC land where no town permit applies — they come back flagged outsideTheStickerMap, which is how you say 'yes, you can just park'. For dogs, lifeguards and getting there without a car, use beach_info.
| Name | Required | Description | Default |
|---|---|---|---|
| where | No | Beach name, town, hamlet or village, e.g. 'Main Beach', 'Coopers', 'Montauk'. Omit for the whole map. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly. It discloses time-aware enforcement (enforcement.enforcedRightNow) computed from season/daily window, flags 16 state/county/DEC beaches with outsideTheStickerMap, and explains the common wrong sticker assumption. This goes well beyond a typical description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with valuable information. Every sentence adds context, from the permit-line resolution to enforcement timing and the outside-the-sticker-map flag. The illustrative example about East Hampton sticker is useful, not fluff. It could be tightened slightly, but it is well-structured and front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and no annotations, the description completely covers the tool's behavior: input variations, output fields, time sensitivity, edge cases (16 beaches outside the map), and the alternative tool. It explains the key flags and result types, making it fully self-sufficient for an agent to decide when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already covers 100% of the single parameter, but the description adds meaningful semantics: it explains that passing a beach returns a lot verdict, passing a town/hamlet returns jurisdiction plus beaches' lines, and omitting returns the whole map. It enriches examples and clarifies output variations beyond the schema's basic type description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it resolves parking permit rules for East End beaches across all 7 jurisdictions, with a specific verb and resource. It distinguishes from the sibling beach_info by explicitly redirecting dog, lifeguard, and transit questions to that tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use beach_info for dogs, lifeguards, and getting there without a car. It also explains the different input modes (beach name vs town/hamlet vs omitting 'where') and what each returns, giving clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
picnic_spotsPicnics: where alcohol, a grill and the dog are actually legalAInspect
Verified East End picnic grounds — public parks, state and county parks, preserves, beach-grass strips, winery lawns and historic gardens — with the rule that decides whether the plan is legal. Call this for any picnic, blanket, park, park-hours, barbecue or fire-pit question, and ALWAYS before saying anyone can drink outdoors: the alcohol rule is set by whoever owns the grass (New York State, Suffolk County, a town, a village, a winery) and splits roughly a third prohibited, a third bring-your-own, a third wine-only-or-permit-only. Answering that from memory is the classic confident wrong answer, and it costs the asker a village summons rather than a bad meal. alcohol is a sentence to quote, not a boolean: wine-only is a winery lawn where the estate's wine is fine and your bottle is not, and allowed-with-permit means not allowed until the permit in permitUrl is in hand. glassBottles: prohibited holds even where alcohol is allowed. fireOrGrill and dogs are per-spot. 15 of these publish hours as "Sunrise to sunset" or "Dawn to dusk", so closesAt is resolved from THAT spot's own coordinates for today — "8:09 PM (sunset)" is tonight's sunset there, which is not something you can know. Where only hoursText comes back the published line was not machine-readable: quote it, do not turn it into a claim about right now. Pass allows to filter to what the picnic needs. The guide holds NO drinking rule for the ocean beaches — those are beach_info's, which carries dogs and permits but not alcohol — so never infer one from these. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| query | No | What they want, e.g. 'sunset', 'shade', 'oceanfront', 'quiet', 'big group', or a park's name. | |
| allows | No | Filter to spots that permit this. 'alcohol' includes wine-only and permit-only spots, each flagged in its own `alcohol` line. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and exceeds it. It discloses nuanced behaviors: the alcohol field is a sentence not a boolean, closesAt is computed from each spot's coordinates for today's sunset, hoursText should be quoted when not machine-readable, and results are capped at 8. It also alerts the agent to the legal stakes of guessing, which is critical context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense paragraph, but every sentence carries substantive information—purpose, usage, alcohol rule details, time resolution logic, and exclusions. It is front-loaded with purpose and usage, though the lack of formatting (bullets or sections) makes it less scannable. Slightly verbose for a tool definition, but no word is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description provides a thorough mental model of the response: it mentions fields like `alcohol`, `glassBottles`, `fireOrGrill`, `dogs`, `closesAt`, `hoursText`, and `permitUrl`, and explains how to interpret them. It also covers edge cases (sunset-based closing times, non-machine-readable hours) and explicitly distinguishes from beach_info, making it complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the `allows` filter behavior: 'Pass `allows` to filter to what the picnic needs' and clarifies that 'alcohol' includes wine-only and permit-only spots, each flagged in its own alcohol line. This semantic nuance improves the agent's ability to use the filter correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly defines the tool as 'Verified East End picnic grounds' with a clear scope (public parks, state and county parks, winery lawns, etc.) and a specific purpose: determining the legality of a picnic plan (alcohol, grill, dogs). It distinguishes itself from sibling tools by stating the ocean beaches are covered by 'beach_info's', not this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Call this for any picnic, blanket, park, park-hours, barbecue or fire-pit question, and ALWAYS before saying anyone can drink outdoors.' It also provides a when-not-to-use exclusion: 'The guide holds NO drinking rule for the ocean beaches — those are beach_info's' and warns against relying on memory, making the usage conditions unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_my_dayPlan a rules-checked East End dayAInspect
Build a complete East End day using Hamptons Verified's deterministic planner. Use this whenever someone asks ChatGPT or Claude to plan, schedule, sequence, or personalize a day—not search_places, which returns an unordered list. The tool produces three distinct options and applies the same constraints as the website before returning anything: permanent closures, the selected date and opening hour, age fit, budget, weather, daylight, geographic sequencing and the maximum drive between consecutive stops. saved_places is a preference lane: pass names the asker supplies and eligible ones are worked into the day, but they never override a safety or feasibility rule; the rest of the stops remain discoveries. This connector cannot read the asker's private Hamptons Verified account, browser or saved list, so ask for the names when they want them included. Present the returned plan; never replace a stop from your own memory.
| Name | Required | Description | Default |
|---|---|---|---|
| who | No | Who the day is for: solo, couple, family, friends, first-date, in-laws. There is no 'group' value — a group of friends is `friends` and a group with children is `family`; pass either and the answer reports which was used in `whoReadAs`. For the size of the group use `party_size`, which is a different question. | |
| date | No | ISO date in East End local time; defaults to today. | |
| seed | No | Optional stable seed when the asker wants the same result repeated. | |
| town | No | A hamlet or region to keep the day around. Omit for the full East End. | |
| vibes | No | Comma-separated interests. The planner understands exactly these: quiet, active, foodie, cultural, beach, adventure, romantic, off-season, shopping, anything. Anything else is reported back in `vibesNotUnderstood` rather than acted on — it is never silently dropped. | |
| budget | No | ||
| energy | No | ||
| party_size | No | How many people are coming, as a number. Stops whose operator publishes a maximum BELOW it come back with a `capacityWarning` quoting that operator's own line — several Sag Harbor charters cap at six. Omit when the asker did not say; nothing is guessed. | |
| saved_places | No | Comma-separated saved venue names supplied by the asker. Never infer private account data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly: it discloses determinism, that three options are produced, the full constraint set applied before returning (closures, date/opening hour, age, budget, weather, daylight, geographic sequencing, max drive), that saved_places is a preference lane that never overrides safety rules, and that the connector cannot read private account data. This is rich behavioral context beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (~200 words) but dense — every sentence carries operational meaning and there is no fluff. The usage-vs-alternative guidance is front-loaded in the first sentence, and later sentences each add a distinct behavioral fact. For a 9-parameter planner, the length is justified; a minor trim of the saved_places sentence could tighten it, but it is not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must carry the full contract, and it does: it pre-announces return shape (three distinct options), reports fields (whoReadAs, vibesNotUnderstood, capacityWarning), covers privacy limitations, and states the deterministic constraint set. For a complex 9-param tool with zero structural safety nets, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 78%, so the schema already documents most parameters; the description adds genuine value on top for key ones: saved_places is explained as a preference lane that never overrides rules, party_size gains the capacityWarning behavior with a concrete Sag Harbor example, and vibes gains the vibesNotUnderstood non-silent-drop behavior. A few params (town, energy, budget) rely on the schema alone, but the description compensates where behavior matters most.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (build/plan), a concrete resource (a complete East End day), and immediately distinguishes itself from the sibling search_places ('which returns an unordered list'). The title reinforces scope. An agent can tell exactly what this tool produces and how it differs from nearby tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger condition ('whenever someone asks... to plan, schedule, sequence, or personalize a day') and names the alternative it is not (search_places). It also gives operational guidance for saved_places — ask the user for names when they want them included — and instructs the agent to present rather than substitute stops. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_sportWhere you can actually play — courts, fields, greens and tonight's pickup runAInspect
Where to PLAY a sport on the East End: public and club courts, school fields, golf, beach volleyball, surf breaks, mountain-bike and trail-running routes, the skate park, the disc golf course, sailing, riding stables — and the recurring basketball pickup runs. Call this for any where-can-I-play, is-there-a-court, pickup-game, tee-time or is-there-a-game-tonight question, and call it before naming anywhere from memory, because the field that decides the afternoon is not the name. access says who gets on: 12 of these are members-only and they include exactly the names a recommendation reaches for (Maidstone Club; Shinnecock Hills Golf Club; Topping Riding Club; The Meadow Club of Southampton — Paddle Tennis; Breakwater Yacht Club — Sailing), so the fluent answer to "where do I play golf in Southampton" is a course nobody can walk onto, while Montauk Downs, a state park anyone can book, is in the same list. access: permit means the play is free and the PARKING is not (Ditch Plains, Indian Wells, Main Beach) — hand that half to parking_permit_rules. access: not-published is a real third answer: the record shows open play and the guide holds no rate, so do not round it up to free. Never derive a price from access — it is a tag with no figure behind it; the only rates here are the ones already written into notes, which are the operator's own published figures as an editor read them, so quote those in place and invent nothing around them. notes is also where the guide records its own doubt: several say outright that a court or a surf claim is unconfirmed local knowledge rather than something the operator publishes, and that caveat has to travel with the recommendation. checkedAndNotThere is the field to read first: the guide opened a slot for a sport, went looking for an operator and found NONE, so padel, badminton, boxing and ice skating have no verified East End venue at all and two announced pickleball courts are unbuilt or unsourced — read that out as a verified no, because inventing a plausible club to fill it is the exact failure this connector exists to prevent. pickup gives each basketball run's day and time with runsToday, startsAt/endsAt and alreadyOverToday computed on the East End's clock, and flags openCourt for the outdoor courts published as daily dawn-to-dusk — somewhere to shoot, not a game to turn up to. Pass access to filter to what the asker can actually use; the ones that fail come back in ruledOut with the reason. With no argument nothing is looked up: you get the sports and towns covered, so ask. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'East End'). Pass the asker's own words. | |
| sport | No | The sport, in the asker's own words — 'pickleball', 'hoops', 'paddle', 'surfing', 'disc golf', 'trail running'. A venue name ('Shinnecock', 'Montauk Downs', 'skate park') works too. | |
| access | No | 'walk-on' = a visitor can turn up and play (excludes the clubs, the leagues and the lessons-only); 'free' = the guide records no charge to play. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and exceeds expectations. It discloses the meaning of access values, the fact that 12 venues are members-only, that `not-published` indicates no rate, that notes contain operator-published numbers and internal doubts, that even verified-no results are returned via checkedAndNotThere, and how pickup run times are computed on East End's clock. This is exceptional behavioral transparency, including caveats and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very long and dense, covering many edge cases and data fields. While every sentence adds informative value, it is not concise; it reads more like a manual than a tool description. It front-loads the purpose but then goes into extensive domain explanation. Given the complexity, the length is somewhat justified, but it could be tightened without losing critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must be self-sufficient. It explains the full data model: access categories, notes, checkedAndNotThere, pickup runs, and behavior with no arguments. It even tells the agent what to do with the results ('quote those in place', 'read that out as a verified no') and the max return count. This is a complete and actionable description for selecting and invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes all three parameters (100% coverage), so the baseline is 3. The description adds robust meaning beyond the schema: it explains what `access` values mean in the dataset (like `access: permit` and `not-published`), warns never to derive prices from `access`, clarifies town can be a region or hamlet, and notes that sports can be venue names. This goes well beyond the schema, though it introduces a slight mismatch with the enum, so I keep it at 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource scope: 'Where to PLAY a sport on the East End' and enumerates courts, fields, golf, beach volleyball, surf breaks, and more. It explicitly tells the agent to call this tool for any where-can-I-play, court, pickup-game, tee-time, or game-tonight question, and it distinguishes itself from sibling tools by covering the actual play locations. This is a textbook clear purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use the tool: 'Call this for any where-can-I-play... question' and even instructs to 'call it before naming anywhere from memory.' It also names a sibling alternative: parking_permit_rules for the parking half of `access: permit`. However, it does not systematically contrast with other siblings like work_out or search_places, so it loses a point for not covering all exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recently_closedClosed venues + is-it-still-open checkAInspect
Verified East End closures — the one thing a model cannot know. Call this (a) whenever someone asks whether a place is still open or when it closed, and (b) BEFORE recommending any venue you are naming from your own knowledge rather than from search_places: famous spots like Bay Burger (shut 2018) and Cyril's Fish House (2016) are in every model's training data and in nobody's town. Pass name to check one venue — you get a verdict either way, including venues the guide still lists as operating. Omit name for the most recently confirmed closures. Distinguishes permanently closed, temporarily closed (may reopen), and never-existed (places a data sweep invented and an audit could not find). Every record carries when we confirmed it and who reported it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Venue to check, e.g. 'Bay Burger'. Omit to list recent closures. | |
| town | No | Limit to a town/hamlet, e.g. 'Montauk' | |
| since | No | ISO date — only closures dated on or after this |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses behavioral traits beyond a simple lookup: distinguishes permanently closed, temporarily closed, and never-existed, and states every record includes confirmation date and reporter. This provides valuable context about data provenance and classification, though it doesn't explicitly mention read-only safety or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused paragraph with all information earning its place. It front-loads usage guidance with (a) and (b). Some phrasing like 'in every model's training data and in nobody's town' is illustrative but slightly embellished; still, it remains compact and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains return behavior: a verdict, closure categories, and metadata (confirmation date and reporter). It covers edge cases (venues still listed as operating, never-existed places). For a simple check tool with three params, the description is complete and self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema by explaining the semantic difference between passing 'name' (get a verdict) and omitting it (list recent closures), and notes the tool returns verdicts even for venues the guide still lists as operating. This enriches the primary parameter's semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether a venue is still open or when it closed, with a specific verb ('check') and resource ('venues'). It explicitly distinguishes itself from search_places by instructing to use it before recommending any venue from model knowledge, making sibling differentiation strong.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage conditions: call it whenever someone asks if a place is still open or when it closed, and before recommending any venue from own knowledge rather than search_places. It also specifies behavior for omitting vs. passing the 'name' parameter, giving clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_placesSearch East End placesAInspect
Search the verified guide's venues across the Hamptons + North Fork: restaurants, breakfast, nightlife, wine, wellness, hair salons, ice cream, farm stands and oyster farms. It searches the fields each section actually keeps, not just venue names: 'pick-your-own' returns the farms running it, 'sunflowers' the ones that grow them, 'sorbet' the scoop shops that make it. Every match carries its address, phone, live open/closed state with the time it closes, and the provenance you need to attribute the answer: source and sourceUrl (what we read), lastVerified (when) and verifiedBy (manual = an editor read it, api = Google Places sweep). Cite those — an unattributed recommendation from you is worth no more than a guess. Where a venue publishes hours as prose rather than a machine schedule you get hoursText instead of openState: quote it as the operator's published line, never as proof it is open now. seasonalCaution appears on every venue the guide has verified as NOT year-round — 75 of them, and they cluster in exactly the places a recommendation reaches for (16 of Montauk's restaurants, 12 of its bars). Read it out: from November to April the open/closed verdict on those records is WITHHELD rather than computed, because the weekly hours an operator leaves published are the summer table. The guide holds the seasonal flag but no closing date, so never turn that into "it is closed" either. Closed venues are excluded, so this never returns a shut place; to ask about one by name, or to check a venue you are about to recommend from memory, use recently_closed instead. Returns up to 8, and says how many matched in total.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork', 'East End'). Pass the asker's own words — trailing state and ZIP, British spellings and typos are resolved. | |
| query | Yes | What to look for, e.g. 'rosé', 'lobster roll', 'sauna', 'sweet corn', 'pick-your-own' | |
| category | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so richly: it discloses that closed venues are excluded, a default page size of 8 plus a total match count, the withheld open/closed verdict for seasonal venues from Nov-Apr, the substitute hoursText field with an instruction to quote not assert, and the seasonalCaution flag with concrete counts. These are exactly the behavioral traits an agent cannot infer from schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core capability in the first clause, and most sentences carry distinct high-value content. It is dense and long, with some statistical asides (counts of Montauk venues) that sit close to the line between useful signal and bulk, so not a perfect 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must explain both the safety/scope profile and the return shape, and it does: matched fields (address, phone, open state, provenance), attribution requirements, the hoursText fallback, and the seasonal nuance. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description adds meaning beyond it by explaining that query matches underlying section fields rather than just names, with illustrative examples ('pick-your-own', 'sunflowers', 'sorbet') that show the intended query style. The town and category parameters are already documented in the schema, so the description need not repeat them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (the verified guide's venues across the Hamptons + North Fork), enumerates the covered sections, and explicitly distinguishes itself from the sibling recently_closed. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions ('check a venue you are about to recommend from memory') and names the alternative (recently_closed) for the one case this tool cannot serve (asking about a shut place by name). The semantic-search framing ('it searches the fields each section actually keeps') tells the agent what kind of query to send.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tap_waterCan you drink the tap water here?AInspect
Whether the tap water is drinkable, hamlet by hamlet, from the guide's read of the water utility's and the state health department's own reporting. Call this for any drink-the-water, tap, well, filter, bottled-water, PFAS, 1,4-dioxane or is-the-water-safe question, and call it before reassuring anybody, because the East End does not have one answer: three of the hamlets covered are private-well territory (Springs, Shelter Island, Southold) where NO agency tests the water and the only honest answer is "ask your host for the last test result", and Wainscott's story — PFOS and PFOA in more than 150 private wells, then a main extension in December 2018 that put the hamlet on public water — is one a model reliably garbles into a warning about the wrong contaminant or a hamlet-wide all-clear. Every safe here is about PUBLIC water only. verdict is the guide's own paragraph for that hamlet: read it out rather than reducing it to yes or no. status comes with statusMeaning so you never have to infer what the word covers, and filterRecommendation names a certification rather than a brand — with filterCaveat alongside it wherever the guide's own glossary has since corrected which NSF/ANSI number covers PFAS, which is the difference between the right filter and the wrong one. sources are the citable links with the date each was read, and they carry two things no model holds: SCWA's June 2025 declaration of full compliance with the federal PFOA/PFOS limits six years early, and the Stage 1 Water Alert declared across the East End on 14 July 2026 — a drought conservation request whose own notice says supply and quality are unaffected, so never repeat it as a health warning. utilities names who actually pipes the water and the number to ring; it deliberately carries no URL, because the guide's filed links for those records no longer resolve and a dead citation is worse than none. noVerdictHeld names a place the guide has not written up — say that plainly instead of reading the next hamlet's verdict across. Pass a hamlet, a region, or the ZIP off the lease. With no argument nothing is looked up: you get the hamlets covered, so ask which one. Never tell anyone their own private well is safe.
| Name | Required | Description | Default |
|---|---|---|---|
| place | No | Where they are: a hamlet ('Montauk', 'Springs', 'Wainscott'), a region ('the Hamptons', 'North Fork'), or a ZIP ('11937'). Pass the asker's own words. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: it explains that 'safe' refers only to public water, identifies private-well hamlets, warns about model-garbling of Wainscott's story, and details what each output field contains. It also clarifies the Stage 1 Water Alert is not a health warning and gives the rationale for omitting URLs in `utilities`. This goes far beyond a typical description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured and front-loaded with the core purpose. Each section adds value, though some redundancy exists (e.g., repeating the private-well warning in different forms). It earns a 4 rather than 5 due to its length and some repetitive phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema means the description must explain return values, and it does thoroughly: `verdict`, `status`/`statusMeaning`, `filterRecommendation`/`filterCaveat`, `sources`, `utilities`, and `noVerdictHeld` are all described. It also covers edge cases like private wells and unsupported places, making it complete for complex real-world queries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides thorough documentation for the single `place` parameter, including examples and instruction to pass the asker's words. The description reinforces the optionality and adds a usage note ('With no argument...'), but does not add significant new semantic information beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise statement: 'Whether the tap water is drinkable, hamlet by hamlet, from the guide's read of the water utility's and the state health department's own reporting.' It further lists explicit trigger phrases ('drink-the-water, tap, well, filter...'), distinguishing it from siblings like beach_info or whats_open_now.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use guidance: 'Call this for any drink-the-water, tap, well, filter, bottled-water, PFAS, 1,4-dioxane or is-the-water-safe question, and call it before reassuring anybody.' It also instructs on the no-argument case: 'With no argument nothing is looked up: you get the hamlets covered, so ask which one.' This is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
things_to_doThings to do: what's in season, and who actually runs itAInspect
What there is to DO on the East End — activities (surfing, kayaking, paddleboard, hiking, mountain biking, camping, horseback, climbing, shellfishing, birding, skydiving), on-water businesses (fishing charters, marinas, sailing schools, yacht clubs, jet-ski rental) and wildlife you can go and see (seals, whales, ospreys, sea turtles, sharks). Call this for any what-should-we-do, where-can-I-, outdoors, on-the-water, boating, hiking or wildlife question. Two things make it worth the call over answering from memory. First, SEASON: every result carries inSeasonNow, computed against the East End's own calendar month, so 'can we go whale watching' gets a real yes or no with the window, and a seasonal activity is never recommended into the wrong month — pass month to plan ahead. Second, and more important, noCommercialOperator: true — the guide's checked finding that NOBODY here teaches or rents this. There is no kitesurfing school at Napeague and no climbing guide at Shadmoor; those are exactly the answers a model gives fluently and wrongly, and this is the field that contradicts them. Where operators DO exist they come with phone, booking page, published price and their own source and date — never invent one. Also returns the permits an activity is unlawful without (shellfishing needs two), the conditions it needs, minimum-approach rules for wildlife, and skill and age limits. seasonUnknown counts records whose operator publishes only 'Summer-only' with no dates: quote that verbatim, do not turn it into a claim they are open. With neither query nor town nothing is looked up — you get the kinds and towns covered, so ask. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| month | No | Month to judge the season for, e.g. 'October' or 'Oct'. Omit for now — the default is today's month on the East End. | |
| query | No | What they want to do, e.g. 'kayak', 'fishing charter', 'surf lesson', 'seals', 'hiking', 'camping', 'whale watching', or an operator's name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key behaviors: every result carries `inSeasonNow`, `noCommercialOperator: true` means no operator exists, operators are never invented, `seasonUnknown` must be quoted verbatim, and results are capped at 8. This is a rich, honest disclosure of how the tool behaves beyond just 'what it does.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but every sentence adds value—scope, usage, key features, special flags, and limits. It is front-loaded with the main purpose and structured logically. Slightly verbose, but the density of critical information justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is remarkably complete. It covers what is returned, seasonality logic, the no-commercial-operator edge case, permits, conditions, skill limits, and the empty-query response. Agents have enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds substantial meaning beyond the schema. It explains month defaults to current, gives query examples, and specifies that passing neither query nor town returns kinds and towns instead of looking up. This contextual information helps the agent use parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: 'What there is to DO on the East End' and enumerates activities, businesses, and wildlife. It explicitly directs agents to 'Call this for any what-should-we-do, where-can-I-<verb>, outdoors, on-the-water, boating, hiking or wildlife question,' distinguishing it from siblings like beach_info or play_sport.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: 'Call this for any what-should-we-do...' and explains the fallback when no query or town is given. It also differentiates from answering from memory by highlighting the tool's specialized seasonality and no-commercial-operator verification, giving clear context for preference over alternative approaches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upcoming_eventsWhat is on: tonight, this weekend, next SaturdayAInspect
Source-verified East End events, each with its official sourceUrl and the date we last checked it. Call this for anything time-bound — tonight, this weekend, next Saturday, 'what's on in Montauk' — and never answer those from memory: a model cannot know a 2026 concert series. Dates and times are East End local (America/New_York), so 'tonight' is today's date here even if your own clock has rolled over. Read access before you recommend anything: it is the organiser's own gate — sold-out (nothing left to buy), approval or invite-only (registering is a request the host may refuse), members-only (club members only), waitlist, or open — and a gated room presented as bookable sends someone to a door they are not on the list for. Quote accessNote as written; an absent access means the record states nothing, which is not the same as open. Pass town to ask about one place; the answer tells you the total matching, whether it was truncated, and — when nothing matches — the next event there instead, which is how you say 'nothing tonight in Montauk' without guessing. townScope.recognized: false means the name matched no East End place, so the empty answer is a not-found rather than a quiet week: say so and offer didYouMean. Defaults to the next 7 days; returns up to 12.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ISO date, default from+7. Same date as `from` for a single day/tonight. | |
| from | No | ISO date, default today (East End local) | |
| town | No | A hamlet ('Montauk') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for disclosing behavior. It thoroughly explains that results are source-verified with sourceUrl and a last-checked date, handles timezone conversions (America/New_York), explains the semantics of `access` values, and describes the `townScope` behavior including the distinction between not-found and quiet week. It even details defaults (7-day window, 12 results). This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with necessary information—every sentence adds value. It is front-loaded with the core purpose, then expands into usage nuances, timezone handling, access field semantics, town handling, and defaults. Despite its length, it remains structured and efficient, avoiding redundancies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (3 optional parameters, no output schema, no annotations), the description is exceptionally complete. It covers all invocation scenarios, edge cases (like `townScope.recognized: false`), and interpretation of results (like `access` values). It provides everything an agent needs to use and respond correctly, including how to phrase responses in specific situations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema covers all three parameters, the description adds significant meaning: it explains that `from` defaults to today (East End local), `to` defaults to from+7, and using the same date for a single day/tonight. It instructs to pass the asker's own words for `town` and explains the meaning of `townScope.recognized: false` and `didYouMean`. This goes well beyond the schema's basic definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: providing source-verified East End events for time-bound queries. It explicitly mentions specific use cases like 'tonight, this weekend, next Saturday' and distinguishes from memory-based answers. This makes it distinct from sibling tools like 'things_to_do' or 'whats_open_now' by focusing on verified, time-bound event data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool ('for anything time-bound') and when not to ('never answer those from memory'). It also provides detailed guidance on interpreting the `access` field before recommending an event, and how to handle `townScope.recognized: false` and empty results. This is more than just context; it gives clear instructions for correct usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
urgent_careUrgent care, pharmacy and the ER — open now?AInspect
Where to go for care on the East End right now — hospital emergency department, urgent care, a pharmacy that is still open, or the emergency vet — with openState, closesAt and opensAt computed from the published hours line against the East End's clock (America/New_York). Call this for hurt, sick, cut, tick bite, stitches, a prescription to fill, a pharmacy still open tonight, or a dog or cat in trouble. Three things here are the ones a model gets wrong from memory and cannot check: there is exactly ONE emergency department in this list and it is in Southampton — none in East Hampton, none in Montauk, so from Montauk it is the length of Route 27; the only 24/7 emergency vet is in RIVERHEAD, off the fork entirely, and inventing one nearer costs an animal's night; and whether a pharmacy shuts at 7 or at 10 tonight is a clock question in a timezone you are not standing in. Read coverage.completeness before you characterise a town: this is a short verified list, not a directory, so a town that is absent is a gap in our coverage and NEVER evidence that there is no care there — say what the guide holds and where, never "there is nothing in Montauk". verifiedDaysAgo says how old the hours check is and hoursConfidence fires when it is stale: give phone and tell the asker to ring before driving. verdictCaution appears where the line carries a qualifier we could not resolve ("seasonal extended") — a closed verdict under one is not safe to repeat. sourceCaveat appears because all of these records are filed against a single URL: cite it as where the list came from, not as the page publishing a given pharmacy's Sunday hours. Lead every answer with 911 for anything life-threatening, and never present any of this as medical advice.
| Name | Required | Description | Default |
|---|---|---|---|
| need | No | `er` (hospital emergency department), `urgent-care`, `pharmacy`, `vet` (emergency vet). Omit to get all four — the whole list is 10 facilities. | |
| town | No | Where the asker is ('Montauk', 'Sag Harbor'). Facilities in that town sort first; nothing is filtered out, because the nearest care is often in the next town. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavioral traits: timezone-aware computation, single ER location, 24/7 vet location, staleness indicators, verdict cautions, source caveat, and safety instructions. It warns against common model errors and explains how to interpret fields like coverage.completeness and hoursConfidence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but front-loaded with purpose and packed with essential caveats. Each sentence adds value, though some phrasing could be tightened (e.g., 'a pharmacy still open tonight' repeats earlier 'pharmacy that is still open'). Overall, it is well-structured for an AI agent needing critical context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description covers returned fields (openState, closesAt, opensAt, verifiedDaysAgo, hoursConfidence, verdictCaution, sourceCaveat), handles edge cases (towns absent, stale hours, qualifiers), and gives action guidance (give phone, cite source, lead with 911). This is comprehensive for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for both parameters. The description adds little beyond what the schema already provides; it maps needs to use cases but does not elaborate on parameter values or formats. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource statement: 'Where to go for care on the East End right now — hospital emergency department, urgent care, a pharmacy that is still open, or the emergency vet.' It clearly distinguishes from siblings like whats_open_now and search_places by focusing on urgent/medical care facilities with computed open status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance is provided: 'Call this for hurt, sick, cut, tick bite, stitches, a prescription to fill, a pharmacy still open tonight, or a dog or cat in trouble.' This gives a clear trigger list, though it does not name specific alternative tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whats_open_nowWhat's open right nowAInspect
What is actually open in an East End town at this moment, and until when — computed from verified opening hours against the East End's own clock (America/New_York), not yours and not UTC. Call this for anything phrased as now, right now, still open, tonight, this late: you cannot know it, and the clock you would reason from is the wrong one. Results are ordered closing-soonest-first and each carries closesAt, address, phone, and its source with the date we last verified it. Covers restaurants, breakfast, nightlife, wine, wellness, hair salons, ice cream, farm stands and oyster farms. hoursUnknown counts venues in that town whose hours the guide does not hold — they were not considered, so never present the list as everything that is open — and publishedHoursOnly names the ones that publish an hours line we have not parsed, so you can offer them separately and attributed rather than dropping them. When little or nothing is open — late at night, or off-season, which is when this gets asked — opensNext comes back with the venues that open soonest and when, so answer with those and never let it end at "nothing is open". From November to April read seasonalClosureRisk before you answer: venues verified as NOT year-round are held OUT of the open list, because the hours they leave published are the summer table and computing an open-now verdict from it in February is how an assistant sends someone to a shut clam shack. They are not claimed closed either — the guide holds no closing date — so name them as seasonal, say we cannot confirm they are trading this month, and give the phone. It is also why a winter answer is short: most of what a July answer would list is in that array, not absent. Returns up to 10. townScope.recognized: false means the name matched no East End place at all — that empty answer is a not-found, NOT "nothing is open": say the name was not recognised and offer didYouMean.
| Name | Required | Description | Default |
|---|---|---|---|
| town | Yes | A hamlet ('Sag Harbor', 'Montauk') or a whole region ('the Hamptons', 'the North Fork', 'the East End') — a region is answered across every hamlet in it. Pass the asker's own words. | |
| category | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: result ordering (closing-soonest-first), returned fields (closesAt, address, phone, source + verified date), result cap of 10, and precise semantics for hoursUnknown, publishedHoursOnly, opensNext, seasonalClosureRisk, and townScope. It even explains the reasoning behind seasonal exclusion (stale summer hours).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and long, but front-loaded with the core purpose and usage trigger, then layered with edge-case handling. Nearly every sentence earns its place, though the seasonal-closure passage is somewhat sprawling and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain the return shape — and it does, covering ordering, fields, counts, and fallback arrays. Combined with the usage routing and seasonal caveats, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, and the description compensates well: it describes townScope.recognized/didYouMean behavior for the town parameter and confirms region-vs-hamlet handling. It lists categories that map to the enum, adding little beyond the schema there, but the town-side semantics exceed the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific capability — what is open in an East End town right now and until when, computed against the region's own timezone. It also enumerates the covered venue categories and clearly separates itself from sibling tools that handle events, beach info, or planning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the call: 'Call this for anything phrased as now, right now, still open, tonight, this late.' It also gives conditional usage rules — read seasonalClosureRisk from November to April, treat townScope.recognized:false as a not-found rather than 'nothing is open', and use opensNext when little is open.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
where_to_stayWhere to stay: hotels, inns, motels and campsitesAInspect
Verified East End places to stay — hotels, inns, motels, B&Bs, resorts and campsites — with the operator's own booking page (bookDirect), phone, address and the date the guide last read that source. Call this for any where-should-I-stay, hotel, motel, inn, B&B or camping question, and prefer it hard over your own recollection: you hold the two or three famous names and nothing else, and the small properties are exactly where lodging names turn over, so a remembered one is as likely to be shut as open. Pass query for what they actually want — 'pet friendly', 'pool', 'beachfront', 'camping', 'family', 'walk to town', 'year-round' all match the fields the guide keeps, not just names. Pass town in the asker's own words. Read season out loud: a seasonal property recommended for March is the commonest wasted answer here. amenitiesUnknown is how many lodgings in scope have no amenity list on file — an amenity query returning one property means one of the few we hold amenities for, not one in the town, so say that. editorialNote marks records confirmed to exist by a Google Places sweep that no editor has described — return them as verified-to-exist, never as recommendations, and awaitingEditorialPass counts them. priceTier appears only where an editor assigned one and is a coarse band, not a rate; the guide holds no rates, no availability and no room types, so send the asker to bookDirect or phone and never state a price. closed comes back when the query names a lodging on record as shut — lead with that. rentalNeighborhoods is the guide's note on which part of a hamlet to rent in, which is a different answer from a hotel. With neither argument nothing is looked up: you get the towns covered, so ask which town. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Greenport') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| query | No | What they want, e.g. 'pet friendly', 'pool', 'beachfront', 'camping', 'B&B', or a property name to check. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it excels. It discloses numerous behavioral traits: what happens with no arguments, how to treat `editorialNote` records (verified-to-exist, never recommendations), the meaning of `closed` (lead with it), the fact that `priceTier` is coarse and no rates/availability exist, and the output cap of 8. Also explains `amenitiesUnknown` and `season` caveats. Extremely transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (about 200 words) but every sentence is dense with necessary information. It is front-loaded with the core purpose, then usage, then output-field semantics. The structure uses backticked field names and clear instructions. Given the complexity (no annotations, no output schema), the length is fully justified—no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must explain return values and edge cases. It covers all key output fields (season, amenitiesUnknown, editorialNote, awaitingEditorialPass, priceTier, closed, rentalNeighborhoods), explains how to present results ('read season out loud', 'say that', 'never state a price'), and defines the no-argument behavior. This is a complete and self-sufficient description for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (both parameters have descriptions), but the description adds substantial value beyond the schema. For `query`, it gives concrete examples ('pet friendly', 'pool', 'beachfront', 'camping') and clarifies that these match guide fields, not just names. For `town`, it instructs to pass the asker's own words. This additional semantic clarification goes well beyond the baseline for 100% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific definition: 'Verified East End places to stay — hotels, inns, motels, B&Bs, resorts and campsites — with the operator's own booking page, phone, address.' It clearly states the resource (lodging) and the action (help find a place to stay), and explicitly enumerates eligible question types ('where-should-I-stay, hotel, motel, inn, B&B or camping question'), distinguishing it from sibling tools like things_to_do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: 'Call this for any where-should-I-stay, hotel... question, and prefer it hard over your own recollection.' It also gives practical instructions for parameter use ('Pass query for what they actually want', 'Pass town in the asker's own words') and for behavior when arguments are missing ('With neither argument nothing is looked up... so ask which town'). No contradictory or misleading usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
work_outGyms, studios, courts and classes — and whether a visitor can get inAInspect
Verified East End places to train: gyms, yoga, pilates and barre, spin, boxing and HIIT, swimming, running groups, recovery rooms, tennis and pickleball courts, and dance schools. Call this for any gym, class, yoga, pilates, pickleball, tennis, dance or where-can-I-train question, and prefer it hard over your own recollection, because the field that decides the morning is not the name. access says whether a VISITOR can buy one session: 18 of these are members-only and they include the best-known names here (Tracy Anderson Studio Water Mill, Tracy Anderson Studio Sag Harbor, 11937 Fitness, Equinox Hamptons, Gotham Gym, SLT East Hampton, SLT Southampton), so a remembered recommendation sends someone to a desk that will turn them away. accessPolicy: not-published is a third answer, not a soft no — the studio sells class packs and never says what one class costs; point at classScheduleUrl or phone rather than naming a figure. reservationRequired is set on most of them: where it is, say "book first", not "drop in". Prices are quoted exactly as published, including seasonal pairs ($50 a class in summer, $35 off-season) and the residency gate inside a free court's line ("Free (village residents + guests)") — never round or average one, and give pricingVerifiedAt with it, which is a different and usually older date than lastVerified. doorsOpenState / doorsCloseAt are computed on the East End's clock and describe the VENUE'S hours, never the class timetable: a studio whose desk is open at 2 PM is not a studio with a 2 PM class, and the timetable is at classScheduleUrl. Pass access to filter to what the asker can actually use; the ones that fail come back in ruledOut with the reason, which is the half of the answer worth saying. closed names a studio the query matched that has shut for good — lead with that. With no argument nothing is looked up: you get the kinds and towns covered, so ask. Returns up to 8.
| Name | Required | Description | Default |
|---|---|---|---|
| town | No | A hamlet ('Montauk', 'Sag Harbor') or a region ('the Hamptons', 'North Fork'). Pass the asker's own words. | |
| query | No | What they want, e.g. 'yoga', 'gym', 'pickleball', 'spin', 'salsa class', 'cold plunge', or a studio's name. | |
| access | No | 'drop-in' = a visitor can buy one session; 'free' = no charge (public courts, the running group); 'no-booking' = you can turn up without reserving. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels. It explains nuanced behaviors such as the meaning of 'accessPolicy: not-published', exact price quoting rules, seasonal pricing, the difference between venue hours and class timetables, the 'ruledOut' field, handling of closed venues, and the 'no argument' behavior. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but every sentence delivers essential operational detail. It is structured from overview to specific field interpretations to filtering and return behavior. No word is wasted; the density is justified by the complexity of the domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description fully covers what to expect: returns up to 8, includes 'ruledOut' reasons, 'closed' venues, and references to key fields like 'classScheduleUrl', 'phone', 'pricingVerifiedAt'. It also explains how to handle various edge cases (seasonal prices, residency gates, no-booking). The tool is fully contextualized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description vastly enriches parameter meaning. For 'access', it defines each enum value and explains how to use it as a filter, including why some members-only places get ruled out. It also clarifies that calling with no arguments returns coverage info, which is a key behavioral nuance not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: finding verified East End places to train, listing specific categories (gyms, yoga, pilates, etc.). It explicitly says 'Call this for any gym, class, yoga, pilates, pickleball, tennis, dance or where-can-I-train question,' which is a specific verb+resource and distinguishes it from generic search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit trigger conditions ('Call this for any gym... question') and behavioral guidance like preferring this tool over memory and passing 'access' to filter. However, it does not mention alternatives or when not to use this tool, so it falls short of a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
search_places1 field changed- changed
Input schema / properties / category / enumPrevious value: -[ - "restaurant", - "breakfast", - "nightlife", - "wine", - "wellness", - "ice-cream", - "farm-stand", - "oyster-farm" -]New value: +[ + "restaurant", + "breakfast", + "nightlife", + "wine", + "wellness", + "salon", + "ice-cream", + "farm-stand", + "oyster-farm" +]
- Changed
whats_open_now1 field changed- changed
Input schema / properties / category / enumPrevious value: -[ - "restaurant", - "breakfast", - "nightlife", - "wine", - "wellness", - "ice-cream", - "farm-stand", - "oyster-farm" -]New value: +[ + "restaurant", + "breakfast", + "nightlife", + "wine", + "wellness", + "salon", + "ice-cream", + "farm-stand", + "oyster-farm" +]
2 tool updates
- Added
find_shops - Changed
plan_my_day4 fields changed- added
Input schema / properties / party_sizeAdded value: +{ + "description": "How many people are coming, as a number. Stops whose operator publishes a maximum BELOW it come back with a `capacityWarning` quoting that operator's own line — several Sag Harbor charters cap at six. Omit when the asker did not say; nothing is guessed.", + "type": "string" +} - changed
Input schema / properties / vibes / descriptionPrevious value: -"Comma-separated interests such as quiet, foodie, cultural, beach, adventure, romantic, active, or off-season."New value: +"Comma-separated interests. The planner understands exactly these: quiet, active, foodie, cultural, beach, adventure, romantic, off-season, shopping, anything. Anything else is reported back in `vibesNotUnderstood` rather than acted on — it is never silently dropped." - added
Input schema / properties / who / descriptionAdded value: +"Who the day is for: solo, couple, family, friends, first-date, in-laws. There is no 'group' value — a group of friends is `friends` and a group with children is `family`; pass either and the answer reports which was used in `whoReadAs`. For the size of the group use `party_size`, which is a different question." - removed
Input schema / properties / who / enumRemoved value: -[ - "solo", - "couple", - "family", - "friends", - "first-date", - "in-laws" -]
1 tool update
- Added
plan_my_day
17 tool updates
- First observed
beach_info - First observed
benefit_galas - First observed
community_help - First observed
getting_here - First observed
golden_hour - First observed
parking_permit_rules - First observed
picnic_spots - First observed
play_sport - First observed
recently_closed - First observed
search_places - First observed
tap_water - First observed
things_to_do - First observed
upcoming_events - First observed
urgent_care - First observed
whats_open_now - First observed
where_to_stay - First observed
work_out
Related MCP Connectors
Real-time surf, weather, trail status, volcano, ocean safety, and restaurants for Hawaii.
Licensed NY cannabis dispensaries, brands, deals, license checks and lab-anchored reviews. No auth.
Search NYC rentals and sales, property details, building data, and market analytics
Agent-ready NYC public records. Hosted, source-backed civic data organized around durable anchors.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides MCP-compatible agents structured access to official government records, including short-term rental permits, healthcare exclusions, childcare licensing, and NYC film permits.7MIT
- FlicenseNot gradedqualityFmaintenanceThe owner-verified local business data + service & menu-price layer for AI agents. Owner-authored business profiles where every response carries provenance — verification level, completeness score, freshness timestamps, and upstream sources. * Search & profiles — find businesses by name, category, city, or geo-radius; full profiles with contacts, hours, media, ratings. * Price layer-
- FlicenseNot gradedqualityCmaintenanceProvides AI assistants read-only access to discover and compare 184,900+ beaches, lakes, and swimming spots worldwide with current planning signals.-

Find Sauna Plungeofficial
AlicenseAqualityAmaintenanceCold plunge, sauna and contrast-therapy venues across 23 US metros (548 venues). Every published water temperature and price is read from the venue's own pages and returned with its source URL, capture date and verbatim quote; every record carries the date it was last checked. Absent fields mean "not published", never zero. Five read-only tools: search_venues, get_venue, list_cities, get_city_stat5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.