united__flight_status
[united · risk:low] Check the status of a United flight
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Flight date (YYYY-MM-DD) | |
| flight_number | Yes | Flight number, e.g. UA123 |
[united · risk:low] Check the status of a United flight
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Flight date (YYYY-MM-DD) | |
| flight_number | Yes | Flight number, e.g. UA123 |
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full burden. It mentions 'risk:low' but fails to disclose behavioral traits such as authentication needs, data freshness, or rate limits. A flight status tool ideally should note these.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, very concise. However, the prefix '[united · risk:low]' is metadata that could be clearer. Still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status check, the description is minimal. It does not mention what the response contains (e.g., delay, gate, time). Given no output schema, some return value context would be helpful. Adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for both parameters. The description adds no additional meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the status of a United flight, using a specific verb and resource. The prefix '[united · risk:low]' reinforces the airline scope, distinguishing it from sibling tools like southwest__flight_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The implication is for United flights, but there are no explicit when-to-use, when-not-to-use, or sibling comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.
Each tool is prefixed with its service name, and within each service, tools have distinct actions (e.g., mail_read vs. availability_find). The duvera tools cover different subdomains like dev, finance, and food with no overlap, making selection unambiguous.
All tools follow the pattern service__action_object or service__category_action, using lowercase with underscores. The order of verb and noun varies slightly (e.g., package_track vs. boardingpass_show), but the naming is highly predictable and readable.
With 52 tools, the server is large but justified as a gateway aggregating many external services. Each tool corresponds to a common task for its service, so no tool feels extraneous, though the total number is high.
The tool surface covers a wide array of services but only provides one or two basic operations per service (mostly read-only). While this suits a quick-lookup gateway, deeper workflows (e.g., creating or updating resources) are missing, leaving gaps for many use cases.