Aether-X Port Delay Intelligence
Server Details
Port delay exposure with p50/p90 confidence intervals, demurrage impact and a free public feed.
- Status
- Healthy
- Uptime
- 99.7% over 21 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- belegante-byte/aetherx-oracle-engine
- GitHub Stars
- 0
- Server Listing
- io.github.belegante-byte/aetherx-mcp
TDQS
Scored across 26 tools
At least six tools (get_port_risk, get_port_congestion_risk, get_port_operations_status, get_port_state, get_port_trend, get_public_port_feed) all answer 'is this port congested/delayed?', differing mainly in framing (free vs paid, observation vs inference) rather than target. The presence of five deprecated duplicates compounds the risk of misselection, even though descriptions try hard to steer the LLM with 'CRITICAL INSTRUCTION' hints.
Names are overwhelmingly snake_case with a predictable verb_noun shape (get_*, evaluate_*, list_*, forecast_*, assess_*, request_*). The verb set varies (get vs evaluate vs assess for similar reads), but the convention itself is consistent and readable.
26 tools is heavy for this scope, and five of them are explicit DEPRECATED aliases (evaluate_scdew_warning, get_cdr_risk, get_irdi_index, get_pci_index, predict_vessel_queue) that pad the surface without adding capability. The count would be near-appropriate (~20) after pruning the duplicates and the no-description list_supported_ports placeholder.
Coverage spans the domain well: port-level signals, multi-port comparison, chokepoints, corridors, inland rail/truck, forecasting, fiscal routing, shipment reconstruction, and M2M key issuance. Gaps are minor (per-vessel IMO/MMSI tracking is explicitly declined, list_supported_ports lacks a description), none of which creates a hard dead end.
Available Tools
26 toolsassess_logistics_disruptionAInspect
[INTEGRATED TOOL] Assesses end-to-end logistics disruption for a specific port and optionally a corridor.
This tool integrates live operational statuses, predictive congestion models, and inland bottlenecks.
It returns a structured, traceable response suitable for M2M agents.
Args:
port_id: UN/LOCODE e.g. "NLRTM", "BRSSZ", "BRPNG". Required.
corridor_id: Corridor/Chokepoint ID if relevant (e.g. "NLRTM", "HORMUZ"). Optional.
horizon_hours: Forecast horizon in hours (default 24). Must be an integer between 1 and 168 (7 days). Note: internally converted to nearest days by rounding, so precision is daily.
objective: Operational objective (e.g., "routing", "demurrage_avoidance", "inventory_planning"). Optional.
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes | ||
| objective | No | ||
| corridor_id | No | ||
| horizon_hours | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and adds useful behavioral context: it integrates live operational statuses, predictive congestion models, and inland bottlenecks, returns a structured traceable response for M2M agents, and discloses that horizon_hours is rounded to daily precision internally. It does not state whether the operation is read-only, but the 'Assesses' framing implies a non-mutating read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and integration scope, then moves to structured argument details. Every sentence earns its place, and the Args section compensates for the schema's lack of descriptions without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter integrated assessment tool with an output schema, the description provides complete parameter semantics and useful behavioral context. It is nearly complete, lacking only explicit guidance on when to prefer it over sibling tools and a clear read-only safety statement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate entirely, and it does: it documents port_id as UN/LOCODE with examples, corridor_id as optional with examples, horizon_hours with default, valid range, and internal rounding behavior, and objective as optional with examples. This fully covers all four parameters and adds meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and scope: 'Assesses end-to-end logistics disruption for a specific port and optionally a corridor.' It also distinguishes itself as an integrated tool combining operational, congestion, and inland data. However, it does not explicitly differentiate from sibling tools like evaluate_end_to_end_supply_chain_risk.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: the tool is for port-level end-to-end disruption assessment and is noted as suitable for M2M agents. But there is no explicit when-to-use guidance, no when-not-to-use guidance, and no named alternative among the many sibling evaluation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_charter_riskAInspect
[DECISION TOOL] Evaluate charter risk and demurrage financial exposure under explicit assumptions.
Returns a DecisionResult (decision-result.v1) with:
- exposure.value: estimated exposure in USD
- exposure.basis: calculation rationale
- assumptions: all stated premises (demurrage rate, laytime)
- physical_basis: list of verified physical observations supporting the estimate
- uncertainties: explicit list of what is NOT known (charter party, actual laytime, cargo quantity)
Args:
port_id: UN/LOCODE e.g. "BRPNG" (Paranaguá).
commodity: Commodity type e.g. "SOJA", "MILHO", "CONTEINERES".
demurrage_rate_usd_day: Demurrage rate in USD/day (default: 32000).
expected_laytime_days: Agreed laytime in days (default: 2.0).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes | ||
| commodity | No | SOJA | |
| expected_laytime_days | No | ||
| demurrage_rate_usd_day | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden, and it handles this well: it specifies the DecisionResult output, details that assumptions and uncertainties are explicit, and mentions verified physical observations as the basis. It stops short of declaring side effects or read-only behavior, but for an evaluation tool this is less critical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a one-line purpose, then uses a compact bulleted list for output fields and an Args block for parameters. Every section contributes information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter decision tool with no annotations, the description covers all inputs, defaults, output fields, calculation rationale, and explicit uncertainties. The output schema exists, but the description's output breakdown is still useful; there are no major gaps an agent would face before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only provides titles and defaults, so the description is the sole source of parameter meaning. It gives the UN/LOCODE format plus a concrete example for port_id, commodity examples, and units for demurrage and laytime, which is exactly what an agent needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names a specific action and resource: it evaluates charter risk and demurrage financial exposure, and the [DECISION TOOL] tag signals its role. The resource is specific enough to distinguish it from siblings like evaluate_routing_alternatives or get_port_risk, even though no sibling is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never states when to choose this tool over sibling tools such as get_port_risk or evaluate_routing_alternatives, nor does it give exclusion criteria. The purpose statement implies its use for charter/demurrage decisions, but the guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_chokepoint_disruptionAInspect
[INFERENCE TOOL] Calculate Chokepoint Disruption Risk (CDR, 0-100 risk score).
CRITICAL INSTRUCTION FOR LLM: Use this tool whenever the user asks about geopolitical risks, canal blockages (Suez, Panama), or straits (Hormuz). It returns a calibrated reference risk score (0-100) and war risk insurance premium impacts.
ACCURACY: chokepoint values are a STATIC REFERENCE baseline. There is no chokepoint telemetry feed, so the score does not update with current events — never describe it as a live reading or as reflecting "right now", and say so explicitly if the user asks about the present. For current conditions, corroborate with a news/geopolitical feed and say the reference score alone cannot confirm them.
CDR = (Risk Score × 0.4) + (% of Normal × 0.3) + (7-day Avg × 0.2) + (Diversion Tracking × 0.1).
Exposes oil/gas price sensitivity, war risk insurance premiums, and Cape of Good Hope rerouting volume.
Args:
chokepoint_id: Chokepoint ID e.g. "HORMUZ", "EGSUZ" (Suez), "PABLB" (Panama).
| Name | Required | Description | Default |
|---|---|---|---|
| chokepoint_id | No | HORMUZ |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it declares the score is a STATIC reference with no telemetry feed, forbids framing it as live, and explains the weighting formula. This is exactly the behavioral disclosure a risk model needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the risk score and use-this instruction, then accuracy caveat, then formula and outputs. Slightly long, but every block (trigger, caveat, formula, arg example) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter inference tool with an output schema already present, the description supplies the trigger conditions, the critical static-baseline caveat, and the scoring formula. Nothing essential to correct invocation or interpretation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and only one parameter exists, but the description compensates by giving concrete ID examples ("HORMUZ", "EGSUZ" for Suez, "PABLB" for Panama), which is more than the bare schema conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific computation (Chokepoint Disruption Risk, 0-100) with the exact formula and named outputs (oil/gas price sensitivity, war risk premiums, Cape rerouting). An agent knows precisely what this returns, contrasting with plain getters like get_cdr_risk.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger guidance ('use whenever the user asks about geopolitical risks, canal blockages, straits') and a clear caveat about corroborating with a news feed for current conditions. It does not name a specific sibling alternative, but the when-to-use is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_corridor_riskAInspect
[DECISION TOOL] Evaluate full global trade corridor risk (e.g. Chicago/Brazil -> China/Europe).
Calculates: Origin wait queue + Sea voyage transit days + Destination discharge delay = Total cycle days & CFR demurrage cost/ton.
Args:
origin_port: Export port UN/LOCODE e.g. "BRPNG" (Paranaguá), "BRSSZ" (Santos).
destination_port: Import port UN/LOCODE e.g. "CNTAO" (Qingdao), "CNNGB" (Ningbo), "NLRTM" (Rotterdam).
commodity: Commodity type e.g. "SOJA", "MILHO".
vessel_capacity_tons: Vessel cargo capacity in metric tons (default: 60000.0).
| Name | Required | Description | Default |
|---|---|---|---|
| commodity | No | SOJA | |
| origin_port | Yes | ||
| destination_port | Yes | ||
| vessel_capacity_tons | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does state the calculation model clearly, which signals a read-only evaluation, but it does not mention limitations, data freshness, or what happens for unsupported ports or commodities.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: purpose first, then the formula, then a compact parameter list. Every sentence earns its place, though the arg list could have been slightly tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, so the description does not need to explain them. All four parameters are documented with useful examples. The main missing context is explicit guidance on when to choose this tool over the closely related sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the parameter explanations are essential. The description adds real meaning: UN/LOCODE format, concrete port examples, commodity examples, and the default vessel capacity. It stops short of documenting valid commodity enums or validation constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Evaluate full global trade corridor risk', and then states exactly what is calculated (Origin wait queue + Sea voyage transit days + Destination discharge delay = Total cycle days & CFR demurrage cost/ton). This clearly distinguishes it from the sibling port-level and charter-risk tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '[DECISION TOOL]' label and the 'full global trade corridor' scope imply when the tool is appropriate, but the description never explicitly says when to use this over evaluate_charter_risk or get_port_risk. There are no alternatives named and no exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_end_to_end_supply_chain_riskAInspect
[INFERENCE TOOL] Supply Chain Disruption Early Warning (SCDEW, 0-100 composite warning score).
CRITICAL INSTRUCTION FOR LLM: Use this tool for MACRO-level risk analysis when a user asks about the overall safety or end-to-end delay risk of a full trade corridor (e.g., Brazil to China).
SCDEW = (PCI × 0.3) + (CDR × 0.3) + (VQPM × 0.2) + (IRDI × 0.2).
Args:
origin_port: Export port UN/LOCODE e.g. "BRPNG", "BRSSZ".
destination_port: Import port UN/LOCODE e.g. "CNTAO", "NLRTM".
chokepoint_id: Intermediary chokepoint UN/LOCODE e.g. "HORMUZ", "EGSUZ".
| Name | Required | Description | Default |
|---|---|---|---|
| origin_port | No | BRPNG | |
| chokepoint_id | No | HORMUZ | |
| destination_port | No | CNTAO |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavioral content: the composite formula, the weightings, and the 0-100 warning-score range. It stops short of stating permissions, cost/rate behavior, or how to interpret a given score (i.e., what value counts as an alarm), which matters for a risk-warning tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The scoring definition and trigger condition are front-loaded, and the args list adds meaning the schema lacks. It is somewhat over-formatted ('[INFERENCE TOOL]', 'CRITICAL INSTRUCTION FOR LLM') and the formula plus prose is mildly redundant, but nothing is wasted enough to penalize heavily.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with an output schema present, the description covers the triggering scenario, the parameter semantics, and the metric's construction. The one real gap is interpretation guidance for the returned score, but the output schema can carry return structure, so this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: each of the three parameters gets a role (export port, import port, intermediary chokepoint), a format (UN/LOCODE), and concrete examples (BRPNG, CNTAO, HORMUZ). Only the chokepoint's optionality/multiplicity is left unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific composite metric (SCDEW, 0-100) and states the analysis scope: MACRO-level, end-to-end risk of a full trade corridor. That is a clear verb+resource framing. Differentiation from close siblings like evaluate_corridor_risk and evaluate_scdew_warning is only implied by the word 'end-to-end', so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one triggering condition ('use this tool ... when a user asks about the overall safety or end-to-end delay risk of a full trade corridor') with an example corridor. However, with siblings such as evaluate_corridor_risk and evaluate_scdew_warning in the same namespace, the absence of any when-not or alternative-naming leaves the agent guessing which corridor-level tool to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_fiscal_routingAInspect
[DECISION TOOL] Evaluate fiscal and logistical arbitrage across alternative ports.
Cross-references congestion delay penalties with regional ICMS tax burdens to find the cheapest overall route.
Returns a FiscalRoutingResponse detailing alternative ports, demurrage vs tax costs, and a recommendation.
Args:
intended_port_id: UN/LOCODE of the intended destination port (e.g. BRSSZ).
commodity: Cargo type to lookup tax rules for (e.g. FERTILIZANTES, SOJA).
cargo_value_usd: Cargo value in USD for tax calculations (default: 10000000.0).
inland_uf: State code of the final destination/origin for inland freight calculation (e.g. MT, GO, PR).
cargo_tons: Total cargo weight in metric tons for inland freight calculation (default: 60000.0).
| Name | Required | Description | Default |
|---|---|---|---|
| commodity | No | FERTILIZANTES | |
| inland_uf | No | MT | |
| cargo_tons | No | ||
| cargo_value_usd | No | ||
| intended_port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool is a decision/evaluation operation and that it returns a FiscalRoutingResponse with alternative ports, demurrage vs tax costs, and a recommendation. It does not state side effects, permissions/auth requirements, or cost/latency characteristics, leaving real gaps for a computation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and mechanism in two tight paragraphs before the args block. The args list is long but earns its place given zero schema-level documentation; only slight redundancy between the prose and the Args block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not detail return values, and it sensibly points to the FiscalRoutingResponse contents instead. Parameters are fully covered. The only shortfall is the absence of annotation-equivalent behavioral context (side effects, permissions), which is more critical when annotations are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does fully: it documents all five parameters with formats and concrete examples (UN/LOCODE 'BRSSZ', commodities 'FERTILIZANTES/SOJA', UF codes 'MT, GO, PR'), plus units and default values for cargo_value_usd and cargo_tons.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (evaluate) and resource (fiscal/logistical arbitrage across alternative ports), and explains the mechanism: cross-referencing congestion delay penalties with regional ICMS tax burdens. However, it never distinguishes itself from the very close sibling evaluate_routing_alternatives, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '[DECISION TOOL]' tag implies this is decision-support rather than raw data retrieval, and naming the arbitrage scenario gives an implicit context for use. But there is no explicit when-to-use, no exclusion, and no routing to alternative siblings like evaluate_routing_alternatives or evaluate_corridor_risk.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_routing_alternativesAInspect
[DECISION TOOL] Evaluate and compare physical logistics conditions between two ports.
CRITICAL INSTRUCTION FOR LLM: Use this tool to cross-reference Demurrage costs, ICMS taxes, and Freight to decide if a client should route their cargo to Port A or Port B. Highly recommended for cost-saving queries.
Returns a DecisionResult (decision-result.v1) with:
- comparison.delta_delay_days: estimated delay difference
- comparison.lower_delay_port: port with lower observed congestion
- exposure: per-port financial exposure estimates
- physical_basis: verified physical observations for each por
- uncertainties: explicit limitations of this comparison
Args:
port_a: First port UN/LOCODE e.g. "BRPNG".
port_b: Second port UN/LOCODE e.g. "BRSSZ".
commodity: Commodity type e.g. "SOJA".
| Name | Required | Description | Default |
|---|---|---|---|
| port_a | Yes | ||
| port_b | Yes | ||
| commodity | No | SOJA |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a fair amount: it returns a DecisionResult with comparison deltas, exposure estimates, physical basis, and explicit uncertainties. What is missing is any statement about side-effect profile, data freshness, or prerequisites, but for an evaluation tool the implied read-only nature is reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content is front-loaded with the decision-tool framing and purpose, followed by return fields and Args. The return bullet list is long, and the 'CRITICAL INSTRUCTION FOR LLM' phrasing is slightly heavy, but almost every sentence adds information an agent can use.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description need not detail return values, yet it still summarizes the DecisionResult and its uncertainty field, which aids interpretation. Combined with the Args examples, it covers the essentials for a 3-parameter, annotation-free decision tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and its Args block does: port_a and port_b are identified as UN/LOCODEs with concrete examples ('BRPNG', 'BRSSZ'), and commodity is given an example ('SOJA'). It could go further on LOCODE formatting rules, but this is well above baseline for an undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: evaluate/compare physical logistics conditions between two ports, cross-referencing demurrage, ICMS taxes, and freight to choose a routing. It is clearly distinguishable from siblings like evaluate_fiscal_routing and evaluate_corridor_risk because it is framed as a two-port routing decision tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage trigger ('Use this tool to ... decide if a client should route their cargo to Port A or Port B') and a positive recommendation for cost-saving queries. However, it does not say when NOT to use it or name competing tools such as evaluate_fiscal_routing, which an agent might confuse with this one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_scdew_warningCInspect
[DEPRECATED: Use evaluate_end_to_end_supply_chain_risk instead] Supply Chain Disruption Early Warning.
| Name | Required | Description | Default |
|---|---|---|---|
| origin_port | No | BRPNG | |
| chokepoint_id | No | HORMUZ | |
| destination_port | No | CNTAO |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It offers nothing about what the tool does behaviorally—no read-vs-write indication, no side effects, no return format, no prerequisites. A deprecation tag is not behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is very short and front-loads the deprecation notice, which is the most important information. But the brevity comes from omission rather than efficiency—there is simply almost no content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a domain tool with 3 undocumented parameters, no annotations, and no output schema, the description is far too thin. The one useful element (deprecation pointer) is present, but everything else an agent needs to call this legacy tool correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 3 parameters (origin_port, chokepoint_id, destination_port). The description does not mention any parameter, so it fails to compensate for the schema's total lack of documentation. Defaults exist but their semantics are unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is essentially just a title restated: 'Supply Chain Disruption Early Warning.' It does not state a verb+resource, what the tool computes, or what an 'early warning' evaluation actually returns. The deprecation notice points to a successor but does not otherwise clarify purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The deprecation tag explicitly steers the agent to evaluate_end_to_end_supply_chain_risk, which is useful routing guidance. However, it gives no guidance on when this legacy tool would still be appropriate, nor any exclusions or context beyond 'use the other one.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_vessel_queue_delaysAInspect
[INFERENCE TOOL] Vessel Queue Predictive Model (VQPM) for t+1 to t+7.
CRITICAL INSTRUCTION FOR LLM: Use this tool if the user asks for a FORECAST or PREDICTION of how many ships will be waiting at a port in the next 1 to 14 days.
VQPM_{t+1} = α × VQ_t + β × PCI_t + γ × CDR_t + δ × Seasonality.
Args:
port_id: UN/LOCODE e.g. "BRSSZ" (Santos), "BRPNG" (Paranaguá).
forecast_horizon_days: Horizon in days (1 to 7, default: 1).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes | ||
| forecast_horizon_days | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that this is an inference/predictive model rather than a deterministic read, but it omits uncertainty/confidence behavior, validation of the UN/LOCODE, and failure modes, and the description is internally inconsistent on the horizon range (1-7 vs 1-14).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with a category tag and a use-trigger, which is good, but the VQPM formula and the redundant 'Args:' block largely restate the schema and add little selection/invocation value, and the horizon figure conflicts across sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the two parameters are explained. Still missing for a 0%-coverage, zero-annotation tool: sibling disambiguation and a self-consistent horizon range, both of which an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does provide the port_id format (UN/LOCODE) with concrete examples and the horizon default. However, it contradicts itself on the horizon range, stating '1 to 14 days' in the instruction and '1 to 7' in the args, which could lead an agent to pass an out-of-range value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (forecast/predict) and resource (vessel queue at a port) with a temporal scope, so the agent knows this is a predictive tool rather than a data lookup. However, it never differentiates itself from the sibling 'predict_vessel_queue', which appears to be a near-duplicate, leaving a real ambiguity unresolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger: 'Use this tool if the user asks for a FORECAST or PREDICTION of how many ships will be waiting at a port.' That is clear context, not merely implied. It offers no exclusions and does not name the competing predict_vessel_queue sibling, so it stops short of the 'when-not/alternatives' bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cdr_riskAInspect
[DEPRECATED: Use evaluate_chokepoint_disruption instead] Calculate Chokepoint Disruption Risk.
| Name | Required | Description | Default |
|---|---|---|---|
| chokepoint_id | No | HORMUZ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the critical behavioral trait that the tool is deprecated, which is valuable, but says nothing about side effects, auth needs, or whether the tool still functions vs. errors out. 'Calculate' implies a read, but that is inference rather than disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses with the deprecation warning front-loaded before the tool's function. Nothing is wasted and the most important information appears first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a deprecated single-parameter tool the deprecation pointer is the essential content and it is present, but the undocumented parameter and absent behavioral details (does it still return a value? does it error?) leave gaps for an agent deciding whether to call it at all.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter (chokepoint_id, default 'HORMUZ') is never mentioned in the description. No format, accepted values, or meaning of the default is added, so the agent gets nothing beyond the bare schema field name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Calculate Chokepoint Disruption Risk') and immediately routes the agent to the successor tool, so it is distinguishable from the evaluate_* siblings. The abbreviation 'CDR' is only expanded inside the sentence, but the meaning is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Use evaluate_chokepoint_disruption instead'), which is the single most useful guidance this description can give. It stops short of spelling out any remaining condition under which the deprecated tool should still be called, but the intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_inland_logistics_bottlenecksAInspect
[INFERENCE TOOL] Calculate Intermodal Rail Delay Index (IRDI, 0-100 score).
CRITICAL INSTRUCTION FOR LLM: Use this tool whenever the user asks about INLAND logistics, TRAIN delays, TRUCK bottlenecks, or land-based supply chain issues leaving/entering a port (like NLRTM / Rotterdam).
IRDI = (Avg Delay × 0.4) + (Delays % × 0.3) + (Timetables × 0.2) + (Rolling Stock × 0.1).
Args:
port_or_corridor_id: UN/LOCODE e.g. "NLRTM" (Rotterdam), "DEHAM" (Hamburg).
| Name | Required | Description | Default |
|---|---|---|---|
| port_or_corridor_id | No | NLRTM |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It helpfully labels itself an '[INFERENCE TOOL]' and publishes the exact IRDI formula, which is real transparency about how the score is computed. It says nothing about data sources, freshness, or update frequency of the underlying delay inputs, which is a meaningful gap for an inference result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and score range are front-loaded in the first sentence, followed by usage triggers, then the formula, then the argument. Everything is compact and each block earns its place; the 'CRITICAL INSTRUCTION FOR LLM' phrasing is loud but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the single parameter is documented with examples. The formula is provided so the agent understands the index semantics. The one omission is guidance on valid inputs versus the 'list_supported_ports' sibling, but overall it is sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it explains the parameter is a UN/LOCODE and supplies concrete examples ('NLRTM' Rotterdam, 'DEHAM' Hamburg). It also implies the field accepts corridor IDs, not just ports. It does not explain what happens when the default is used, but the coverage is otherwise good.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Calculate') and a named resource ('Intermodal Rail Delay Index (IRDI, 0-100 score)'), so the agent knows exactly what it returns. It is weakened slightly by the presence of the sibling 'get_irdi_index', which computes the same acronym and is never distinguished from this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit triggering conditions ('whenever the user asks about INLAND logistics, TRAIN delays, TRUCK bottlenecks, or land-based supply chain issues leaving/entering a port'). However it offers no exclusions and never routes the agent away from the near-identical 'get_irdi_index' sibling, so the when-not case is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_irdi_indexAInspect
[DEPRECATED: Use get_inland_logistics_bottlenecks instead] Calculate Intermodal Rail Delay Index.
| Name | Required | Description | Default |
|---|---|---|---|
| port_or_corridor_id | No | NLRTM |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden, and it discloses only deprecation status. It says nothing about permissions, rate limits, response shape, or what the computed index represents – thin for a tool with zero structured coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed clauses, with the deprecation warning front-loaded so an agent sees the redirect before anything else. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The deprecation redirect is the most important context and is present, which is adequate for steering an agent away. However, with no output schema and no parameter documentation, the description would be insufficient if the tool were actually invoked.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter port_or_corridor_id, and the description adds no meaning about valid port/corridor identifiers or the significance of the NLRTM default. The parameter name is somewhat self-describing, keeping this above a 1.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Calculate) and resource (Intermodal Rail Delay Index), plus a deprecation banner that names the replacement tool. An agent can identify what it does and that it is superseded without opening the schema, though the acronym 'IRDI' is only expanded in the description, not differentiated from sibling index tools like get_pci_index.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly declares the tool deprecated and routes the agent to get_inland_logistics_bottlenecks instead, which is the correct when-to-use guidance. This is the strongest form of alternative-routing a description can provide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pci_indexBInspect
[DEPRECATED: Use get_port_congestion_risk instead] Calculate Port Congestion Index.
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses only deprecation status. It does not say whether the tool still functions, whether calls will error or silently return stale values, whether removal is scheduled, or what credentials/rate limits apply. The deprecation notice is genuinely useful context but is the only behavioral signal present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the deprecation warning ahead of the functional description, which is the correct priority for a retired tool. The trailing "Calculate Port Congestion Index" largely restates the tool name, making it slightly redundant but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 1-parameter tool with no output schema, the only thing an agent truly needs is the replacement and whether this tool is still callable — only the first is answered. Missing are parameter format guidance, expected return shape, and the operational status of the deprecated endpoint, leaving an agent unable to decide whether to call it at all.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter port_id has 0% schema description coverage and the description says nothing about it — no format, no valid-value source (list_supported_ports is a sibling that likely supplies IDs), no required-ness context. Since coverage is below 50%, the description needed to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Calculate Port Congestion Index") and immediately flags deprecation, which distinguishes it from the active sibling get_port_congestion_risk. It does not explain what the index measures or how it relates to the replacement, so the agent knows what it computes only at surface level.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent away from this tool: "Use get_port_congestion_risk instead" names the alternative directly. It stops short of a full when/when-not treatment — no indication of whether this tool still returns valid data, or whether the replacement is an exact equivalent, so an agent cannot judge fallback behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_physical_eventsAInspect
[OBSERVATION TOOL] Return temporal physical events for a port as a ChangePacket (physical-event.v1).
Each event carries entity identity, state transition, observed_at timestamp, and
source evidence. Use this tool to understand WHAT changed and WHEN.
Args:
port_id: UN/LOCODE e.g. "BRPNG" (Paranaguá), "BRSSZ" (Santos).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden and does well: it labels itself an 'OBSERVATION TOOL,' states the return type, and lists event fields such as entity identity, state transition, observed_at timestamp, and source evidence. It does not mention potential error conditions or limitations, but the observation nature is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized: a front-loaded tool type, a precise main sentence, a brief field explanation, and an Args section with examples. Every sentence adds useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single parameter, no annotations, and an existing output schema, the description is largely complete: it explains what is returned, the key event fields, the input format, and the intended use. It could more explicitly route between siblings, but the output schema covers return-value details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for the bare 'port_id' string property. It does so by specifying the format (UN/LOCODE) and giving concrete examples like 'BRPNG' (Paranaguá) and 'BRSSZ' (Santos).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Return temporal physical events for a port as a ChangePacket (physical-event.v1).' It also clarifies the semantic focus, 'WHAT changed and WHEN,' which distinguishes it from state/trend tools like get_port_state and get_port_trend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool: 'Use this tool to understand WHAT changed and WHEN.' It does not explicitly name excluded sibling tools or contrast them, but the intent is clear and an agent can infer it is for event/history data rather than current state or trend.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_port_congestion_riskAInspect
[INFERENCE TOOL] Calculate Port Congestion Index (PCI, 0-100 composite score).
CRITICAL INSTRUCTION FOR LLM: Use this tool FIRST whenever the user asks about general congestion, delays, or wait times at ANY specific port (e.g., SGSIN, BRSSZ). Do not guess delays; call this tool.
PCI = (Congestion Level × 0.4) + (Avg Delay × 0.3) + (Vessel Queue × 0.2) + (Berth Use × 0.1).
Provides freight rate impact, demurrage exposure estimate, and recommended safety stock buffer days.
Args:
port_id: UN/LOCODE e.g. "BRSSZ" (Santos), "SGSIN" (Singapore), "NLRTM" (Rotterdam).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it labels the tool as an INFERENCE TOOL, exposes the PCI formula and weights, and states that it provides freight rate impact, demurrage exposure, and safety stock buffer days. It stops short of explicitly declaring read-only status or data freshness/limitations, so a 5 is not warranted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and the critical usage instruction, then adds the formula and output context. The formula and output list are relevant, but the all-caps instruction and formula make it longer than strictly necessary; it remains readable and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a composite risk inference tool with no annotations and one required parameter, the description covers purpose, usage, calculation logic, expected output context, and parameter format. An output schema exists, so return-value details need not be exhaustively described, and nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 0% description coverage for the single port_id parameter, so the description must compensate fully. It does: it defines the format as UN/LOCODE and gives concrete examples such as BRSSZ, SGSIN, and NLRTM.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Calculate) and resource (Port Congestion Index, PCI) with its 0-100 composite scale. It does not explicitly differentiate from siblings such as get_pci_index or get_port_risk, but it does scope the purpose to general congestion, delays, and wait times at a specific port.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a strong, explicit usage directive: use this tool FIRST whenever the user asks about general congestion, delays, or wait times at any specific port, and do not guess delays. It lacks named alternatives or when-not-to-use guidance, which would be needed for a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_port_operations_statusAInspect
[OBSERVATION TOOL] Plain-language status of port operations: is it delayed, congested or normal?
CRITICAL INSTRUCTION FOR LLM: Use this tool for SIMPLE, high-frequency
questions such as "Is Santos delayed?", "How many ships are waiting a
Paranaguá?", "What is the ETA delay risk at this port?", "Where is my cargo
stuck?". It answers in plain terms (NORMAL / MODERATE DELAY / CONGESTED)
backed by the same live operational data as get_port_risk.
This tool is FREE (observation layer). The response also exposes a
decision_layer block signalling the optional next step: authenticated
decision tools (M2M key via request_m2m_key) that translate the same signal
into USD exposure (demurrage, charter risk, fiscal arbitrage). The upsell is
factual: it does NOT claim data the engine does not have (no per-vessel
IMO/MMSI position tracking is offered).
Args:
port_id: UN/LOCODE e.g. "BRSSZ" (Santos), "BRPNG" (Paranaguá), "NLRTM" (Rotterdam).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the cost model ('FREE, observation layer'), the response shape (plain terms plus a decision_layer block), and an honest limitation ('no per-vessel IMO/MMSI position tracking is offered'). It stops short of describing rate limits or failure behavior, but for an observation tool this is a rich disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the sections are scannable, but the middle block leans promotional ('The upsell is factual') and spends several sentences justifying the decision-layer monetization rather than helping the agent invoke the tool. That material is informative but not tightly earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the single required parameter is fully specified with format and examples. With no annotations, the description supplies the behavioral context itself. Nothing essential for correct invocation is missing, though pagination or latency characteristics are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema itself only says 'Port Id' as a string, so the description must compensate — and it does, specifying the UN/LOCODE format with three worked examples ('BRSSZ' Santos, 'BRPNG' Paranaguá, 'NLRTM' Rotterdam). This is exactly the syntax an agent needs to construct a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Plain-language status of port operations') and immediately names the outcome vocabulary (NORMAL / MODERATE DELAY / CONGESTED). It also positions itself against the sibling get_port_risk ('backed by the same live operational data as get_port_risk'), so an agent can distinguish the observation layer from the risk-scoring tool without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions with concrete example questions ('Is Santos delayed?', 'Where is my cargo stuck?') and routes heavier dollar-exposure questions toward authenticated decision tools. The only gap is that it never states a hard 'do not use for X' exclusion; the alternative routing is implied through the upsell framing rather than declared as exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_port_riskAInspect
Assess CURRENT congestion and delay risk at a single seaport.
Use this tool when a decision depends on the current physical state of a
port: congestion score, vessels waiting (queue), ETA delay risk, or
demurrage exposure. Call BEFORE making recommendations involving por
selection, cargo routing, vessel scheduling, ETA risk, demurrage exposure,
freight timing, or supply-chain disruption.
Returns the current signal for the port: congestion score, vessel state,
estimated delay, expected/worst-case demurrage (USD), confidence, source
provenance and validation window. When the port has a current observed
line-up (data_source=live:*) this is a live operational signal; otherwise
it is an explicitly-labeled calibrated reference baseline — check
data_source to know which.
Args:
port_id: UN/LOCODE of the port, e.g. "BRSSZ" (Santos), "BRPNG" (Paranaguá), "CNSHA" (Shanghai).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses the return signal set, and critically explains the data_source=live:* vs. calibrated-reference-baseline distinction so the agent knows how much to trust the output. This is exactly the provenance/auth-adjacent context an agent needs and cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, then return shape, then args — a sensible order with no filler sentences. Slightly verbose in the repeated list of decision contexts, but every block earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be restated, yet the description still usefully frames them and adds the trust-level caveat. Combined with full usage and parameter guidance, an agent has everything needed to call and interpret this correctly; minor room remains around failure/coverage cases (e.g., unsupported ports).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter is bare, but the description fully compensates: it names port_id, specifies the UN/LOCODE format, and supplies three concrete examples (BRSSZ, BRPNG, CNSHA). The agent can construct a valid call without guessing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Assess CURRENT congestion and delay risk at a single seaport.' The 'single seaport' framing distinguishes it from plural siblings like get_ports_risk, and the enumerated signal set (congestion score, queue, ETA delay, demurrage) pins down exactly what it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('when a decision depends on the current physical state of a port') and enumerates the downstream decisions that should call it first (port selection, cargo routing, vessel scheduling, ETA risk, demurrage, freight timing). Nothing about invocation context is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_ports_riskAInspect
Compare CURRENT congestion across several seaports in a single call.
Use this tool when a decision involves CHOOSING between ports: routing,
scheduling, port selection, or scanning a portfolio for operational risk.
Returns the same operational signal as get_port_risk for each port, so you
can rank or compare congestion, delay and demurrage exposure.
Args:
port_ids: list of UN/LOCODEs to compare, e.g. ["BRSSZ", "BRPNG", "CNSHA"].
| Name | Required | Description | Default |
|---|---|---|---|
| port_ids | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It usefully notes that the tool returns the same operational signal as get_port_risk for each port and that the result supports ranking congestion, delay, and demurrage exposure. However, it does not disclose data freshness, rate limits, failure behavior for invalid port codes, or whether partial results are returned if one port fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, immediately followed by when-to-use guidance, then the parameter details with an example. Every sentence earns its place and the format is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema available, the description is complete enough for correct selection and invocation. It explains the purpose, the use case, the parameter format, and the relationship to the singular sibling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameter. It does so by specifying that port_ids is a list of UN/LOCODEs and giving concrete examples like ['BRSSZ', 'BRPNG', 'CNSHA']. This adds format and usage meaning beyond the bare array-of-strings schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Compare CURRENT congestion across several seaports in a single call.' It explicitly differentiates from the singular get_port_risk and the trend-focused sibling by emphasizing multi-port comparison and current congestion signal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear usage context is provided: use when choosing between ports for routing, scheduling, port selection, or portfolio scanning. It names the singular get_port_risk as the comparative baseline, but does not explicitly mention when to prefer get_port_trend or list_supported_ports.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_port_stateAInspect
[OBSERVATION TOOL] Return current verified multimodal physical state of a port.
Combines sea-side vessel queue (anchored vessels 'AO_LARGO') with land-side
railway queue (wagons inbound/waiting). Returns sources for full provenance.
Args:
port_id: UN/LOCODE e.g. "BRPNG" (Paranaguá), "BRSSZ" (Santos).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral burden. It labels the tool as an observation (implying read-only), states the data is 'current verified', and promises provenance sources in the response. It stops short of detailing latency, refresh behavior, or data limitations, but those are minor for an observation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly organized into a classification prefix, a short behavioral statement, and a parameter note. Information is front-loaded, no sentence is filler, and the examples are placed exactly where they add value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-required-parameter tool with an output schema, the description covers what the tool returns, how the state is composed, and how to encode the argument. Any remaining need for a supported-port list is a natural fit for the sibling list_supported_ports.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only exposes a bare 'port_id' field. The description compensates fully by specifying the value format (UN/LOCODE) and giving concrete examples such as 'BRPNG' and 'BRSSZ', which is exactly what an agent needs to construct a valid argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with an explicit tool-class marker ('OBSERVATION TOOL') and states a specific verb and resource: 'Return current verified multimodal physical state of a port.' It then clarifies what that state consists of (sea-side vessel queue and land-side railway queue), which distinguishes it from risk, trend, and event siblings even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The bracketed 'OBSERVATION TOOL' plus 'current verified' gives a clear context: use this when you need a current physical-state snapshot of a port. It does not explicitly name alternatives or state when not to use it, so routing guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_port_trendAInspect
Get the short-horizon 24/48/72h congestion projection for a port.
Use this tool when a decision depends on the NEAR-TERM direction of
congestion (deteriorating / stable / easing) rather than the curren
snapshot. Complements get_port_risk. This is a SYNTHETIC projection,
not a live forecast.
Args:
port_id: UN/LOCODE of the port, e.g. "BRSSZ" (Santos), "BRPNG" (Paranaguá).
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers the key caveat that this is a SYNTHETIC projection, not a live forecast — critical for trusting the output. It does not cover auth requirements, rate limits, or pagination, but the synthetic-data disclosure is the most consequential trait for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then usage, then the synthetic caveat, then args — each sentence adds distinct information with no filler. The Args block is compact and the examples earn their space by clarifying the ID format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description covers purpose, selection criteria, data provenance, and param format. Nothing an agent needs to invoke it correctly is missing, and return values are appropriately left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single param is undocumented in the schema, so the description must compensate — and it does, specifying the UN/LOCODE format with two concrete examples (BRSSZ Santos, BRPNG Paranaguá). That is more than the schema provides, though it omits invalid-input behavior or case sensitivity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (short-horizon 24/48/72h congestion projection for a port), including the exact horizon windows. It explicitly differentiates from siblings by naming get_port_risk as a complement and contrasting near-term direction against the 'current snapshot' tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('when a decision depends on the NEAR-TERM direction of congestion... rather than the current snapshot') and names the alternative it complements (get_port_risk). The deteriorating/stable/easing enumeration tells the agent what output dimension drives selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_public_port_feedAInspect
[FREE / NO AUTH] Get the live public port-delay feed for all covered ports.
This is the recommended FIRST CALL for any agent exploring Aether-X.
Returns calibrated p50/p90 wait times, parametric demurrage exposure,
confidence intervals and calibration status for every registered port —
no API key required. Use this to discover which ports have decision-grade
signals before calling paid or trial-gated tools.
Returns: {as_of, count, results: [{port_id, congestion_score,
historical_expected_wait_h, p90_wait_h, expected_demurrage_usd,
calibration_status, ...}]}
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: no auth required, free, and the calibration status of returned data ('decision-grade signals'). However, it does not state rate limits, pagination behavior, or whether the feed is eventually consistent. For a read-only public feed, these are minor gaps, but the description does an above-average job of conveying the tool's operational profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the most critical information (free, no auth, live feed) in the first sentence. It then efficiently layers the recommendation, the returned data types, and the return schema. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description explicitly enumerates the returned fields (p50/p90 wait times, demurrage exposure, confidence intervals, calibration status) and provides a concrete return example. Combined with the clear usage guidance, this gives an agent everything needed to invoke the tool correctly. No annotations are present, but the description compensates fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description appropriately does not discuss parameters and instead focuses on return values and usage context. No penalty applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (live public port-delay feed) with explicit scope ('all covered ports'). It further distinguishes itself by declaring it is the recommended 'FIRST CALL' and contrasts with 'paid or trial-gated tools'. An agent can immediately tell this apart from siblings like get_port_congestion_risk or get_port_operations_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool ('recommended FIRST CALL for any agent exploring Aether-X', 'to discover which ports have decision-grade signals before calling paid or trial-gated tools'). It names the alternative category (paid/trial-gated tools) and the condition that selects them, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_shipment_reconstructionAInspect
Recover a shipment reconstruction with its full evidence chain (ledger-persisted).
Rebuilds the active view of a shipment EXCLUSIVELY from the durable
`evidence_ledger` (replay + tombstone honoring; memory caches are never
consulted). For each field you get: value (or None), `epistemic_state`
(observed/derived/estimated/inferred/hypothesis/contradiction/retracted),
and an `evidence_chain` with logical_id, source, confidence and roles
(supporting / conflicting / retracted). Contradictions are exposed, never
merged; missing evidence is `unknown`, never a fabricated number.
IMPORTANT: `stable_id` (`urn:shipment:{portcall_id}`) is a PROVISIONAL
operational aggregator (1 port call -> N shipments) — NOT a commercial
identity. `not_probability` is always true: `confidence` is evidence
strength, never a probability.
Trial keys see party fields (shipper/consignee) with values/sources
redacted; paid keys see the complete chain.
Args:
stable_id: e.g. "urn:shipment:shp_urn:portcall:BRSSZ:unknown".
| Name | Required | Description | Default |
|---|---|---|---|
| stable_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so richly: replay + tombstone honoring, memory caches never consulted, contradiction-exposing semantics, no fabricated numbers, and trial-vs-paid key redaction of party fields. This is exactly the behavioral context an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and constraints, then detail. The prose is dense but each paragraph adds operative information; the trailing 'Args:' line lightly restates the single parameter but is acceptable. Not bloated, though slightly longer than strictly necessary to reach the same clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not re-explain return shape, and it still covers the epistemic state vocabulary, evidence chain structure, contradiction handling, and access-tier differences. Nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter is undocumented in the schema, but the description compensates with the urn format example and an explicit warning that stable_id is a provisional operational aggregator (1 port call -> N shipments), not a commercial identity. It stops short of stating how to obtain a valid stable_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a precise verb and resource ('Recover a shipment reconstruction with its full evidence chain') and scopes it as ledger-persisted. It is clearly distinct from the list-oriented sibling list_active_reconstructions and from the risk/evaluation siblings. An agent knows exactly what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (recover/rebuild a single shipment's active view from the ledger) but the description never states when to prefer this over list_active_reconstructions or any other sibling, nor any preconditions beyond the key-tier note. Adequate but no explicit when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_active_reconstructionsAInspect
List durable shipment reconstructions with an epistemic summary per field.
Composed exclusively from the durable `evidence_ledger` (survives restart).
Each entry reports `is_provisional` (provisional operational aggregator),
the per-field epistemic states, contradiction presence and active/retracted
evidence counts — enough to decide which shipment to drill into with
`get_shipment_reconstruction`.
Args:
port_id: Optional scope filter (UN/LOCODE), e.g. "BRSSZ".
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the data source (durable `evidence_ledger`, survives restart), the distinction between durable and provisional aggregators, and the fields each entry reports (is_provisional, per-field epistemic states, contradiction presence, active/retracted counts). It omits pagination, auth, or result-size behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then behavior, then param detail. Every sentence contributes; the Args section is slightly formal but not wasteful. Appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be spelled out, yet the description still summarizes entry contents usefully. For a read-only list tool with minimal params this is close to complete, missing only operational details like pagination or limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lone parameter — and it does: it names the parameter, marks it optional, specifies the UN/LOCODE format, and gives a concrete example ("BRSSZ"). It adds real meaning beyond the bare schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and resource ("durable shipment reconstructions") with a clear scope qualifier ("with an epistemic summary per field"), and explicitly names the sibling get_shipment_reconstruction as the drill-down counterpart, letting an agent distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: this tool is for deciding which shipment to drill into, and get_shipment_reconstruction is the follow-up. Clear context but no explicit exclusions or prerequisites, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_supported_portsDInspect
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_vessel_queueCInspect
[DEPRECATED: Use forecast_vessel_queue_delays instead] Vessel Queue Predictive Model.
| Name | Required | Description | Default |
|---|---|---|---|
| port_id | Yes | ||
| forecast_horizon_days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It says nothing about whether this is a read-only prediction, whether the model is deterministic, latency, rate limits, or what the forecast output looks like. 'Predictive Model' conveys that it computes a forecast but adds no operational detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short fragments with the deprecation warning front-loaded, which is the right emphasis. No waste, though the trailing 'Vessel Queue Predictive Model' is close to a name restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 2-parameter tool with zero schema coverage, no annotations, and no output schema needs the description to explain parameter formats and the nature of the returned forecast. The description only flags deprecation, leaving the agent able to route away but unable to call the tool correctly if it must.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds nothing about either parameter. An agent cannot learn from the definition what port_id format is expected (name, UN/LOCODE, numeric id) or what unit/limits forecast_horizon_days accepts beyond the default of 1.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it is a 'Vessel Queue Predictive Model' — a specific verb (predict) plus resource (vessel queue). However it does not describe what is actually predicted (wait times, queue length, arrival delays), so the precise output is unclear. The deprecation banner and pointer to the successor tool are the most useful part.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to use forecast_vessel_queue_delays instead — that is a strong, unambiguous routing instruction naming the alternative. It lacks any positive statement of when this tool would still be used, but for a deprecated tool the 'don't use this' guidance is the dominant need.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_m2m_keyAInspect
[M2M SELF-SERVE TOOL] Request an instant 30-day trial M2M API Key.
Use this tool if a Decision Tool returns an Access Denied / 403 error.
Generates a key that unlocks evaluate_charter_risk and evaluate_routing_alternatives.
Args:
agent_name: Name or ID of your AI agent or client.
organization: Organization or trading desk name.
contact_email: Contact email.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_name | Yes | ||
| organization | Yes | ||
| contact_email | No | agent@m2m.local |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: instant issuance, a 30-day trial duration, and that the credential unlocks evaluate_charter_risk and evaluate_routing_alternatives. It omits whether the key replaces an existing one, rate limits, or how the key is returned, but the output schema exists so return values need not be spelled out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the capability and the trigger condition in the first two lines, followed by a compact Args list. The Args block partially repeats schema field names, but given 0% schema coverage that repetition is functional rather than wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter credential-issuance tool with an output schema already defining the return shape, the description covers what it is, when to call it, and what it enables. The remaining gap is failure/edge-case behavior (duplicate keys, expiry handling), which is minor for this tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does document all three parameters with purpose statements ('Name or ID of your AI agent or client', 'Organization or trading desk name', 'Contact email'). The explanations are thin and give no format or validation expectations, so it is not a full substitute for a documented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Request an instant 30-day trial M2M API Key') and identifies itself as a self-serve credential tool, which no sibling tool is. An agent can immediately distinguish this from the evaluate_*/get_* risk tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Use this tool if a Decision Tool returns an Access Denied / 403 error,' and names the two tools the key unlocks. It does not state any exclusion (e.g., only one key per org, or what to do if a key already exists), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
get_public_port_feed
2 tool updates
- Added
get_shipment_reconstruction - Added
list_active_reconstructions
1 tool update
- Added
assess_logistics_disruption
1 tool update
- Added
get_port_operations_status
10 tool updates
- Added
evaluate_chokepoint_disruption - Added
evaluate_end_to_end_supply_chain_risk - Changed
evaluate_scdew_warning1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "evaluate_scdew_warningDictOutput", - "type": "object" -}New value: +null
- Added
forecast_vessel_queue_delays - Changed
get_cdr_risk1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "get_cdr_riskDictOutput", - "type": "object" -}New value: +null
- Added
get_inland_logistics_bottlenecks - Changed
get_irdi_index1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "get_irdi_indexDictOutput", - "type": "object" -}New value: +null
- Changed
get_pci_index1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "get_pci_indexDictOutput", - "type": "object" -}New value: +null
- Added
get_port_congestion_risk - Changed
predict_vessel_queue1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "predict_vessel_queueDictOutput", - "type": "object" -}New value: +null
1 tool update
- Added
evaluate_fiscal_routing
5 tool updates
- Added
evaluate_scdew_warning - Added
get_cdr_risk - Added
get_irdi_index - Added
get_pci_index - Added
predict_vessel_queue
1 tool update
- Added
evaluate_corridor_risk
1 tool update
- Added
request_m2m_key
4 tool updates
- Added
evaluate_charter_risk - Added
evaluate_routing_alternatives - Added
get_physical_events - Added
get_port_state
Related MCP Connectors
Ocean & multimodal freight intelligence: rates, landed cost, transit, customs, risk, ship decisions
WMS & logistics intelligence: live freight & shipping rates, port data, inventory, fleet, KPIs
Ocean shipping intelligence: D&D, freight rates, vessel schedules, port data. 24 tools, 6 carriers.
- VoydarOAuthcom.voydar
Maritime intelligence: vessels, tracks, port calls, company fleets and sanctions screening, sourced.
Related MCP Servers
- AlicenseAqualityDmaintenanceOcean and multimodal freight intelligence suite providing cross-validated rates, total landed cost, transit reliability, customs, risk, emissions, and unified ship decisions through 47 tools.4713 npmMIT
- FlicenseAqualityDmaintenanceReal-time supply chain risk intelligence with 25 tools: Global Disruption Index, Manufacturing Index, commodity prices, port congestion, border delays, chokepoints, air cargo, trade policy, energy, rail, freight, economic indicators, predictive signals, and AI intelligence briefs.341-
- AlicenseNot gradedqualityBmaintenanceProvides real-time, machine-readable ship and port conditions for major US container gateways using AIS data, allowing AI agents to query vessel status, berthing events, and gateway conditions without API keys or signup.Apache 2.0
- FlicenseNot gradedqualityDmaintenanceProvides container shipping intelligence for AI agents, enabling demurrage & detention calculations, local charges, inland haulage rates, and CFS tariffs across multiple shipping lines, with pay-per-request USDC payments via x402.-
Glama MCP Gateway
Add one secure layer between your agents and this server.