Skip to main content
Glama

Overwing

Server Details

Guardrails for LLM output: pass / fail / review verdicts and one recommended action. Hosted or npm.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
frod27/overwing-mcp
GitHub Stars
0
Server Listing
overwing-mcp

TDQS

A3.7/5.0

Scored across 27 tools

Disambiguation4/5

Most tools target a clearly distinct resource+action: the atlas_* trio (browse registry, resolve a UA string, public stats), the beacon_* trio (start a paid check, fetch its report, read a free sample), the rule/evaluation group, and the tower_* governance group. Only mild overlaps exist (beacon_sample vs beacon_report, get_usage vs whoami), and the descriptions do enough to separate them.

Naming Consistency4/5

Naming is consistently snake_case and verb-first, with per-product prefixes (atlas_, beacon_, tower_) that make grouping obvious. The deviation is that the evaluation/rule-set family (evaluate, get_rule_set, list_evaluations, whoami) carries no prefix while the other families do, so the pattern isn't uniform across the whole set.

Tool Count3/5

At 27 tools the surface is heavy, but it is split across four genuinely different products (Atlas registry, Beacon checks, rule-set evaluation, Tower agent governance), so the count is defensible rather than padded. Still, an agent must first work out which family it needs, and several tower_* tools could be merged.

Completeness4/5

Tower has full lifecycle coverage (load template, create/list/revoke agents, capabilities, decide, submit, poll action, compensate, receipts, verify chain) and Beacon covers start/retrieve/sample. The rule-set side is missing update_rule_set and delete_rule_set, so custom rule sets can be created but never edited or removed, which is a real gap.

Available Tools

27 tools
atlas_agentsSearch the agent registry (Overwing Atlas)A
Read-only
Inspect

Browse or search Overwing Atlas, the registry of AI crawlers, fetchers and browser agents: name, operator, user-agent tokens, Web Bot Auth key directory, robots.txt behaviour, and (with Atlas Pro or Team) purpose class, verification, evasion flags and traffic shares. Filter by free text, purpose (training, search, browser, coding), operator, or verification. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFree text over name, operator, user agents, description
limitNo
purposeNoPurpose substring, e.g. training, search, browser, fetcher, coding
operatorNo
verificationNoe.g. 'Web Bot Auth', 'spoofable', 'Unattributable'

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds genuinely useful behavioral context beyond that: an organization key is required, and several fields (purpose class, verification, evasion flags, traffic shares) are gated behind Atlas Pro or Team. It stops short of covering pagination or result caps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and returned fields are front-loaded, followed by filters and then the auth prerequisite. It is a single dense sentence with little waste, though the parenthetical Pro/Team list makes it slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a read-only annotation, the description carries the right load: it describes the browsable record fields, the filter surface, and the auth requirement. The main omission is pagination/limit behavior, which is minor for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description compensates by naming exactly the filter axes that map to q, purpose, operator and verification, including example purpose values (training, search, browser, coding). Only limit is left unexplained in prose, though the schema supplies its default and bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (browse or search) and resource (Overwing Atlas, the registry of AI crawlers, fetchers and browser agents), then enumerates the record fields. It is clearly distinguishable from siblings atlas_lookup and atlas_summary by describing registry-wide search rather than a single-entity fetch or aggregate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete filtering dimensions (free text, purpose, operator, verification) and notes the organization-key requirement, which implies when it applies. However, it never states when to prefer this over atlas_lookup or atlas_summary, so routing among the Atlas siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

atlas_lookupIdentify a user agent (Overwing Atlas)A
Read-only
Inspect

Say what a User-Agent string claims to be and whether the claim can be trusted, from the Overwing Atlas registry of AI crawlers, fetchers and browser agents. Returns the claimed agent, operator, purpose class, verification method (Web Bot Auth signature, user-agent string only, or unattributable) and a trust note. Works with no API key: 10 lookups a day. With a key, metered per day by Atlas tier: free 100, Pro 10,000, Team 100,000. Use it when deciding whether to serve, block, or pay-gate a request, or to understand who is hitting a site.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_agentYesThe User-Agent header value to identify

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries most of the operational burden — and it delivers: no-API-key access at 10 lookups/day, tier metering (free 100, Pro 10,000, Team 100,000), and the verification-method taxonomy it returns. It adds real context beyond annotations, though it doesn't explain failure modes or what happens on rate-limit exhaustion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, then return shape, then cost/usage — nothing is filler. The rate-limit enumeration is dense but genuinely decision-relevant, so it earns its space, though the paragraph is slightly heavy for a single-parameter lookup.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must describe the return — and it does: claimed agent, operator, purpose class, verification method, and trust note. With one fully documented parameter and no annotations gap, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 100% schema description coverage, so the schema already documents 'user_agent' fully. The description adds no syntax, format, or edge-case guidance beyond what the schema provides — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('say what a User-Agent string claims to be and whether the claim can be trusted') and scopes it to the Overwing Atlas registry. It is distinguishable from siblings atlas_agents and atlas_summary without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger — 'when deciding whether to serve, block, or pay-gate a request, or to understand who is hitting a site.' That is explicit usage context, but it never disambiguates against the sibling Atlas tools (atlas_agents, atlas_summary), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

atlas_summaryAgent traffic and spending summary (Overwing Atlas)A
Read-only
Inspect

Public numbers from Overwing Atlas: registry counts by purpose and verification, published browser-agent traffic shares, sector field-scan headlines (e.g. how many OSINT sites carry AI-crawler rules or any agent-payable surface), and the summary of the Agent Consumers report on what agents actually spend. No key needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, so the safety profile is covered. The description adds value beyond that by disclosing 'No key needed,' an authentication fact absent from annotations, plus that all returned data is public/aggregate — relevant context for a zero-parameter tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence beginning with the source and scope, followed by a compact enumeration of the four content categories. It is dense with parenthetical examples but no sentence is wasted; a bulleted structure would be marginally easier to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters, no output schema, and read-only semantics, the description carries the return-value burden itself and does so by listing the four data categories and giving a sample metric. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so per the rubric the baseline is 4. Nothing in the description needs to explain arguments, and it correctly provides none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates the specific content the tool returns: registry counts by purpose/verification, browser-agent traffic shares, field-scan headlines, and the Agent Consumers spending summary. This is a resource-specific read tool clearly distinct from atlas_agents and atlas_lookup, though it describes output content rather than using a crisp verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the content list — an agent can infer it should call this for public aggregate stats — and 'No key needed' signals low-friction invocation. However, there is no explicit when-to-use guidance or naming of alternatives such as atlas_lookup or atlas_agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

beacon_reportRead a reachability report (Overwing Beacon)A
Read-only
Inspect

The report for a check started with beacon_start, once it is paid for: score (0 to 100), verdict (yes, partly, no), the find / read / use answers, every check with what was found and a fix when it did not pass, and top_fixes ranked by value. Answers 402 with the checkout link while unpaid and 202 while the check is running (ask again in a few seconds). The id is the credential: whoever holds it can read the report. No key needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe check id from beacon_start

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint already covering the safety profile, the description still adds substantial behavior: the 402/202 interim states and retry guidance, and critically the security model ('the id is the credential: whoever holds it can read the report. No key needed'). That is disclosure an agent cannot get from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph, front-loaded with the report contents before the status/retry and security notes. Nearly every clause earns its place, though the run-on structure mixes payload, status codes, and credential notes without visual separation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description enumerates the returned fields (score, verdict, answers, checks, fixes, top_fixes), covers the interim 402/202 states, and states the auth model. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the pattern is documented, so the baseline would be 3, but the description adds meaning the schema lacks: the id is a bearer credential granting read access, which shapes how cautiously the agent should handle it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('the report for a check started with beacon_start') and enumerates the report's payload (score, verdict, find/read/use answers, per-check findings, top_fixes). This clearly distinguishes it from beacon_start and beacon_sample.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent this reads the output of beacon_start and gives the exact conditions for retry (402 while unpaid, 202 while running, 'ask again in a few seconds'). It does not explicitly contrast with beacon_sample, so it falls just short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

beacon_sampleSample reachability report (Overwing Beacon)A
Read-only
Inspect

A real Overwing Beacon report, free: the check of overwing.ai itself, refreshed daily. Read it to see exactly what a paid check returns before starting one. No key needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already declaring safety, the description adds genuinely useful behavioral context: no API key required, refreshed daily, and that the content is a real report rather than a mock. It does not describe pagination or size, but for a zero-param sample read this is solid added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core fact and followed by the reason to call it. Slight marketing framing ('free', 'A real ... report') but no filler that wastes an agent's budget.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read with no output schema, the description covers what the agent gets (a real sample report of overwing.ai), why to call it (preview before a paid check), and access requirements (no key). Nothing essential to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so per the baseline there is nothing for the description to document. It correctly spends no words on parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource — a real Overwing Beacon reachability report, pre-populated with the check of overwing.ai itself. It is distinguishable from the paid report/start tools, though it never names those siblings directly, leaving the reader to infer the relationship.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read it to see exactly what a paid check returns before starting one' gives a clear use context: inspect before committing to a paid run. It stops short of explicitly naming beacon_start or beacon_report as the alternative, so routing is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

beacon_startStart a reachability check (Overwing Beacon)AInspect

Overwing Beacon answers one question about a site: is this product reachable by agents? It checks robots.txt rules for AI agents, llms.txt, the sitemap, the MCP server card, hosted MCP endpoint and MCP Registry listing, the A2A agent card, OpenAPI discovery, and how the home page reads to a model, then returns three answers (find, read, use), a score and the fixes worth making. This tool starts a check paid by card ($5): it returns a checkout_url for a person to open and a check id. Nothing runs and nothing is charged until that payment completes; then call beacon_report with the id. An agent with a wallet can instead pay $1 in USDC over x402 and get the report in one call: GET /api/x402/beacon?url=. Call beacon_sample first to see a real report. No key needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe site to check: a domain (example.com) or an https URL. Public sites only.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, so an agent knows this mutates state. The description adds genuinely useful context beyond that: the $5 card price, that nothing runs or is charged until payment completes, and that the result is a checkout_url plus a check id. However, it doesn't disclose expiry of the check id, refund behavior, or rate limits, so it adds value but not rich behavioral coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and mostly front-loaded, but the opening sentences describe what the product checks before stating what the tool itself does, slightly delaying the action. Every sentence still carries real routing or pricing information, so there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining the return values (checkout_url and check id) and the follow-up call. For a one-parameter, payment-gated tool, an agent has everything needed to call it correctly and knows what to do next.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema already documents url including format and public-only constraint. The description adds no syntax or format detail beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ("starts a check") on a named product and explains exactly what the check covers (robots.txt, llms.txt, MCP card, A2A, OpenAPI, etc.). It clearly distinguishes itself from siblings beacon_report and beacon_sample by describing its role in the flow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sequencing: call beacon_sample first to see a sample, this tool to start a paid check, then beacon_report with the id once payment completes. It also names the x402 alternative with its exact endpoint for wallet-holding agents. When-to-use and alternatives are both spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_rule_setCreate rule setBInspect

Create a custom rule set with 1 to 25 rules. Each rule is a choice (pick one option), score (position on an ordered scale), or noul (yes/no) question with a fail condition, optional review threshold, weight, and action (block, redact, or review) that callers should take when it fails. Rule instructions may reference context.* fields that callers pass with each evaluation. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
slugYes
rulesYes
descriptionNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, so the description should carry the behavioral load; it does add useful domain context (the 1-25 rule bound, what a fail condition/action means, that rule instructions may reference context.* fields). However, it says nothing about slug uniqueness, conflict/error behavior on creation, or immutability of a created rule set after evaluations reference it, which matters for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences that front-load the core action and size constraint before drilling into rule anatomy. No filler, though the second sentence packs several concepts and could break more cleanly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a complex nested rules payload, the description covers the important structure (types, fail conditions, actions, context references) so an agent can construct a valid call. It stops short of describing slug semantics and creation failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does add real meaning: the three question types, the fail-condition shape per type, the optional review threshold, weight, and the block/redact/review action semantics. It does not explain name vs slug roles or the slug format constraint, leaving a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Create a custom rule set') and scopes the output ('with 1 to 25 rules'), and the domain detail about rule types makes it unmistakably the write-side counterpart to get_rule_set/list_rule_sets. It does not name siblings explicitly, but the verb alone routes correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states a prerequisite ('Needs an organization key') but gives no when-to-use/when-not guidance and never contrasts itself with evaluate, get_rule_set, or list_rule_sets. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluateEvaluate textAInspect

Works with no API key: 10 evaluations a day, inputs up to 2,000 characters, and the text is not stored. With a key: 250 a day and up, inputs up to 100,000 characters, and your own rule sets. Score any text (typically an LLM's output) against an Overwing rule set. The text can be in any language; tested in Spanish, Portuguese, French, German, Japanese, Chinese, Korean, Arabic and Hindi. Results come back in English. Returns an aggregate verdict of pass, fail, or review, a recommended_action (block, redact, review, or allow), and per-rule answers with probability, confidence, and the rule's action. Act on recommended_action: block means do not send, redact means remove the flagged content and resend, review means ask a human or a slower model, allow means proceed. Prebuilt sets: 'content-safety' (toxicity, PII, self-harm, sexual content, severity) and 'outbound-message', which also takes a context object (recipient, channel, owns_contact_info) so PII that the recipient already owns is not flagged.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesThe text to evaluate, in any language
storeNoWith a key: false runs the check without keeping the input text or the context (the verdict and metadata are still recorded). Without a key the text is never stored.
contextNoFacts the rules may reference: recipient, channel, whether you own the data, sender, purpose. Sent to the model alongside the text (max 8 KB). Use the 'outbound-message' rule set to have it honoured.
metadataNoOpaque data stored with the evaluation and echoed in webhooks (max 8 KB)
rule_setNoRule set slugcontent-safety

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (only readOnlyHint=false), so the description carries the load and does so richly: quota tiers with/without a key, input-size limits, storage behavior ('the text is not stored'), language coverage, and the exact semantics of each recommended_action. This is far beyond what the single annotation conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The tier/quota framing is front-loaded, results come next, and action semantics follow, so information is well ordered. It is dense and slightly long, but nearly every sentence conveys actionable detail rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully compensates by describing the return payload (verdict, recommended_action, per-rule probability/confidence/action) and the behavioral envelope (quotas, storage, languages). An agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains that 'store' affects retention of text and context, and that 'context' is only honoured with the 'outbound-message' rule set (e.g., PII the recipient already owns is not flagged). It stops short of explaining rule_set slugs and metadata behavior beyond what the schema shows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Score any text ... against an Overwing rule set'), clarifies the typical input (an LLM's output), and enumerates the concrete outputs (aggregate verdict, recommended_action, per-rule answers). An agent can distinguish this from siblings like evaluate_batch or get_evaluation without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear downstream guidance ('Act on recommended_action: block means do not send...') and notes that the 'outbound-message' rule set requires the context object. However, it never says when to choose this tool over evaluate_batch or other siblings, so tool-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_batchEvaluate many textsAInspect

Score up to 50 texts against one rule set in a single call. Each item counts as one evaluation. Returns a summary plus per-item verdicts; items can fail independently. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
storeNoFalse runs every item without keeping its input text or context
contextNoFacts the rules may reference: recipient, channel, whether you own the data, sender, purpose. Sent to the model alongside the text (max 8 KB). Use the 'outbound-message' rule set to have it honoured.
rule_setNocontent-safety

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false signalling a write/quota-consuming operation, the description adds genuine context beyond the annotation: per-item billing ('Each item counts as one evaluation'), partial-failure semantics ('items can fail independently'), and an auth prerequisite ('Needs an organization key'). It still omits whether evaluations persist by default (the 'store' flag), which is a meaningful behavioral trait left to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the scale and batching constraint, each carrying distinct information (scale, quota, return/failure shape, auth). No filler, though 'Returns a summary plus per-item verdicts' and 'items can fail independently' are partially redundant with each other.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly covers the return shape and partial-failure behavior, and flags the auth requirement for a 4-parameter nested-object tool. The main remaining gap is behavioral detail around the 'store' default and how batch-level versus per-item context interact, which the schema only partly explains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so half the parameters get only schema-level documentation. The description reinforces 'rule set' and the item-count ceiling, and 'organization key' is an auth note rather than a parameter (there is no such parameter in the schema). 'store' and the batch-level 'context' interplay are not addressed in prose, so it does not fully compensate for the coverage gap; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Score) and resource (texts) with an explicit scale bound (up to 50) and batching constraint ('in a single call'). This implicitly distinguishes it from the sibling 'evaluate' (single-item), though it never names that sibling, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the batch framing and the quota note ('Each item counts as one evaluation'), giving an agent a cost signal for choosing this over repeated 'evaluate' calls. However, there is no explicit when-to-use/when-not guidance or named alternative, so it remains implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evaluationGet evaluationA
Read-only
Inspect

Fetch a stored evaluation by id, including the original input and per-rule results. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so read safety is covered; the description adds an auth requirement (organization key) and describes the returned content (original input, per-rule results), which is useful context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and the returned payload, with the prerequisite stated last. Nothing is wasted or repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-resource read with no output schema and one simple parameter, the description covers what is fetched, what it returns, and the auth requirement. It could mention error behavior for an unknown id, but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter at 0% schema description coverage, so the description must carry the load. It tells the agent that id selects a stored evaluation, but adds no format, source, or lookup details beyond that minimal framing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (a stored evaluation) plus the lookup key (id) and the payload returned (original input and per-rule results). It is clearly distinguishable from list_evaluations even though it does not name that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by 'by id' — retrieve a single evaluation when you already have its identifier. A prerequisite (organization key) is given, but there is no explicit when-to-use vs alternatives guidance such as pointing to list_evaluations for browsing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rule_setGet rule setA
Read-only
Inspect

Fetch a rule set with its full rule definitions. Use 'content-safety' as a worked example when writing your own. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=true), so the bar is lower. The description adds real context beyond that: the call returns full rule definitions, requires an organization key, and can serve as a template reference. It does not describe error behavior for an unknown slug, but for a read-only single-fetch tool this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and zero padding. The 'worked example' sentence is slightly tangential to the fetch operation but still earns its place by steering authoring use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers only that full rule definitions come back — no sense of structure, size, or pagination. With one undocumented parameter and no return detail, it is adequate but leaves gaps an agent would have to discover by calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden for the single 'slug' parameter. The 'content-safety' example implicitly conveys the identifier format, which is genuine added value, but the description never names the parameter or explains what a slug identifies or where to obtain one.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Fetch a rule set with its full rule definitions.' The 'full rule definitions' phrase implies a detail-fetch that contrasts with list_rule_sets, though it never names that sibling explicitly. An agent can broadly distinguish it, but the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a usage hint — use 'content-safety' as a worked example when authoring your own rules, which implicitly links it to create_rule_set — and states a prerequisite ('Needs an organization key'). However, it never says when to call this versus list_rule_sets or what condition selects one over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usageGet usageB
Read-only
Inspect

Daily usage, remaining quota for today, and plan limits. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds a genuine behavioral constraint the annotation does not: it requires an organization key. It stops short of describing response format, quota windowing, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight clauses, front-loading the returned data ahead of the prerequisite, with no filler. Efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must carry return-value meaning, and it does list the returned fields, plus the auth requirement. But it omits the sole parameter's meaning and any response shape details, leaving an agent under-informed for the one configurable input.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'days' (integer, 1-90) has 0% schema description coverage and the description never mentions it. 'Daily usage' gestures at day granularity but does not explain what the days parameter controls, so the gap is unfilled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and enumerates the returned data: daily usage, remaining quota today, and plan limits. That is concrete and distinguishable from siblings like list_plans or whoami, though it doesn't explicitly say what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (query your quota/usage) and the auth prerequisite 'Needs an organization key' is stated, but there is no explicit when-to-use, when-not, or alternative routing versus siblings such as list_plans.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_evaluationsList evaluationsA
Read-only
Inspect

List recent evaluations, newest first, with optional verdict and rule set filters. Use next_cursor to page. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNo
verdictNo
rule_setNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint, so the description carries extra weight and does add real context: default sort order (newest first), cursor-based pagination, and an organization key requirement. The 'next_cursor' name mismatch slightly muddies the paging instruction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero padding, with the core action front-loaded and secondary details (ordering, filters, paging, auth) ordered by importance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so return shape need not be explained, but with four parameters at 0% schema coverage the description leaves limit undocumented and misnames cursor. Filtering, ordering and paging are covered, making it adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names verdict and rule_set filters and implies paging, but omits the limit parameter entirely and refers to the cursor as 'next_cursor', leaving half the parameters weakly or wrongly documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List recent evaluations') plus ordering ('newest first') and filter scope. It clearly contrasts with get_evaluation (single record) by being a listing, but never names the sibling it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the use case (browsing recent evaluations) and mentions filters and paging, but gives no explicit when-to-use vs get_evaluation or evaluate. The guidance is weakened by naming 'next_cursor' when no such parameter exists, and by claiming an organization key requirement that no parameter reflects.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_plansList plansA
Read-only
Inspect

Public plan catalog: prices, daily limits, and per-minute burst limits. No API key needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so safety is covered. The description adds genuinely new behavioral context: that the endpoint is public and requires no API key authentication, plus the dimensions of the data returned (prices, daily and burst limits). It does not describe pagination or response shape, but none is likely needed for a static catalog.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loading the resource and then the authentication fact. Every clause carries information; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, annotation-backed, output-schema-less read tool, the description supplies what an agent needs: what the catalog contains and that no credentials are required. Minor omission is any hint about response structure or size.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to explain; baseline is 4 by the rubric.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (a public plan catalog) and enumerates its contents: prices, daily limits, and per-minute burst limits. That is specific enough to tell it apart from the usage/evaluation-oriented siblings, though it leans on the tool name rather than a verb like 'list'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states 'No API key needed,' which is a useful precondition for calling it, but gives no explicit when-to-use vs. alternatives (e.g., get_usage for account-specific limits rather than catalog limits). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rule_setsList rule setsB
Read-only
Inspect

List the prebuilt and custom rule sets available to this organization. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_inactiveNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered and the bar is lower. The description adds only that results span prebuilt and custom sets for the organization; it says nothing about pagination, result volume, or ordering, which are the remaining behavioral unknowns for a list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action, with no filler. The second sentence is brief but ambiguous about how the organization key is provided, which slightly undercuts its economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no parameter documentation, and only a readOnly annotation, the description should carry more weight. It establishes the org-scoped list of prebuilt and custom rule sets but omits filter semantics and any sense of the returned collection's shape or size.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is exactly one parameter (include_inactive), yet the description never explains what it does or what the default behavior is. The only schema-adjacent claim ('Needs an organization key') refers to a field that is not present in the input schema, so it adds confusion rather than clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List the ... rule sets') and scopes it to the organization, which separates it from get_rule_set and create_rule_set well enough. It does not explicitly name a sibling or contrast its scope, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no contrast with the sibling get_rule_set (single fetch) or create_rule_set. The one prerequisite, 'Needs an organization key,' is asserted without saying how it is supplied, and no such parameter exists in the schema, leaving the agent to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_capabilitiesWhat may I do? (Overwing Tower)A
Read-only
Inspect

List the operations this agent is allowed to call, each with the JSON Schema its input must match and its compensating operation. Call this first; build inputs from the schema rather than guessing. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real behavioral context beyond that: the auth requirement ("Needs an agent key") and the shape of each returned entry (allowed operations with their input schemas and compensating operations). Return format is not otherwise documented, so this context is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, followed by the invocation instruction and the auth prerequisite. No padding or repetition; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with no output schema, the description covers purpose, invocation order, return contents, and the auth prerequisite. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly does not invent parameter guidance; the stated auth key requirement is the only input-relevant detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: list the operations this agent is allowed to call. It also specifies what each entry contains (JSON Schema plus compensating operation), which distinguishes it clearly from sibling listing tools like tower_list_agents or whoami. It does not, however, explicitly name a sibling it differs from, so it falls just short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Call this first" gives explicit ordering guidance and "build inputs from the schema rather than guessing" prescribes how to use the result, which is strong when-to-use direction. It stops short of naming alternatives or exclusions (e.g., when not to call this versus whoami).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_compensateUndo an executed action (Overwing Tower)A
Idempotent
Inspect

Run the compensating operation for an executed action, for example cancel the order that create_order made. Runs once; repeating it returns the first outcome. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
action_idYesThe executed action's id (a UUID)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and idempotentHint=true; the description reinforces and explains idempotency with 'Runs once; repeating it returns the first outcome,' and adds a real behavioral constraint not in the annotations ('Needs an agent key'). It does not, however, describe failure modes for non-compensable actions or the shape of the returned outcome.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, all front-loaded: what it does, the idempotency caveat, then the auth requirement. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation with no output schema, the definition covers purpose, idempotency and auth. It stops short of saying what the compensating operation returns or what errors to expect when no compensation exists, which an agent would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single action_id parameter is fully documented in the schema with a UUID pattern. The description adds no syntax or format detail beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('run the compensating operation') on a specific resource ('an executed action') and grounds it with a concrete example ('cancel the order that create_order made'). An agent can distinguish it from tower_submit_action, which creates actions rather than undoing them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly signals the when: undoing an already-executed action. It does not explicitly name the counterpart tool to use for the forward operation or state exclusions (e.g. what happens if the action was never executed or is not compensable), so the routing guidance is good but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_create_agentCreate an agent identity (Overwing Tower)AInspect

Create a scoped agent identity and its key. Needs an organization key. Scope it to the operations the agent needs (for example create_order) rather than * where you can. The key (ow_agent_...) is returned once, in this result, and this server does not keep it: store it, and send it as the Authorization bearer on a connection used for the agent tools (tower_capabilities, tower_decide, tower_submit_action, tower_get_action, tower_compensate, tower_get_receipt, tower_verify_receipts). Revoke with tower_revoke_agent.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesA name a person will recognise in the review queue, e.g. order-intake-bot
scopesYesOperation names this agent may call, or ["*"] for all
use_for_sessionNoIgnored here. The hosted server has no session and never holds a key; it exists so calls written for the npm package still validate.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry readOnlyHint=false. The description adds high-value behavior the annotations cannot: the key is returned exactly once and never stored server-side, its prefix (ow_agent_...), the requirement to send it as an Authorization bearer on a connection, and the set of agent tools it unlocks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then prerequisites, then key handling, then revoke path. The parenthetical list of seven agent tools is long but genuinely useful for chaining; a shorter reference would be slightly tighter without losing much.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining what the result contains (the one-time key). An agent has everything needed to create, store, and use the identity correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes beyond it by advising narrowing scopes rather than using '*', which is real guidance the schema's neutral 'Operation names this agent may call, or [*] for all' does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (create an agent identity and its key) with the 'scoped' qualifier that distinguishes it from generic identity creation. Sibling names like tower_revoke_agent and tower_list_agents make the lifecycle boundary clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit prerequisites (needs an organization key), explicit guidance on how to choose scopes ('scope it to the operations the agent needs rather than *'), and explicit routing to tower_revoke_agent for the reverse operation. When-to-use and when-to-do-something-else are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_decideWould this be allowed? (Overwing Tower)A
Read-only
Inspect

Ask Tower how it would rule on an operation without doing anything: auto (would execute), review (a person must approve), or reject. Returns the score, the reason, and each policy question's answer. No side effects. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesThe operation's input, matching its input_schema from tower_capabilities
operationYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true. The description reinforces the safety profile ('No side effects') and adds context annotations don't carry: the required credential ('Needs an agent key') and the return contents (score, reason, per-question answers).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core purpose and outcome vocabulary before the return/auth details. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully describes the return shape and states the auth requirement, covering the main gaps. The only omission is clarification of the 'operation' identifier semantics for the nested-input call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; the 'input' param is documented, but the 'operation' string param has no schema description and the description only alludes to it ('an operation'). It does not explain the operation-name format or that tower_capabilities enumerates valid names, so it barely compensates for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (ask/decide) and resource (Tower's ruling on an operation) and enumerates the exact outcomes (auto, review, reject). The phrase 'without doing anything' cleanly distinguishes it from the mutating sibling tower_submit_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The dry-run framing ('without doing anything') gives a clear use context — preview a ruling before committing an action — but no alternative is named explicitly (e.g., 'use tower_submit_action to actually execute'). Clear context, no stated exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_get_actionGet an action (Overwing Tower)A
Read-only
Inspect

Status and result of an action: pending, approved, executed, failed, compensated, or rejected. Use it to poll an action that is waiting on human review. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
action_idYesThe action's id (a UUID), from tower_submit_action

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds value beyond them with the auth requirement (\"Needs an agent key\") and the full enumeration of terminal and intermediate states, which tells the agent what outcomes to expect when polling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the return semantics, then usage, then the auth precondition. No filler and nothing repeated from annotations or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must carry the return-value burden, and it does name status plus the state space and mentions \"result.\" It is nearly complete for a simple single-id getter, missing only a hint about the result payload shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single action_id parameter is fully documented in the schema, including its UUID pattern and provenance (from tower_submit_action). The description adds nothing about parameters, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (an action) and what is returned (status and result), and enumerates the possible lifecycle states. It is clearly a read of an existing action, distinguishable from tower_submit_action and tower_decide, though it never explicitly contrasts itself with the nearby tower_get_receipt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit condition: \"Use it to poll an action that is waiting on human review.\" That tells the agent when this tool is the right one. It does not name which sibling to use instead once the action leaves the pending state, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_get_receiptGet a receipt (Overwing Tower)A
Read-only
Inspect

Fetch one signed receipt by id or by sequence number: the payload, its hash, the previous link's hash, and the Ed25519 signature. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReceipt id or sequence number

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=true, so the description's explicit 'Needs an agent key' auth requirement is real added value beyond the structured data. It also discloses that the receipt is signed and chained via the previous link's hash, giving behavioral context about what the object is; it stops short of noting failure behavior for unknown ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, followed by a tight colon-list of returned fields. No filler, no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating return fields, and it covers the auth prerequisite for a read-only single-record fetch. For a one-parameter lookup this is nearly complete; only error/not-found behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'id' parameter is already documented as 'Receipt id or sequence number'. The description restates that dual-form lookup without adding syntax, format, or resolution rules, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (one signed receipt), and enumerates exactly what is returned: payload, hash, previous link's hash, and Ed25519 signature. That return payload also implicitly distinguishes it from the sibling tower_verify_receipts (verification) without needing the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Identifies the two lookup modes (by id or by sequence number) and states the auth prerequisite ('Needs an agent key'), which is genuine usage context. However, it never says when to prefer this over tower_verify_receipts or tower_get_action, so routing guidance remains implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_list_agentsList agents (Overwing Tower)A
Read-only
Inspect

List this organization's Tower agents with scopes, status and last use. Keys are never returned. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description reinforces that with 'Keys are never returned' — a meaningful data-exposure guarantee that annotations alone do not convey. It also discloses the auth requirement (organization key), adding real context beyond the safety hint, though it says nothing about volume or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences; the returned fields come first and the security/auth caveats follow. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list operation with no output schema, the description covers what is returned and the security boundary of keys not being exposed. Only pagination or result-size behavior is left unspecified, a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description's mention of an organization key clarifies the implicit auth context rather than a schema parameter, which is appropriate given the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (Tower agents) and previews the returned fields: scopes, status, last use. It does not name or contrast with a sibling such as atlas_agents or tower_capabilities, so an agent must infer the boundary from the 'Tower' qualifier alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Needs an organization key' gives an implied precondition, which is the only usage guidance present. There is no statement of when to prefer this over atlas_agents, tower_capabilities, or whoami, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_load_templateLoad the starter workflow (Overwing Tower)A
Idempotent
Inspect

Set up Overwing Tower for this organization by loading the starter workflow: email purchase order to order entry, with create_order, update_order and cancel_order against a mock IBM i system, and a starter policy. Idempotent. Needs an organization key. Returns a sample input you can submit. Next: tower_create_agent.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and idempotentHint=true; the description restates idempotency (consistent, no contradiction) and adds meaningful extra context: the exact artifacts created, the environment it touches ('mock IBM i system'), the auth precondition ('Needs an organization key'), and the return value ('Returns a sample input you can submit'). No annotation contradiction, and it goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the setup purpose, then the concrete contents, the idempotency/auth notes, and finally the next step. Dense but each clause earns its place. The single long appositive sentence listing the workflow contents is slightly heavy, keeping it just under a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param setup tool with no output schema, the description covers what is created, the auth precondition, idempotency, and the return value, plus the natural next tool. Little an agent needs to invoke it correctly is missing, though a note on behavior when the org is already set up would complete it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so per the baseline rule this is a 4. The description does hint at an out-of-band input ('Needs an organization key') that is not represented in the schema, which is useful context even though there is nothing to document parameter-wise.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ('Set up Overwing Tower... by loading the starter workflow') and enumerates exactly what is provisioned: email-po-to-order entry, three specific tools (create_order, update_order, cancel_order) against a mock IBM i system, and a starter policy. This is far more specific than its siblings and lets an agent recognize it as the bootstrap/setup entry point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real contextual guidance: it operates 'for this organization', 'Needs an organization key', and explicitly routes forward with 'Next: tower_create_agent'. It does not state when NOT to use it (e.g., against an already-configured org other than via idempotency), so it falls short of a full 5, but the when/where and next-step are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_revoke_agentRevoke an agent (Overwing Tower)A
DestructiveIdempotent
Inspect

Revoke an agent. Its key stops working at once and cannot be restored. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent's id (a UUID), from tower_list_agents

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=true, but the description adds genuinely useful context: the effect is immediate ('Its key stops working at once') and irreversible ('cannot be restored'), plus an auth requirement. It stops short of explaining the response or the practical meaning of idempotency here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and its consequence. Every sentence carries distinct information with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-param tool, annotations cover the safety profile and the schema covers the parameter, leaving the description to add immediacy and irreversibility, which it does. Only the return behavior is unaddressed (no output schema).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema coverage, the schema already documents agent_id fully (including its UUID pattern and provenance from tower_list_agents). The description adds no format or sourcing detail beyond it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Revoke an agent') that is immediately distinguishable from the sibling tower_create_agent. It does not, however, explicitly name or contrast against any alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a prerequisite ('Needs an organization key') but no statement of when to reach for this tool versus siblings like tower_create_agent or tower_compensate. Usage is only implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_submit_actionRequest an action (Overwing Tower)A
Idempotent
Inspect

Ask Tower to perform an operation on the legacy system. Tower decides: executed (done), pending (a person must approve; poll tower_get_action, do not resubmit), or rejected (do not retry unchanged). Always send an idempotency_key that is stable for this business request, such as the source message id: repeating a key returns the original outcome instead of acting twice. Set dry_run to see the decision without executing. Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesThe operation's input, matching its input_schema from tower_capabilities
dry_runNo
operationYes
idempotency_keyYesStable per business request; reuse it on retries

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-readonly and idempotent, and the description goes further by enumerating the three possible outcomes (executed/pending/rejected), explaining idempotency_key reuse semantics ('repeating a key returns the original outcome instead of acting twice'), disclosing the dry_run behavior, and noting an agent key is required. This is rich context well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and outcome model, then idempotency, then dry_run, then the auth requirement. Dense and largely waste-free, though the semicolon-chained pending/rejected clauses make it slightly heavy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and a nested input object, the description covers the outcome states, retry semantics, and dry-run preview, plus the agent-key requirement. The only remaining gap is that valid operation names and input shapes are delegated to tower_capabilities, which is acceptable but worth noting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description must compensate: it clarifies idempotency_key with a concrete example ('such as the source message id') beyond the schema note, and explains dry_run's effect ('see the decision without executing'). Operation and input are left to the schema/other tools, so it is strong but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Ask Tower to perform an operation on the legacy system.' It clearly differs from siblings like tower_get_action (used only for polling) and tower_capabilities (used to discover operations), so an agent can route without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit decision-dependent guidance: on 'pending' poll tower_get_action and do not resubmit; on 'rejected' do not retry unchanged. This is a rare case where the description tells the agent precisely when to call this tool and what to do with each outcome.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tower_verify_receiptsVerify the receipt chain (Overwing Tower)A
Read-only
Inspect

Recompute every hash and check every signature over a range of this organization's receipts. Reports the first break, if any. Defaults to the whole chain (up to 5,000 links per call). Needs an agent key.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNo
fromNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds genuinely useful non-schema behavior: the default scope, the 5,000-link per-call cap, the fail-fast reporting of the first break, and the agent-key requirement. It does not say whether the receipt range is inclusive or what it costs, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler, front-loading the core action, then defaults, then the prerequisite. Every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers the essential return signal ('first break, if any') plus auth and scope limits. It is nearly complete for a 2-param read-only tool; only the precise from/to semantics remain thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description bears the burden for 'from' and 'to'. It implies a range and a default ('defaults to the whole chain') and uses the integer bounds indirectly, but never states which parameter is the start versus end, or whether the bounds are inclusive. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Recompute every hash and check every signature over ... receipts.' It is clearly distinct from read-oriented siblings such as tower_get_receipt, and gives an outcome ('Reports the first break, if any').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides operating context (defaults to whole chain, up to 5,000 links per call, needs an agent key) but never says when to choose this over tower_get_receipt or tower_capabilities. Usage is implied by the verification purpose rather than stated with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWho am IA
Read-only
Inspect

Identify the organization and plan behind the API key on this request, the key's scope, the organization's data settings (store_inputs, retention_days), and whether billing is set up. Needs an organization key.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

with readOnlyHint=true already declared, the description earns credit for disclosing the auth requirement and the specific fields the call reveals (store_inputs, retention_days, billing setup). No note on rate limits or error behavior, but the safety profile is well covered by both annotation and text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the purpose and layers secondary detail afterward. Every clause names a distinct returned fact; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned fields and the auth precondition. It omits response shape/format, but for a parameterless diagnostic tool this is close to sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the schema carries nothing to describe; baseline of 4 applies. The description correctly implies the only input is the ambient request key rather than a parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (identify) and enumerates exactly what is identified: organization, plan, key scope, org data settings, and billing status. An agent can tell it apart from siblings like tower_capabilities or get_usage without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one real precondition — 'Needs an organization key' — which is useful, but gives no when-to-use context relative to siblings or what to do when only a non-org key is available. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • Addedbeacon_report
    • Addedbeacon_sample
    • Addedbeacon_start
  2. 24 tool updates
    • First observedatlas_agents
    • First observedatlas_lookup
    • First observedatlas_summary
    • First observedcreate_rule_set
    • First observedevaluate
    • First observedevaluate_batch
    • First observedget_evaluation
    • First observedget_rule_set
    • First observedget_usage
    • First observedlist_evaluations
    • First observedlist_plans
    • First observedlist_rule_sets
    • First observedtower_capabilities
    • First observedtower_compensate
    • First observedtower_create_agent
    • First observedtower_decide
    • First observedtower_get_action
    • First observedtower_get_receipt
    • First observedtower_list_agents
    • First observedtower_load_template
    • First observedtower_revoke_agent
    • First observedtower_submit_action
    • First observedtower_verify_receipts
    • First observedwhoami

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a pre-flight/post-flight firewall for LLM calls with comprehensive detection, classification, policy enforcement, reversible redaction, output safety, and immutable audit logging.
    1
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Guardrails service for AI agents that evaluates every tool call for safety and alignment before execution, providing default-deny policy, LLM safety evaluation, and audit trail.
    131 PyPI
    22
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides real-time content security for large language models by identifying and intercepting risks across compliance, ethics, and safety dimensions. It enables secure input and output monitoring through a customizable policy engine using an SSE-based interface.
    1
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.